# Beyond 0-100: How Confidence Scale Design Shapes Verbalized Confidence and LLM Metacognition Authors: Yuyang Dai and Yuxia Wang Affiliation: INSAIT Status: Research preprint, 2026 ## Canonical summary This paper studies whether the numerical scale used to elicit confidence from a large language model changes the quality of its verbalized uncertainty. Across six primary LLMs and three multiple-choice datasets, the standard 0-100 scale produces severe confidence discretization. The three most frequent values account for 78.2%-92.1% of responses, and models use only 15-28 of the 101 available integers. Within the evaluated models, datasets, prompts, and decoding conditions, a 0-20 scale yields higher metacognitive efficiency than 0-100. Aggressive lower-bound compression degrades performance, and round-number preferences persist under irregular ranges. ## Scope Primary models: GPT-5.2; Gemini 3.1 Pro; LLaMA-4-Maverick; LLaMA-4-Scout; Qwen3-235B-A22B-Instruct; Qwen3-30B-A3B-Instruct. Additional granularity baseline: Llama-3-8B-Instruct. Datasets: MMLU; GSM8K; TruthfulQA. Measures: Expected Calibration Error; AUROC; meta-d-prime; metacognitive efficiency; round-number preference; range violations. ## Important qualification The results apply to the evaluated settings. Generalization to open-ended generation, other prompting methods, and future model versions is not established. ## Resources - Full paper: paper.pdf - Human-readable project page: index.html - LaTeX source: paper/ - Citation: CITATION.cff and metadata/paper.bib - Structured metadata: metadata/paper.json