of reports are covered by only the three most frequent values.
Research preprint · LLM calibration · 2026
Beyond 0-100
How Confidence Scale Design Shapes Verbalized Confidence and LLM Metacognition
01 · Direct answer
Does confidence-scale design affect LLM calibration?
Yes, within the evaluated settings. Across six primary LLMs and MMLU, GSM8K, and TruthfulQA, the standard 0-100 scale produces heavily discretized confidence reports. A 0-20 scale yields higher metacognitive efficiency than 0-100, while aggressive lower-bound compression degrades performance.
02 · Evidence
The signal is far less continuous than the scale suggests.
The findings below are bounded to the tested models, datasets, prompts, and decoding conditions.
distinct integers are used out of 101 possible values.
separate reported confidence from conditional accuracy at the dominant anchor.
03 · Baseline
Confidence discretization under 0-100
| Model | Dominant value | Top-3 coverage | Values used | Mratio |
|---|---|---|---|---|
| GPT-5.2 | 95 42.1% | 85.3% | 21 | 0.92 |
| LLaMA-4-Maverick | 90 35.6% | 78.2% | 28 | 0.82 |
| Qwen3-235B | 95 51.2% | 88.6% | 19 | 0.78 |
| LLaMA-4-Scout | 90 38.9% | 81.5% | 24 | 0.76 |
| Gemini 3.1 Pro | 100 68.4% | 92.1% | 17 | 0.74 |
| Qwen3-30B | 100 55.8% | 90.3% | 15 | 0.62 |
04 · Experimental design
Three interventions, complementary measures.
Granularity
Compare five scales from 0-5 to 0-100 to test the trade-off between resolution and anchor-driven noise.
Boundary shifting
Raise the lower bound from 0 to 60 while keeping the upper bound at 100.
Irregular ranges
Use non-standard bounds such as 0-73, 14-86, and 3-38 to test semantic adaptation.
Primary models
GPT-5.2, Gemini 3.1 Pro, LLaMA-4-Maverick, LLaMA-4-Scout, Qwen3-235B-A22B-Instruct, and Qwen3-30B-A3B-Instruct.
Benchmarks
MMLU, GSM8K, and TruthfulQA. Llama-3-8B-Instruct is an additional dense baseline in the granularity analysis.
Measures
ECE, AUROC, meta-d-prime, metacognitive efficiency, round-number preference, and range violations.
05 · Questions
Frequently asked
Why do LLM confidence scores cluster around round numbers?
The evaluated models rely heavily on salient numerical anchors such as 90, 95, and 100. Irregular ranges reduce but do not eliminate this behavior.
Is a finer confidence scale always better?
No. Performance is non-monotonic in these experiments: 0-5 is too coarse, while 0-100 does not outperform 0-20 for any evaluated model.
What should practitioners test?
Test 0-20 as an alternative to 0-100, report a discrimination measure alongside ECE, and inspect the empirical confidence distribution before interpreting calibration results.
Can this result be generalized to every LLM task?
No. The experiments focus on three multiple-choice benchmarks. Generalization to open-ended generation, other prompts, and future model versions is not established.
06 · Cite this work
Machine-readable and human-readable citation files are included.
@article{daiwang2026confidencescale,
title = {Beyond 0--100: How Confidence Scale Design Shapes
Verbalized Confidence and LLM Metacognition},
author = {Dai, Yuyang and Wang, Yuxia},
year = {2026},
note = {Preprint}
}
CITATION.cff · BibTeX · JSON metadata · LLM-readable summary