Research preprint · LLM calibration · 2026

Beyond 0-100

How Confidence Scale Design Shapes Verbalized Confidence and LLM Metacognition

Yuyang Dai and Yuxia Wang · INSAIT

01 · Direct answer

Does confidence-scale design affect LLM calibration?

Yes, within the evaluated settings. Across six primary LLMs and MMLU, GSM8K, and TruthfulQA, the standard 0-100 scale produces heavily discretized confidence reports. A 0-20 scale yields higher metacognitive efficiency than 0-100, while aggressive lower-bound compression degrades performance.

02 · Evidence

The signal is far less continuous than the scale suggests.

The findings below are bounded to the tested models, datasets, prompts, and decoding conditions.

78.2–92.1%

of reports are covered by only the three most frequent values.

15–28

distinct integers are used out of 101 possible values.

10–29 pts

separate reported confidence from conditional accuracy at the dominant anchor.

Overview of confidence discretization, meta-d-prime measurement, and the three scale manipulations: granularity, boundary shifting, and non-standard ranges.
Study overview: the observed failure mode, the metacognitive measure, and the three controlled scale manipulations.

03 · Baseline

Confidence discretization under 0-100

ModelDominant valueTop-3 coverageValues usedMratio
GPT-5.295 42.1%85.3%210.92
LLaMA-4-Maverick90 35.6%78.2%280.82
Qwen3-235B95 51.2%88.6%190.78
LLaMA-4-Scout90 38.9%81.5%240.76
Gemini 3.1 Pro100 68.4%92.1%170.74
Qwen3-30B100 55.8%90.3%150.62

04 · Experimental design

Three interventions, complementary measures.

G

Granularity

Compare five scales from 0-5 to 0-100 to test the trade-off between resolution and anchor-driven noise.

B

Boundary shifting

Raise the lower bound from 0 to 60 while keeping the upper bound at 100.

N

Irregular ranges

Use non-standard bounds such as 0-73, 14-86, and 3-38 to test semantic adaptation.

Primary models

GPT-5.2, Gemini 3.1 Pro, LLaMA-4-Maverick, LLaMA-4-Scout, Qwen3-235B-A22B-Instruct, and Qwen3-30B-A3B-Instruct.

Benchmarks

MMLU, GSM8K, and TruthfulQA. Llama-3-8B-Instruct is an additional dense baseline in the granularity analysis.

Measures

ECE, AUROC, meta-d-prime, metacognitive efficiency, round-number preference, and range violations.

05 · Questions

Frequently asked

Why do LLM confidence scores cluster around round numbers?

The evaluated models rely heavily on salient numerical anchors such as 90, 95, and 100. Irregular ranges reduce but do not eliminate this behavior.

Is a finer confidence scale always better?

No. Performance is non-monotonic in these experiments: 0-5 is too coarse, while 0-100 does not outperform 0-20 for any evaluated model.

What should practitioners test?

Test 0-20 as an alternative to 0-100, report a discrimination measure alongside ECE, and inspect the empirical confidence distribution before interpreting calibration results.

Can this result be generalized to every LLM task?

No. The experiments focus on three multiple-choice benchmarks. Generalization to open-ended generation, other prompts, and future model versions is not established.

06 · Cite this work

Machine-readable and human-readable citation files are included.

@article{daiwang2026confidencescale,
  title  = {Beyond 0--100: How Confidence Scale Design Shapes
            Verbalized Confidence and LLM Metacognition},
  author = {Dai, Yuyang and Wang, Yuxia},
  year   = {2026},
  note   = {Preprint}
}

CITATION.cff · BibTeX · JSON metadata · LLM-readable summary