Self-Calibrating Language Models via Test-Time Discriminative Distillation. Language models are confidently wrong. SECL teaches a model to trust its own judgment instead of its own bravado, at test time, with no labels. Scroll to see how, then run the mechanism on any model you like.
This is systematic overconfidence, and alignment training tends to make it worse by rewarding agreement over truthfulness. In a review of 519 healthcare LLM studies, 95% measured accuracy and only 1.2% measured calibration. Confidence you cannot trust is a safety problem.
A model that cannot reliably produce the right answer can often still recognize a wrong one. Judging is easier than generating. Theory backs this up: generative error is lower-bounded by roughly twice the matching discriminative error.
That gap is free supervision. The token probability of "True" when you ask the model "is this answer correct?" is better calibrated than the confidence it states out loud. SECL distills the reliable signal into the unreliable one. Flip the tabs:
A reliability diagram plots stated confidence against actual accuracy. Perfect calibration sits on the diagonal: when the model says 70%, it is right 70% of the time. The gap between the bars and the diagonal, averaged, is Expected Calibration Error (ECE).
SECL only acts when the world changes. An entropy gate watches the question stream and triggers a short calibration burst on a distribution shift. Inside a burst, each question runs this loop:
The update target is ĉ = c + α·clip(c*−c, −δ, δ) with α=0.5 and δ=0.15. Conservative, bounded steps toward the better signal. Only ~0.01–0.02% of weights (LoRA on late layers) move, and task accuracy is preserved within a point.
Pick a model, paste a key, and watch the mechanism work on real outputs. For each question SECL gets the model's answer and stated confidence, scores that answer against distractors to get NormP(True), and nudges the confidence toward it. Points land on a live reliability diagram so you can watch ECE fall.
| Question | Model | OK? | c | c* | Gate | ĉ |
|---|---|---|---|---|---|---|
| No runs yet. | ||||||
Tip: OpenRouter and Google Gemini both work cleanly from the browser. If a provider blocks requests with a CORS error, switch to OpenRouter. The gap is widest on TruthfulQA-style misconception questions, where models are most overconfident.
Across four small models and four domains, SECL cuts ECE by 56–78% while preserving task accuracy, training on only 6–26% of the stream and at lower cost than the signal it learns from.
| Model | Verbalized ECE | SECL ECE | Reduction | Cost (FWD-eq) |
|---|---|---|---|---|
| Llama 3.2-3B | 0.170 | 0.050 | −71% | 1.8–4.6 |
| Llama 3.1-8B | 0.225 | 0.083 | −63% | 2.8 |
| Gemma 2-2B | 0.256 | 0.056 | −78% | 1.8 |
| Phi 3.5-Mini | 0.251 | 0.110 | −56% | 2.1 |
SECL is the first method to apply test-time training to calibration. It surpasses its own supervision signal, stays cheaper than the P(True) Norm baseline, and needs no labels. The same principle generalizes: wherever a model can evaluate better than it can generate, that gap can be distilled back into its outputs.