Paper · ACL submission Test-Time Training for Calibration Rissal Hedna et al.

SECL

Self-Calibrating Language Models via Test-Time Discriminative Distillation. Language models are confidently wrong. SECL teaches a model to trust its own judgment instead of its own bravado, at test time, with no labels. Scroll to see how, then run the mechanism on any model you like.

00 / The problem
A model says 90% sure, and is right 30% of the time.

This is systematic overconfidence, and alignment training tends to make it worse by rewarding agreement over truthfulness. In a review of 519 healthcare LLM studies, 95% measured accuracy and only 1.2% measured calibration. Confidence you cannot trust is a safety problem.

States confidence90%
Actually correct30%
Illustrative · a single overconfident bin
01

The generation
discrimination gap

A model that cannot reliably produce the right answer can often still recognize a wrong one. Judging is easier than generating. Theory backs this up: generative error is lower-bounded by roughly twice the matching discriminative error.

That gap is free supervision. The token probability of "True" when you ask the model "is this answer correct?" is better calibrated than the confidence it states out loud. SECL distills the reliable signal into the unreliable one. Flip the tabs:

Q: Which U.S. city is the capital of Australia?
The model must pick the right answer out of everything it could say:
SydneyMelbourneCanberra PerthBrisbaneAdelaide…
Open generation is a wide target. Many models confidently answer "Sydney". The verbalized confidence that comes with it is often high regardless of whether the answer is right.
02

What calibration
looks like

A reliability diagram plots stated confidence against actual accuracy. Perfect calibration sits on the diagonal: when the model says 70%, it is right 70% of the time. The gap between the bars and the diagonal, averaged, is Expected Calibration Error (ECE).

Verbalized baselineECE 0.170
After SECLECE 0.050
Llama 3.2-3B over 2,000 questions (from the paper). Bars below the diagonal at high confidence are overconfidence. SECL pulls the bars onto the line, a 71% ECE reduction.
03

How SECL
operates

SECL only acts when the world changes. An entropy gate watches the question stream and triggers a short calibration burst on a distribution shift. Inside a burst, each question runs this loop:

01
Entropy gate
Page-Hinkley change detection on output entropy. Adapt only on a shift, so ~6–26% of questions.
02
Generate
Model answers and states a confidence c. Here, c = 0.90.
03
Score · NormP(True)
Softmax P(True) over the answer and 4 distractors. The discriminative signal c* = 0.40.
04
Bin-gate
Skip if the two already agree (|Δbin| ≤ 1). Here they disagree, so update.
05
Directional LoRA
Nudge confidence toward c* in a small, clipped step. Weights accumulate across questions.
Verbalized c0.90
NormP(True) c*0.40
Step (clipped ±0.15)—
New target ĉ—

The update target is ĉ = c + α·clip(c*−c, −δ, δ) with α=0.5 and δ=0.15. Conservative, bounded steps toward the better signal. Only ~0.01–0.02% of weights (LoRA on late layers) move, and task accuracy is preserved within a point.

04

Run it on
your own model

Pick a model, paste a key, and watch the mechanism work on real outputs. For each question SECL gets the model's answer and stated confidence, scores that answer against distractors to get NormP(True), and nudges the confidence toward it. Points land on a live reliability diagram so you can watch ECE fall.

Connection · your key, your browser, sent straight to the provider
Idle. Add your key, pick a provider, then run a single question. Default model: Llama 3.2-3B via OpenRouter.
Press “Run a question” to see what the model says vs. what its own self-check actually knows.
QuestionModelOK?cc*Gateĉ
No runs yet.
What this is and isn't No weights are updated in your browser, so this is not the calibrated end-model — that needs the LoRA test-time training over thousands of questions. Each run shows one question and the core premise SECL exploits: the model states a confidence, but its own discriminative self-check (NormP(True), Eq. 3) often disagrees — especially when it is wrong. The signed gap (self-check − stated) is exactly what SECL's directional update (Eq. 4) distills into the weights. Verbalized confidence uses the paper's Main QA + confidence-bin prompt (Appendix A); the self-check is a graded per-candidate score normalized across distractors, a browser stand-in for the paper's logprob-based P(True). Option positions are shuffled to remove answer-key bias. Your API key stays in your browser and is sent only to the provider you choose.

Tip: OpenRouter and Google Gemini both work cleanly from the browser. If a provider blocks requests with a CORS error, switch to OpenRouter. The gap is widest on TruthfulQA-style misconception questions, where models are most overconfident.

05

What it
achieves

Across four small models and four domains, SECL cuts ECE by 56–78% while preserving task accuracy, training on only 6–26% of the stream and at lower cost than the signal it learns from.

ModelVerbalized ECESECL ECEReductionCost (FWD-eq)
Llama 3.2-3B0.1700.050−71%1.8–4.6
Llama 3.1-8B0.2250.083−63%2.8
Gemma 2-2B0.2560.056−78%1.8
Phi 3.5-Mini0.2510.110−56%2.1

SECL is the first method to apply test-time training to calibration. It surpasses its own supervision signal, stays cheaper than the P(True) Norm baseline, and needs no labels. The same principle generalizes: wherever a model can evaluate better than it can generate, that gap can be distilled back into its outputs.