Writing
September 17, 2026 · 7 min read

Confidence should come from experience, not from resampling the present

A new paper, XConf, estimates an LLM's confidence by retrieving graded past episodes instead of resampling the current answer, matching or beating 10-sample self-consistency at a tenth of the cost and lifting selective-prediction accuracy by up to 8.7 points on agent tasks.

llm-evaluationconfidence-calibrationai-agentsretrieval-augmented-generationresearch

The idea: confidence needs a track record

Every confidence estimator I've deployed in production shares the same blind spot: it only looks at the current inference. Ask the model to introspect on its own chain of thought, score token probabilities, or resample the same prompt ten times and vote — all of these methods are betting that the present attempt contains enough signal to judge itself. A new paper out of Cambridge, "Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents" by Caiqi Zhang, Xiaochen Zhu, Chengzu Li, Yulong Chen, Dharshan Kumaran, and Nigel Collier, argues that premise is the problem. A single inference, however carefully you interrogate it, is not a sufficient basis for confidence. What's missing is memory of how similar attempts actually turned out.

The method is called XConf, for eXperiential Confidence. Instead of squeezing more signal out of the current generation, it gives the model a bank of its own past episodes to consult. Each episode records five things: the task, the model's reflection on how it approached it, the confidence it stated at the time, the graded outcome, and a lesson written once that grade came in. That last field matters — it's not just a log of right and wrong, it's a distilled note to the model's future self about why it was right or wrong.

When a new task arrives, XConf runs in two stages:

  • Recall. Retrieve past episodes that resemble the new task and that were met with a similarly-stated confidence level, then read off their historical success rate. This is the part that resampling can never give you: an empirical base rate for "how often was I actually right when I felt this sure about tasks like this one."
  • Reflect. Show the model this retrieved record, have it name its recurring failure mode on that class of task, and let it restate a confidence that's now informed by its own history rather than just its current chain of thought.

The image below shows the flow end to end, alongside the cost comparison that I think is the more practically important result.

XConf pipeline: new task retrieves similar graded episodes, reads their historical success rate, reflects on the recurring failure mode, and restates a calibrated confidence, at one generation versus ten for self-consistency

What I like about this design is that it's format-general by construction. It doesn't need logit access, it doesn't need weight updates, and it costs exactly one answer generation — the same one you were already going to produce. Everything else is retrieval and a second pass of reasoning over retrieved text, which works the same whether you're calling a closed API or running local weights.

Why self-consistency was never going to be enough

Ten-sample self-consistency — generate the same answer ten times, take the majority vote, treat agreement as confidence — has been the default cheap calibration trick for a while now. It's popular because it's simple and it doesn't need any infrastructure beyond a temperature setting. But it has a structural weakness: if the model has a systematic bias on a task type, resampling ten times mostly reproduces that bias ten times. Agreement across samples measures the model's internal consistency, not its correctness. A model can be very consistently wrong.

Experience-based retrieval sidesteps that because the signal comes from outside the current inference — from what actually happened the last N times the model faced something like this. That's a genuinely different source of information, and it's why the paper frames this as a new paradigm for confidence estimation rather than an incremental tweak to introspection or resampling.

The results

The authors evaluate across nine benchmarks — spanning reasoning, coding, multimodal QA, and interactive agent tasks — and four models across three model families. Two numbers stood out to me:

  • XConf beats or matches 10-sample self-consistency on discrimination (AUROC) in 23 of 24 comparisons, with substantially lower calibration error (ECE), while using roughly one-tenth the generation cost.
  • Used for selective prediction — abstaining on the 10% of episodes where the model is least confident — XConf raises delivered success rate by up to 8.7 points on agent tasks.

That second number is the one I'd pay attention to if you're running anything agentic in production. Agent tasks are exactly where a single bad step compounds into a wasted trajectory, and knowing when to stop and escalate to a human — or to a slower, more expensive path — is often worth more than raw accuracy. An 8.7-point lift from abstaining on the worst 10% is a large, cheap win if the confidence signal driving that abstention is trustworthy.

My take

The part of this paper that I find most convincing isn't the benchmark table, it's the framing: current-inference-only confidence estimation was always going to plateau, because it's trying to extract more information than a single forward pass actually contains. Retrieval-augmented generation solved a parallel problem for factual grounding by giving models access to external knowledge; XConf is doing the same move for self-knowledge, giving the model access to its own graded history instead of asking it to reason its way to calibration from scratch each time.

The practical cost of adopting something like this is real but manageable: you need infrastructure to grade outcomes, write lessons, and index episodes for retrieval — which is more plumbing than a temperature-sampling loop. But if you're already running evals or human review on production outputs, you likely have the graded-outcome half of that pipeline sitting unused. The gap between "we log outcomes" and "we retrieve outcomes to calibrate the next decision" is smaller than it looks, and this paper is a solid argument for closing it, especially anywhere you're using confidence to decide what to autonomously ship versus what to send back for a second look.

References
  1. 01Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents (arXiv:2609.17708)