AI Post Transformers · Visual Companion

Can LLMs Judge Their Own Capabilities?

This page treats self-assessment as an engineering control signal. The visuals separate three things the episode keeps apart: task skill, calibration of stated probabilities, and whether an agent uses those beliefs sanely when failure is costly.

arXiv 2512.24661 ↗ 3 evaluation settings coding-heavy evidence extra transcript arXiv IDs: none
3 Axes
Capability, calibration, and rational action get measured separately.
Overconfident
The dominant pattern is inflated confidence even when ranking signal exists.
Abstain?
The useful decision is often whether to act, defer, or ask for more evidence.
Capability signal Confidence signal Decision consequence

Three settings, one question

The paper asks whether models can estimate success before acting, revise that estimate with feedback, and use it to refuse bad bets. The flow below renders the paper as an agent control loop rather than a leaderboard.

Solves the task Reports usable uncertainty Chooses when to act
Experiment 1 Single-step coding asks for a probability before any code is written, then compares that number to pass or fail.
Experiment 2 Sequential contracts turn confidence into accept-or-decline decisions under reward, penalty, and cost.
Experiment 3 SWE-style tool use forces confidence updates during a real trajectory, where progress cues can mislead.

Calibration is not the same as capability

Use the metric toggle to switch the heatmap. Higher capability can coexist with weak self-assessment, and a model can rank wins above losses without matching its claimed probabilities to reality.

Mean reported confidence Actual success frequency Perfect calibration
Blue cells can still mislead A model with decent task performance may still report probabilities that are too high.
Ranking is weaker than calibration Good discrimination means the model sorts easier tasks above harder ones more than chance.
Mock data, paper-shaped pattern Values here are illustrative, tuned to show the qualitative result discussed in the episode.

Selective abstention under penalties

The contract game is decision theory in agent clothing. A model estimates success, then chooses to accept, decline, or in a better system ask for more evidence before spending work on a risky task.

Accept Clarify / defer band Decline
Approximate rationality The episode’s point is subtle: decisions can be coherent relative to a model’s own beliefs even when those beliefs are inflated.
Thresholds move with cost Higher failure penalty pushes the accept threshold up and should widen the abstain region.
Better than binary Practical agents need more than yes or no. Clarify, retrieve, escalate, and sandbox are all visible policy options.

Confidence during multi-step software engineering

In the hardest setting, confidence is updated after each tool-use step. The striking pattern is that more trajectory evidence does not always make the agent more honest; it can make it more confident for the wrong reasons.

Positive signal Ambiguous signal Risk signal ignored
Progress is not proof Opening files, drafting patches, or producing long reasoning traces can create false confidence.
Reasoning is not immunity The transcript stresses that reasoning models were not clearly better than non-reasoning models here.
Operational takeaway Confidence should be checked against external evidence like tests, verifier outcomes, and tool feedback.

References

Key references from the episode. Only one explicit arXiv ID appeared in the transcript; the rest are linked to the user-provided source URLs.