This page treats self-assessment as an engineering control signal. The visuals separate three things the episode keeps apart: task skill, calibration of stated probabilities, and whether an agent uses those beliefs sanely when failure is costly.
The paper asks whether models can estimate success before acting, revise that estimate with feedback, and use it to refuse bad bets. The flow below renders the paper as an agent control loop rather than a leaderboard.
Use the metric toggle to switch the heatmap. Higher capability can coexist with weak self-assessment, and a model can rank wins above losses without matching its claimed probabilities to reality.
The contract game is decision theory in agent clothing. A model estimates success, then chooses to accept, decline, or in a better system ask for more evidence before spending work on a risky task.
In the hardest setting, confidence is updated after each tool-use step. The striking pattern is that more trajectory evidence does not always make the agent more honest; it can make it more confident for the wrong reasons.
Key references from the episode. Only one explicit arXiv ID appeared in the transcript; the rest are linked to the user-provided source URLs.