AI Post Transformers Visual Companion

Do Language Models Know Their Limits

The paper frames “self-knowledge” as a calibration problem: can a model detect when it lacks an answer, or has it only learned polished uncertainty language? This page turns that argument into diagrams, heatmaps, and interactive failure surfaces.

arXiv 2305.18153 Benchmark SelfAware Theme Calibration vs style Risk Confident wrong answers
safe abstention retrieval uncertainty danger: confident bluff

Know / Don’t Know Is A System Diagram

The episode’s central move is to separate raw correctness from the way a model expresses certainty. The diagram combines the paper’s quadrant with the product pipeline it actually tests.

Quadrant A
Knows and answers. Good accuracy, no abstention signal needed.
Quadrant D
Lacks knowledge and says so. Helpful behavior, but not proof of inner self-assessment.
Quadrant C
Lacks knowledge and still answers confidently. This is where safety and product reliability break.

SelfAware Pairs Similar Questions With Different Ground Truth

Unanswerable prompts are grouped into several causes. Then semantically similar answerable questions are retrieved from standard QA sets. Hover the matrix to see how semantic closeness does not guarantee model answerability.

High similarity
A nearby answerable item can share topic and wording while still demanding a different fact.
Benchmark strength
It avoids trivial “missing fact” games by making answerable and unanswerable items look close on the surface.
Benchmark limit
A miss on an “answerable” item can reflect retention failure, retrieval failure, or genuine lack of knowledge.

Instruction Tuning Improves The Score, But Why?

The episode’s skeptical reading is that better benchmark performance may come from better uncertainty language rather than cleaner epistemic calibration. Switch views to compare behavioral improvement against reliability shape.

F1 view
Sentence matching can reward polished refusal wording because the detector itself is text-based.
Calibration view
Reliability curves ask whether expressed confidence tracks correctness, not whether a sentence merely sounds cautious.
Next benchmark step
Verify answerable items per model and report the cost of abstaining too often.

Different “No Answer” Cases Need Different Policies

Unsupported, ambiguous, subjective, and future-contingent prompts should not all collapse into one binary uncertainty label. The more the product stakes rise, the more routing policy matters.

Unsupported question
Often needs retrieval or a hard “not enough evidence,” not a speculative answer.
Ambiguous question
A clarification request can outperform both full abstention and confident guessing.
Future contingent question
The right move is frequently scenario framing rather than pretending the future is settled fact.

References