From “Can It Find?” to “Can It Learn?”
The key move is separating retrieval from temporary knowledge formation. This diagram places long-context understanding, few-shot in-context learning, and CL-BENCH-style context learning on one visual spectrum.
Benchmark Construction Flow
Each context is designed as a miniature world: legal code, product manual, lab notebook, or synthetic rule system. Tasks then probe whether the model can build a working theory of that world.
Task Mix
CL-BENCH varies the kind of reasoning being demanded, not just the topic area. Hover segments to inspect the distribution used in this illustrative mockup.
Binary Rubrics Turn Partial Understanding Into Visible Friction
The benchmark grades dense answer requirements, so missing one mandated condition can zero out a task. Use the toggle to compare single-turn and sequential settings across rubric families.
Why heatmaps here?
What the colors mean
What to inspect
Strict Solve Rate vs Partial Credit
The all-or-nothing headline is harsh by design. This chart uses realistic illustrative data to show how the story changes when partial rubric credit is separated from perfect task completion.
Benchmark Family Comparison
CL-BENCH pushes farther from literal retrieval than earlier long-context suites. The farther right, the more the test demands task-local rule induction rather than document lookup.
Position Stress
“Lost in the Middle” style effects remain relevant. This synthetic curve shows why larger windows alone do not guarantee stable use of buried evidence.
Failure Propagation Across Sequential Tasks
In sequential contexts, an early rule misunderstanding can infect later steps. Use the slider to watch how a small miss at step 1 amplifies through carryover, retrieval pressure, and rubric completeness.
What a Better Harness Changes
The paper tests prompt-only models. External scaffolding can still reduce some failures without proving the model learned the local world on its own.
Podcast Cross-Links
This episode connects to prior discussions on induction heads, context rot, million-token systems, test-time adaptation, and divide-and-conquer reasoning.
References
Primary paper plus the immediate benchmark and in-context learning lineage discussed in the episode. Additional arXiv IDs found in transcript text: 2602.03587.