Inference Pipeline Diagram
Hover the stages. The key transition is where demonstration selection becomes curriculum design.
A visual tour of the paper’s central tension: long prompts are not just retrieval containers. Under the right model, task, ordering, and demonstration style, they can behave more like a temporary lesson plan executed at inference time.
The transcript frames prompt construction as lesson sequencing rather than nearest-neighbor search. The visual sections below turn that claim into diagrams, heatmaps, and mock experimental dashboards.
Instead of treating demonstrations as isolated neighbors, the paper’s framing says the prompt can become a structured workspace. This view only pays off when the model can actually use the reasoning traces in sequence.
Hover the stages. The key transition is where demonstration selection becomes curriculum design.
Rows show design choices. Columns show the paper’s claimed consequences when prompts become long.
The paper argues that more demonstrations are not uniformly better. The visual below contrasts reasoning-friendly and non-reasoning settings while exposing growing order sensitivity as prompt length increases.
Toggle between outcome and instability. Mock data mirrors the episode’s reported directional pattern.
Cells estimate how much performance swings across random demo orderings as demonstration count rises.
CDS is the paper’s answer to conceptual whiplash. Rather than choosing only nearby examples, it tries to create a smooth route through reasoning space.
The left path shows semantic nearest-neighbor hopping. The highlighted curve shows a staged curriculum.
Best-case improvement from smoother ordering becomes most visible at long prompt length.
Two knobs shape the mock prompt: demonstration count and procedural compatibility. The token map shows why “more” can become clutter instead of instruction.
Hover cells to inspect where the context window is spending tokens and which spans are actually teachable.
A compact scorecard converts the visual budget into expected accuracy, latency, and fragility.