Step through the pipeline. Each call builds on the previous role's output.
Hover a dot. Editorial reading of the contrast drawn in the episode.
One teacher family, one student, five benchmarks.
Sample depth is the global-diversity lever. Meta-prompts per node-set is the local lever.
Node combinations (global) × meta-prompts (local). Log scale.
Every sampled node path is the provenance record.
The QDC framing from Havrilla et al.'s survey, drawn.
Each rung isolates one component. Click a row, or climb the ladder.
Gemma 3 4B fine-tuned on Simula data, by dataset size. Curves are mock data shaped to the reported findings.
2% · 9% · 9% · 61%. Hover a bar.
The generating model itself.
Toy model: the critic shares the generator's knowledge, so it misses what the generator misses. Not the paper's numbers.
Taxonomy, critic, complexity scorer, teacher: all Gemini 2.5 Flash.
What was checked, and what was not. Hover cells.
Trust in the taxonomy, critic and complexity verdicts tracks how well the teacher knows the domain.
Also discussed in the episode: Self-Instruct (Wang et al., 2022), Promptbreeder (Fernando et al., 2024), ReAct (Yao et al., 2022), Datasheets for Datasets (Gebru et al., 2018), LEXam (Fan et al., 2025).