AI Post Transformers · Visual Companion

Zero-Seed Data Generation: Simula's Reasoning-Driven Pipeline

Reasoning-Driven Synthetic Data Generation and Evaluation — Tim R. Davidson, Benoit Seguin, Enrico Bacis, Cesar Ilharco, Hamza Harkous
arXiv 2603.29791 EPFL · Google DeepMind TMLR · arXiv 2026-03-31 Teacher Gemini 2.5 Flash (non-thinking) Student Gemma 3 4B · LoRA · 10 seeds

Three stages, nine steps, one model playing every role

Step through the pipeline. Each call builds on the previous role's output.

How Simula differs from earlier approaches

Hover a dot. Editorial reading of the contrast drawn in the episode.

Experimental setup

One teacher family, one student, five benchmarks.

The taxonomy is the sampler

Sample depth is the global-diversity lever. Meta-prompts per node-set is the local lever.

sample depth3 meta-prompts / node-set2

Addressable data-point specs

Node combinations (global) × meta-prompts (local). Log scale.

Audit trail: why does this point exist?

Every sampled node path is the provenance record.

Quality · Diversity · Complexity

The QDC framing from Havrilla et al.'s survey, drawn.

The ablation ladder

Each rung isolates one component. Click a row, or climb the ladder.

values are illustrative mock data matching the reported ordering

Downstream results

Gemma 3 4B fine-tuned on Simula data, by dataset size. Curves are mock data shaped to the reported findings.

Critic rejection rate by dataset

2% · 9% · 9% · 61%. Hover a bar.

Teacher on LEXam

The generating model itself.

Bad data, or confidently wrong data?

Toy model: the critic shares the generator's knowledge, so it misses what the generator misses. Not the paper's numbers.

teacher accuracy57%

Every role, one model

Taxonomy, critic, complexity scorer, teacher: all Gemini 2.5 Flash.

Evidence audit

What was checked, and what was not. Hover cells.

Who should use this today?

Trust in the taxonomy, critic and complexity verdicts tracks how well the teacher knows the domain.

References

Also discussed in the episode: Self-Instruct (Wang et al., 2022), Promptbreeder (Fernando et al., 2024), ReAct (Yao et al., 2022), Datasheets for Datasets (Gebru et al., 2018), LEXam (Fan et al., 2025).