Toggle the teacher setup to see how visible multi-agent structure changes the same core loop: propose, attack, revise, filter, distill.
Outcome-only supervision, trajectory augmentation, and process-aware distillation differ mainly in how much of the teacher process becomes trainable signal.
Mock quantitative summary inspired by the episode: all methods help over a plain student; PAD usually wins, especially on harder reasoning and robustness tasks.
This panel separates supported engineering claims from unresolved causal claims about whether the student truly internalizes something uniquely multi-agent.