ASI-Evolve for Data,
Architectures, and RL

A visual companion for a podcast episode about whether one agentic loop can improve AI itself across three hard targets: pretraining data, model architectures, and RL algorithm design.
Episode focus: AI-for-AI Paper: Xu et al., 2026 arXiv:2603.29640 Episode viz link

One outer loop, three inner objects

Traditional AutoML usually tunes within a bounded setup. This diagram shows the episode’s core distinction: ASI-Evolve is framed as automating chunks of the research loop itself.

cognition / priors proposal + mutation expensive evaluation analysis + memory update

Why this is harder than benchmark agents

Interactive heatmap: move across columns to see how feedback quality degrades as you go from toy tasks to frontier ML research.

Axes show mock but realistic dimensions discussed in the episode: feedback delay, cost, ambiguity, variance, and proxy mismatch.

Switch the search object

The same orchestration loop points at very different editable artifacts: code for architectures, code for dataset pipelines, and code for RL updates.

Action-space matrix

Hover cells to inspect how “unified” can still hide thick domain scaffolding. Brighter cells indicate stronger dependence on domain-specific interfaces.

Reported results dashboard

Toggle between absolute scores and deltas. Values are mock reconstructions anchored to the episode’s reported magnitudes, built to show the shape of the claims rather than exact paper tables.

Architecture search trajectory

Mock search trace of 1,773 rounds. Hover points to see a few candidate checkpoints. Dense local winners can mean fertile search space—or benchmark gaming.

Skepticism map

Click a lens to reweight the risk profile. The episode repeatedly returns to one question: is this a genuine general research engine, or a shared shell wrapped around domain engineering?

Failure-mode heatmap

Common ways AI-for-AI claims can look stronger than they are. This panel converts the episode’s objections into an interactive matrix.

References in this companion

  1. Xu et al. (2026), ASI-Evolve: AI Accelerates AI
  2. He, Zhao, Chu (2021), AutoML survey
  3. Yang et al. (2024), Large Language Models as Optimizers
  4. Lu et al. (2024), The AI Scientist
  5. Novikov et al. (2025), AlphaEvolve
  6. Zoph & Le (2017), Neural Architecture Search with Reinforcement Learning
  7. Real et al. (2019), Regularized Evolution
  8. Liu, Simonyan, Yang (2019), DARTS
  9. Tan & Le (2019), EfficientNet
  10. Gao et al. (2020), The Pile
  11. Hoffmann et al. (2022), Chinchilla / compute-optimal training
  12. FineWeb (2024), Hugging Face collaborators
  13. DCLM / DataComp-LM (2024)
  14. Eysenbach et al. (2021), Discovering Reinforcement Learning Algorithms
  15. Schulman et al. (2017), PPO
  16. DeepSeek-AI (2024), DeepSeekMath / GRPO
  17. Hendrycks et al. (2021), MMLU
Additional arXiv IDs extracted from the transcript using the requested pattern: 2603.29640 only.