Learning, Fast and Slow: LLMs That Adapt Without Forgetting
arXiv:2605.12484Rishabh Tiwari et al., 2026UC Berkeley · Mila · UT AustinFast-Slow Training (FST)
Catastrophic forgetting and plasticity loss are two separate costs of RL
post-training. This paper interleaves slow weight updates (RLVR / CISPO)
with a fast, prompt-evolving channel (GEPA) — reaching RL's peak accuracy
with up to 3× fewer samples, drifting up to 70% less from the base
model, and staying plastic enough to learn the next task where pure RL stalls.