← All episodes MetaClaw: Just Talk and Continual Agent Adaptation

MetaClaw: Just Talk and Continual Agent Adaptation

Mar 28, 2026
This episode explores a paper on long-lived AI agents that keep adapting to changing real-world tasks without being taken offline. It explains the paper’s central idea of combining two learning timescales: fast updates through an evolving skill library and slower policy improvement through parameter-efficient weight tuning such as LoRA. The discussion unpacks why agent learning is harder than ordinary one-shot language modeling, since failure happens across whole action trajectories involving tools, recovery strategies, and multi-step decisions. Listeners would find it interesting because the episode connects this proposal to broader debates about memory retrieval, skill libraries, and continual meta-learning, while questioning whether dynamic skill evolution alone can already deliver substantial behavioral improvement.
Sources:
1. MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild — Peng Xia, Jianwen Chen, Xinyu Yang, Haoqin Tu, Jiaqi Liu, Kaiwen Xiong, Siwei Han, Shi Qiu, Haonian Ji, Yuyin Zhou, Zeyu Zheng, Cihang Xie, Huaxiu Yao, 2026
http://arxiv.org/abs/2603.17187
2. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks — Chelsea Finn, Pieter Abbeel, Sergey Levine, 2017
https://scholar.google.com/scholar?q=Model-Agnostic+Meta-Learning+for+Fast+Adaptation+of+Deep+Networks
3. Meta-Learning with Sparse Experience Replay for Lifelong Language Learning — Nithin Holla, Pushkar Mishra, Helen Yannakoudakis, Ekaterina Shutova, 2020
https://scholar.google.com/scholar?q=Meta-Learning+with+Sparse+Experience+Replay+for+Lifelong+Language+Learning
4. La-MAML: Look-ahead Meta Learning for Continual Learning — Gunshi Gupta, Karmesh Yadav, Liam Paull, 2020
https://scholar.google.com/scholar?q=La-MAML%3A+Look-ahead+Meta+Learning+for+Continual+Learning
5. When Meta-Learning Meets Online and Continual Learning: A Survey — Jaehyeon Son, Soochan Lee, Gunhee Kim, 2025
https://scholar.google.com/scholar?q=When+Meta-Learning+Meets+Online+and+Continual+Learning%3A+A+Survey
6. Reflexion: Language Agents with Verbal Reinforcement Learning — Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, Shunyu Yao, 2023
https://scholar.google.com/scholar?q=Reflexion%3A+Language+Agents+with+Verbal+Reinforcement+Learning
7. Voyager: An Open-Ended Embodied Agent with Large Language Models — Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, Anima Anandkumar, 2023
https://scholar.google.com/scholar?q=Voyager%3A+An+Open-Ended+Embodied+Agent+with+Large+Language+Models
8. ExpeL: LLM Agents Are Experiential Learners — Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, Gao Huang, 2023
https://scholar.google.com/scholar?q=ExpeL%3A+LLM+Agents+Are+Experiential+Learners
9. Solving Math Word Problems with Process- and Outcome-Based Feedback — Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, Irina Higgins, 2022
https://scholar.google.com/scholar?q=Solving+Math+Word+Problems+with+Process-+and+Outcome-Based+Feedback
10. Let's Verify Step by Step — Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, Karl Cobbe, 2023
https://scholar.google.com/scholar?q=Let%27s+Verify+Step+by+Step
11. Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations — Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, Zhifang Sui, 2024
https://scholar.google.com/scholar?q=Math-Shepherd%3A+Verify+and+Reinforce+LLMs+Step-by-step+without+Human+Annotations
12. Process Reward Models for LLM Agents: Practical Framework and Directions — Sanjiban Choudhury, 2025
https://scholar.google.com/scholar?q=Process+Reward+Models+for+LLM+Agents%3A+Practical+Framework+and+Directions
13. ReAct: Synergizing Reasoning and Acting in Language Models — Shunyu Yao, Jeffrey Zhao, Dian Yu, et al., 2022
https://scholar.google.com/scholar?q=ReAct%3A+Synergizing+Reasoning+and+Acting+in+Language+Models
14. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, et al., 2021
https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models
15. Proximal Policy Optimization Algorithms — John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, Oleg Klimov, 2017
https://scholar.google.com/scholar?q=Proximal+Policy+Optimization+Algorithms
16. Trajectory-Informed Memory Generation for Self-Improving Agent Systems — Gaodan Fang et al., 2026
https://scholar.google.com/scholar?q=Trajectory-Informed+Memory+Generation+for+Self-Improving+Agent+Systems
17. Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models — Rishabh Tiwari et al., 2026
https://scholar.google.com/scholar?q=Reward+Under+Attack%3A+Analyzing+the+Robustness+and+Hackability+of+Process+Reward+Models
18. Process Reinforcement through Implicit Rewards — Ganqu Cui et al., 2025
https://scholar.google.com/scholar?q=Process+Reinforcement+through+Implicit+Rewards
19. Reinforcement Learning for Self-Improving Agent with Skill Library — Jiongxiao Wang et al., 2025
https://scholar.google.com/scholar?q=Reinforcement+Learning+for+Self-Improving+Agent+with+Skill+Library
20. When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail — Xiaoxiao Li, 2026
https://scholar.google.com/scholar?q=When+Single-Agent+with+Skills+Replace+Multi-Agent+Systems+and+When+They+Fail
21. AI Post Transformers: Agentic Context Engineering: Evolving Contexts for Self-Improving LLMs — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/agentic-context-engineering-evolving-contexts-for-self-improving-llms/
22. AI Post Transformers: Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models — Hal Turing & Dr. Ada Shannon, Wed,
https://podcast.do-not-panic.com/episodes/agentic-context-engineering-evolving-contexts-for-self-improving-language-models/
23. AI Post Transformers: Experiential Reinforcement Learning: Internalizing Reflection for Better Policy Training — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/experiential-reinforcement-learning-internalizing-reflection-for-better-policy-t/
24. AI Post Transformers: Metacognition and Skill Discovery in LLM Math Reasoning — Hal Turing & Dr. Ada Shannon, Mon,
https://podcast.do-not-panic.com/episodes/metacognition-and-skill-discovery-in-llm-math-reasoning/
25. AI Post Transformers: NeurIPS 2025: SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data — Hal Turing & Dr. Ada Shannon, Sat,
https://podcast.do-not-panic.com/episodes/neurips-2025-serl-self-play-reinforcement-learning-for-large-language-models-wit/
26. AI Post Transformers: MATTRL: Collaborative Test-Time Reinforcement Learning for Multi-Agent Reasoning — Hal Turing & Dr. Ada Shannon, Mon,
https://podcast.do-not-panic.com/episodes/mattrl-collaborative-test-time-reinforcement-learning-for-multi-agent-reasoning/
27. AI Post Transformers: A Framework for LLM Application Safety Evaluation — Hal Turing & Dr. Ada Shannon, Mon,
https://podcast.do-not-panic.com/episodes/a-framework-for-llm-application-safety-evaluation/