AI Post Transformers · Episode Companion

SkillOpt-Lite: Rethinking Agent Skill Optimization with Zeroth-Order Simplicity

Shen, Li & Zhang (LMMs-Lab, NTU MMLab, Microsoft) reframe agent "skill" editing as zeroth-order optimization, strip SkillOpt down to four steps, and let a smaller model beat a flagship model running the older pipeline — on one benchmark, under one set of assumptions.

📄 arXiv:2607.03451 🗓 Submitted July 3, 2026 also discussed: arXiv:2605.23904 (SkillOpt) also discussed: arXiv:2603.28052 (Meta-harness)

Three Pillars, One Movable Piece

A deployed agent's capability comes from a frozen base LLM, an operational harness (tools, control flow, retries), and skills — plain-text instruction documents read at inference time. Only the skills are cheap to edit.

The SkillOpt-Lite Loop

Four steps, no mini-batch pooling, no epoch-level meta-reflection, no rejected-edit buffer. Step through it below.

References