SkillOpt-Lite: Rethinking Agent Skill Optimization with Zeroth-Order Simplicity
Shen, Li & Zhang (LMMs-Lab, NTU MMLab, Microsoft) reframe agent "skill" editing as
zeroth-order optimization, strip SkillOpt down to four steps, and let a smaller model beat a flagship
model running the older pipeline — on one benchmark, under one set of assumptions.
A deployed agent's capability comes from a frozen base LLM, an operational harness (tools, control
flow, retries), and skills — plain-text instruction documents read at inference time. Only the
skills are cheap to edit.
The SkillOpt-Lite Loop
Four steps, no mini-batch pooling, no epoch-level meta-reflection, no rejected-edit buffer. Step
through it below.
Table 1 — Skill Editing as Zeroth-Order Optimization
Five prior techniques, each mapped onto a classical zeroth-order optimization concept. Click a row to
see the concept it borrows from.
Benchmark Deltas
Reasoning-heavy, deterministic tasks (LiveMath, Spreadsheet, DocVQA) move a lot. SearchQA, ALFWorld,
and OfficeQA stay basically flat between SkillOpt and SkillOpt-Lite.
Endpoint deltas for LiveMath (36.6→73.6, GPT-5.5) and Spreadsheet (39.9→79.4,
GPT-5.4) are as stated in the episode. SkillOpt intermediate values and the DocVQA/SearchQA/ALFWorld/
OfficeQA figures are illustrative reconstructions of the described "flat vs. moving" pattern, not
exact paper numbers.
Convergence Over Ten Batches
SkillOpt-Lite climbs faster early and never gives the lead back — SkillOpt is capped at 4
epochs or 10 batches, SkillOpt-Lite at a strict 10 batches.
Where Did the Gain Actually Come From?
On SpreadsheetBench, HarnessOpt without any skill layer already scores 0.7651 — close
to the entire headline result.
Generalization Risk: Gated vs. Ungated
Section 3.1's pilot skipped validation gating entirely. Redder = further below (or below) baseline;
bluer = safely above it.