AI Post Transformers · Visual Companion

OpenSkill for Open-World Self-Evolution in LLM Agents

Most agent papers prove they can use tools. OpenSkill asks a nastier question: can an agent leave deployment with a durable playbook built from public evidence, its own proxy tests, and no weight updates at all?

Zhiling Yan et al. Posted June 4, 2026 Transcript arXiv scan: no extra IDs beyond 2606.06741
OpenSkill · arXiv 2606.06741 Docs -> skills -> proxy tests Weights stay fixed Verifier remains the bottleneck
Mode
Step
task and deployment boundary open-world evidence and knowledge durable skill artifact proxy rehearsal and diagnosis
View
low / weak middle / ambiguous high / hot
View
OpenSkill baseline or transfer source verifier or component pressure
View
verbal critique / judge loops environment or trajectory-heavy skill learning memory and skill-structure systems open-world grounded skill evolution
References

Paper Graph

Core papers and evaluation threads cited in the episode. arXiv links point directly to the preprints when available; later 2026 systems without public arXiv IDs use the provided Scholar lookup URLs.