We welcome issues and pull requests for missing work on harness agents, agent skills, memory, self-improvement, agent RL, evaluation, and safety.
News
- [2026-06-25] 🎉 Release: Our survey is now available on OpenReview.
Citation
If you find this survey or paper list helpful, please cite our work:
@article{jiang2026selfimprovingagents,
title={Self-Improving Agents in the Era of Experience: A Survey of Self- to Meta-Evolution},
author={Che Jiang and Jincheng Zhong and Yu Fu and Kai Tian and Junlin Yang and Kaikai Zhao and Yuchong Wang and Tianwei Luo and Weizhi Wang and Yuxin Zuo and Guoli Jia and Xingtai Lv and Dianqiao Lei and Sihang Zeng and Yuru Wang and Zhenzhao Yuan and Xinwei Long and Ermo Hua and Can Ren and Xin Jiang and Shulei Xie and Yuanchun Zheng and Youbang Sun and Biqing Qi and Ning Ding and Kaiyan Zhang and Bowen Zhou},
journal={OpenReview Archive},
year={2026},
url={https://openreview.net/pdf?id=IUltZSgLMm}
}
Contents
Overview
This repository collects papers, systems, benchmarks, and resources for studying how deployed agentic AI systems become more capable after deployment.
We organize the landscape around the harness agent: a deployed runtime system whose behavior is jointly shaped by a base model, a mutable harness, a user-facing interface, and an environment-facing interface. Under this view, self-improvement first appears as fast runtime adaptation over external surfaces such as skills, memory, context, tools, and execution environments. Repeated experience may later be consolidated into model parameters through reinforcement learning, fine-tuning, or continual learning.
We organize the survey into four parts:
- Paradigm shift: from task-bounded tool-use loops to deployed harness-centered runtime systems.
- External path: skills, memory, context, tools, and environments as editable runtime adaptation surfaces.
- Parameter path and meta-evolution: agent RL, continual learning, and meta-agents for durable learning and update orchestration.
- Conditions and limits: evaluation, safety, governance, and open problems for reliable post-deployment improvement.
Paper List
This paper list follows the references cited by the LaTeX manuscript. It currently includes 331 unique cited entries from 379 unique manuscript citation keys and 379 cited BibTeX records.
Every row includes a date, display name, title, and at least one public source badge. Cited entries whose public source URL still needs verification are omitted from this table until complete metadata is available (44 currently omitted; 4 duplicate cited records collapsed).
Foundations and Surveys
| Date |
Name |
Title |
Paper |
Github |
| 2026-06 |
OpenSkill |
OpenSkill: Open-World Self-Evolution for LLM Agents |
 |
- |
| 2026-05 |
huang2026rawexperience |
From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills |
 |
- |
| 2026-05 |
SkillOpt |
SkillOpt: Executive Strategy for Self-Evolving Agent Skills |
 |
- |
| 2026-04 |
cursor2026cursor3 |
Meet the New Cursor |
 |
- |
| 2026-04 |
neuralcomputers2026_paper |
Neural Computers |
 |
- |
| 2026-04 |
SWE-chat |
SWE-chat: Coding Agent Interactions From Real Users in the Wild |
 |
- |
| 2026-01 |
AI Agent Systems |
AI Agent Systems: Architectures, Applications, and Evaluation |
 |
- |
| 2025-08 |
fang2025selfevolvingagents |
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems |
 |
- |
| 2025-07 |
A Survey of Self-Evolving Agents |
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence |
 |
- |
| 2025-06 |
chen2025compoundaisystems |
From Standalone LLMs to Integrated Intelligence: A Survey of Compound AI Systems |
 |
- |
| 2025-02 |
anthropic2025claudecode |
Claude 3.7 Sonnet and Claude Code |
 |
- |
| 2025-01 |
silver2025eraexperience |
Welcome to the Era of Experience |
 |
- |
| 2024-05 |
WildChat |
WildChat: 1M ChatGPT Interaction Logs in the Wild |
 |
- |
| 2024-04 |
tao2024selfevolutionsurvey |
A Survey on Self-Evolution of Large Language Models |
 |
- |
| 2024-01 |
langchain2024langgraph |
LangGraph |
 |
- |
| 2023-09 |
xi2023riseagentsurvey |
The Rise and Potential of Large Language Model Based Agents: A Survey |
 |
- |
| 2023-08 |
wang2023llmagentsurvey |
A Survey on Large Language Model based Autonomous Agents |
 |
- |
| 2022-01 |
ahn2022can |
Do as i can, not as i say: Grounding language in robotic affordances |
 |
- |
Harness and Runtime Architecture
| Date |
Name |
Title |
Paper |
Github |
| 2026-08 |
Ouroboros |
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution |
 |
 |
| 2026-06 |
recursive2026automatedresearch |
First Steps Toward Automated AI Research |
 |
- |
| 2026-06 |
osmani2026loopengineering |
Loop Engineering |
 |
- |
| 2026-06 |
Traj-Evolve |
Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection |
 |
- |
| 2026-05 |
Agent Harness Engineering |
Agent Harness Engineering: A Survey |
 |
- |
| 2026-04 |
meng2026agentharness |
Agent Harness for Large Language Model Agents: A Survey |
 |
- |
| 2026-04 |
boeckeler2026harnessengineering |
Harness Engineering for Coding Agent Users |
 |
- |
| 2026-04 |
martin2026harnessfailures |
Most AI Agent Failures Are Harness Failures |
 |
- |
| 2026-04 |
Scaling Managed Agents |
Scaling Managed Agents: Decoupling the Brain from the Hands |
 |
- |
| 2026-04 |
xu2026futureagentsopensource |
The Future of Agents is Open Source (part 1 of 2) |
 |
- |
| 2026-03 |
cursor2026composer2 |
Composer 2 Technical Report |
 |
- |
| 2026-03 |
yue2026workflowsurvey |
From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents |
 |
- |
| 2026-02 |
Unlocking the Codex harness |
Unlocking the Codex harness: how we built the App Server |
 |
- |
| 2026-01 |
openharness2026_repo |
OpenHarness |
 |
 |
| 2025-11 |
young2025effectiveharnesses |
Effective Harnesses for Long-Running Agents |
 |
- |
| 2025-07 |
mei2025contextengineering |
A Survey of Context Engineering for Large Language Models |
 |
- |
| 2025-05 |
Darwin Godel Machine |
Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents |
 |
- |
| 2025-05 |
openai2025codex |
Introducing Codex |
 |
- |
| 2020-01 |
Artificial Intelligence |
Artificial Intelligence: A Modern Approach |
 |
- |
| 2003-01 |
Goedel Machines |
Goedel Machines: Self-Referential Universal Problem Solvers Making Provably Optimal Self-Improvements |
 |
- |
Skills and Skill Libraries
| Date |
Name |
Title |
Paper |
Github |
| 2026-07 |
Skill-SP |
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills |
 |
 |
| 2026-06 |
li2026agenticenvironmentengineering |
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application |
 |
- |
| 2026-05 |
zhou2026cta |
Counterfactual Trace Auditing of LLM Agent Skills |
 |
- |
| 2026-05 |
Group of Skills |
Group of Skills: Group-Structured Skill Retrieval for Agent Skill Libraries |
 |
- |
| 2026-05 |
MIND-Skill |
MIND-Skill: Quality-Guaranteed Skill Generation via Multi-Agent Induction and Deduction |
 |
- |
| 2026-05 |
MUSE-Autoskill |
MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation |
 |
- |
| 2026-05 |
OpenClaw Research |
OpenClaw Research: A Systematic Survey of Large Language Model Agents in Open Deployment |
 |
- |
| 2026-05 |
SkillEvolver |
SkillEvolver: Skill Learning as a Meta-Skill |
 |
- |
| 2026-05 |
SkillRAE |
SkillRAE: Agent Skill-Based Context Compilation for Retrieval-Augmented Execution |
 |
- |
| 2026-05 |
SkillRet |
SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents |
 |
- |
| 2026-04 |
CoEvoSkills |
CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification |
 |
- |
| 2026-04 |
Corpus2Skill |
Corpus2Skill: Distilling Document Corpora into Hierarchical Skill Directories for Agent Navigation |
 |
- |
| 2026-04 |
Externalization in LLM Agents |
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering |
 |
- |
| 2026-04 |
hivemind2026_repo |
Hivemind: Continual Learning Layer that Distills Coding-Agent Session Trajectories into Reusable Skills |
 |
 |
| 2026-04 |
skillswild2026realistic |
How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings |
 |
- |
| 2026-04 |
SKILL0 |
SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization |
 |
- |
| 2026-04 |
SkillX |
SkillX: Automatically Constructing Skill Knowledge Bases for Agents |
 |
- |
| 2026-03 |
bi2026repositorymining |
Automating Skill Acquisition through Large-Scale Mining of Open-Source Agentic Repositories: A Framework for Multi-Agent Procedural Knowledge Extraction |
 |
- |
| 2026-03 |
AutoSkill |
AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution |
 |
- |
| 2026-03 |
d2skill2026dynamic |
Dynamic Dual-Granularity Skill Bank for Agentic RL |
 |
- |
| 2026-03 |
EvoSkill |
EvoSkill: Automated Skill Discovery for Multi-Agent Systems |
 |
- |
| 2026-03 |
From Model to Agent |
From Model to Agent: Equipping the Responses API with a Computer Environment |
 |
- |
| 2026-03 |
MetaClaw |
MetaClaw: Just Talk – An Agent That Meta-Learns and Evolves in the Wild |
 |
- |
| 2026-03 |
SkillNet |
SkillNet: Create, Evaluate, and Connect AI Skills |
 |
- |
| 2026-03 |
SkillRouter |
SkillRouter: Retrieve-and-Rerank Skill Selection for LLM Agents at Scale |
 |
- |
| 2026-03 |
SWE-Skills-Bench |
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering? |
 |
- |
| 2026-03 |
Trace2Skill |
Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills |
 |
- |
| 2026-03 |
XSkill |
XSkill: Continual Learning from Experience and Skills in Multimodal Agents |
 |
- |
| 2026-02 |
Skill-Pro |
Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents |
 |
- |
| 2026-02 |
SkillRL |
SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning |
 |
- |
| 2026-02 |
SkillsBench |
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks |
 |
- |
| 2026-01 |
AutoRefine |
AutoRefine: From Trajectories to Reusable Expertise for Continual LLM Agent Refinement |
 |
- |
| 2026-01 |
hermesagent2026_repo |
Hermes Agent |
 |
 |
| 2026-01 |
li2026singleagentskills |
When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail |
 |
- |
| 2025-12 |
sage2025selfimproving |
Reinforcement Learning for Self-Improving Agent with Skill Library |
 |
- |
| 2025-10 |
composeincontext2025skills |
Can Language Models Compose Skills In-Context? |
 |
- |
Memory and Context Management
| Date |
Name |
Title |
Paper |
Github |
| 2026-05 |
Auto-Dreamer |
Auto-Dreamer: Learning Offline Memory Consolidation for Language Agents |
 |
- |
| 2026-05 |
MemORAI |
MemORAI: Memory Organization and Retrieval via Adaptive Graph Intelligence for LLM Conversational Agents |
 |
- |
| 2026-05 |
MEMOREPAIR |
MEMOREPAIR: Barrier-First Cascade Repair in Agentic Memory |
 |
- |
| 2026-05 |
zou2026demem |
Remember the Decision, Not the Description: A Rate-Distortion Framework for Agent Memory |
 |
- |
| 2026-05 |
SAGE |
SAGE: A Self-Evolving Agentic Graph-Memory Engine for Structure-Aware Associative Memory |
 |
- |
| 2026-04 |
APEX-MEM |
APEX-MEM: Agentic Semi-Structured Memory with Temporal Reasoning for Long-Term Conversational AI |
 |
- |
| 2026-04 |
Cognis |
Cognis: Context-Aware Memory for Conversational AI Agents |
 |
- |
| 2026-04 |
DeltaMem |
DeltaMem: Towards Agentic Memory Management via Reinforcement Learning |
 |
- |
| 2026-04 |
zhang2026lightmem |
Lightweight LLM Agent Memory with Small Language Models |
 |
- |
| 2026-03 |
zhang2026amac |
Adaptive Memory Admission Control for LLM Agents |
 |
- |
| 2026-03 |
AriadneMem |
AriadneMem: Threading the Maze of Lifelong Memory for LLM Agents |
 |
- |
| 2026-03 |
Chronos |
Chronos: Temporal-Aware Conversational Agents with Structured Event Retrieval for Long-Term Memory |
 |
- |
| 2026-03 |
CLAG |
CLAG: Adaptive Memory Organization via Agent-Driven Clustering for Small Language Model Agents |
 |
- |
| 2026-03 |
GAAMA |
GAAMA: Graph Augmented Associative Memory for Agents |
 |
- |
| 2026-03 |
Memori |
Memori: A Persistent Memory Layer for Efficient, Context-Aware LLM Agents |
 |
- |
| 2026-03 |
Memory for Autonomous LLM Agents |
Memory for Autonomous LLM Agents: Mechanisms, Evaluation, and Emerging Frontiers |
 |
- |
| 2026-03 |
PlugMem |
PlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents |
 |
- |
| 2026-02 |
CAST |
CAST: Character-and-Scene Episodic Memory for Agents |
 |
- |
| 2026-02 |
From Lossy to Verified |
From Lossy to Verified: A Provenance-Aware Tiered Memory for Agents |
 |
- |
| 2026-02 |
Live-Evo |
Live-Evo: Online Evolution of Agentic Memory from Continuous Feedback |
 |
- |
| 2026-02 |
MemoryArena |
MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks |
 |
- |
| 2026-02 |
xinlewu2026umem |
Towards Autonomous Memory Agents |
 |
- |
| 2026-02 |
UI-Mem |
UI-Mem: Self-Evolving Experience Memory for Online Reinforcement Learning in Mobile GUI Agents |
 |
- |
| 2026-02 |
xMemory |
xMemory: Beyond RAG for Agent Memory – Retrieval by Decoupling and Aggregation |
 |
- |
| 2026-01 |
Active Context Compression |
Active Context Compression: Autonomous Memory Management in LLM Agents |
 |
- |
| 2026-01 |
Agentic Memory |
Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents |
 |
- |
| 2026-01 |
EMemBench |
EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents |
 |
- |
| 2026-01 |
H-Mem |
H-Mem: A Novel Memory Mechanism for Evolving and Retrieving Agent Memory via a Hybrid Structure |
 |
- |
| 2026-01 |
baldelli2026hangman |
LLMs Can’t Play Hangman: On the Necessity of a Private Working Memory for Language Agents |
 |
- |
| 2026-01 |
Mem2ActBench |
Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents |
 |
- |
| 2026-01 |
MemRL |
MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory |
 |
- |
| 2026-01 |
SimpleMem |
SimpleMem: Efficient Lifelong Memory for LLM Agents |
 |
- |
| 2025-11 |
WebCoach |
WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance |
 |
- |
| 2025-08 |
Nemori |
Nemori: Self-Organizing Agent Memory Inspired by Cognitive Science |
 |
- |
| 2025-03 |
Search-R1 |
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning |
 |
- |
| 2024-10 |
LongMemEval |
LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory |
 |
- |
| 2024-04 |
zhang2024memorymechanism |
A Survey on the Memory Mechanism of Large Language Model Based Agents |
 |
- |
| 2023-12 |
empoweringworkingmemory2023 |
Empowering Working Memory for Large Language Model Agents |
 |
- |
| 2023-10 |
MemGPT |
MemGPT: Towards LLMs as Operating Systems |
 |
- |
| 2023-03 |
Reflexion |
Reflexion: Language Agents with Verbal Reinforcement Learning |
 |
- |
| 2020-05 |
lewis2020rag |
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks |
 |
- |
| 2016-12 |
kirkpatrick2017ewc |
Overcoming catastrophic forgetting in neural networks |
 |
- |
| Date |
Name |
Title |
Paper |
Github |
| 2026-05 |
zhong2026executablebenchmark |
An Executable Benchmarking Suite for Tool-Using Agents |
 |
- |
| 2026-05 |
wu2026chemcost |
Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning |
 |
- |
| 2026-05 |
CUA-Gym |
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents |
 |
- |
| 2026-05 |
MANTRA |
MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents |
 |
- |
| 2026-05 |
PhysicianBench |
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments |
 |
- |
| 2026-05 |
SkillSmith |
SkillSmith: Compiling Agent Skills into Boundary-Guided Runtime Interfaces |
 |
- |
| 2026-05 |
When Simulation Lies |
When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents |
 |
- |
| 2026-04 |
Agent-World |
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence |
 |
- |
| 2026-04 |
Agentic World Modeling |
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond |
 |
- |
| 2026-04 |
Gym-Anything |
Gym-Anything: Turn any Software into an Agent Environment |
 |
- |
| 2026-04 |
zhou2026sandmle |
Synthetic Sandbox for Training Machine Learning Engineering Agents |
 |
- |
| 2026-04 |
ToolMisuseBench |
ToolMisuseBench: An Offline Deterministic Benchmark for Tool Misuse and Recovery in Agentic Systems |
 |
- |
| 2026-02 |
CLI-Gym |
CLI-Gym: Scalable CLI Task Generation via Agentic Environment Inversion |
 |
- |
| 2026-02 |
shen2026wac |
World-Model-Augmented Web Agents with Action Correction |
 |
- |
| 2026-01 |
xiang2026selfevolvingcoevolution |
A Systematic Survey of Self-Evolving Agents: From Model-Centric to Environment-Driven Co-Evolution |
 |
- |
| 2026-01 |
AG-UI |
AG-UI: The Agent-User Interaction Protocol |
 |
 |
| 2026-01 |
a2a_spec_2026 |
Agent2Agent (A2A) Protocol |
 |
 |
| 2026-01 |
CLI-Anything |
CLI-Anything: Making ALL Software Agent-Native |
 |
 |
| 2026-01 |
Harbor |
Harbor: A Framework for Running Agent Evaluations and Creating and Using RL Environments |
 |
 |
| 2026-01 |
LiteCoder-Terminal |
LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents |
 |
- |
| 2026-01 |
MEnvAgent |
MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering |
 |
- |
| 2026-01 |
openclaw2026_repo |
OpenClaw |
 |
 |
| 2026-01 |
SearchGym |
SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation |
 |
- |
| 2026-01 |
Terminal-Bench |
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces |
 |
- |
| 2026-01 |
WebGym |
WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks |
 |
- |
| 2025-08 |
SEAgent |
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience |
 |
- |
| 2025-06 |
proceduraltooluse2025_paper |
Procedural Environment Generation for Tool-Use Agents |
 |
- |
| 2025-05 |
DeepResearchGym |
DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research |
 |
- |
| 2025-05 |
MLE-Dojo |
MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering |
 |
- |
| 2025-03 |
chezelles2025browsergym |
The BrowserGym Ecosystem for Web Agent Research |
 |
- |
| 2025-01 |
mcp_spec_2025 |
Model Context Protocol Specification |
 |
 |
| 2024-12 |
pan2024swegym |
Training Software Engineering Agents and Verifiers with SWE-Gym |
 |
- |
| 2024-05 |
AndroidWorld |
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents |
 |
- |
| 2024-04 |
OSWorld |
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments |
 |
- |
| 2024-03 |
CRADLE |
CRADLE: General Computer Agents with Tool Creation and Knowledge Discovery |
 |
- |
| 2024-03 |
DeepSeek-VL |
DeepSeek-VL: Towards Real-World Vision-Language Understanding |
 |
- |
| 2024-02 |
tang2024worldcoder |
WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment |
 |
- |
| 2024-01 |
AppWorld |
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents |
 |
- |
| 2024-01 |
WorkArena |
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks? |
 |
- |
| 2023-10 |
SWE-bench |
SWE-bench: Can Language Models Resolve Real-World GitHub Issues? |
 |
- |
| 2023-07 |
WebArena |
WebArena: A Realistic Web Environment for Building Autonomous Agents |
 |
- |
| 2023-04 |
Generative Agents |
Generative Agents: Interactive Simulacra of Human Behavior |
 |
- |
| 2023-03 |
MM-ReAct |
MM-ReAct: Prompting ChatGPT for Multimodal Reasoning and Action |
 |
- |
| 2022-10 |
ReAct |
ReAct: Synergizing Reasoning and Acting in Language Models |
 |
- |
| 2020-10 |
ALFWorld |
ALFWorld: Aligning Text and Embodied Environments for Interactive Learning |
 |
- |
Agent RL and Continual Learning
| Date |
Name |
Title |
Paper |
Github |
| 2026-04 |
Skill-SD |
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents |
 |
- |
| 2026-03 |
cursor2026realtimerl |
Improving Composer through real-time RL |
 |
- |
| 2026-03 |
OpenClaw-RL |
OpenClaw-RL: Train Any Agent Simply by Talking |
 |
- |
| 2026-02 |
xue2026acurl |
Autonomous Continual Learning of Computer-Use Agents for Environment Adaptation |
 |
- |
| 2026-02 |
liu2026empo2 |
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization |
 |
- |
| 2026-02 |
ye2026opcd |
On-Policy Context Distillation for Language Models |
 |
- |
| 2025-12 |
wei2025selfplayswerl |
Toward Training Superintelligent Software Agents through Self-Play SWE-RL |
 |
- |
| 2025-11 |
MemSearcher |
MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning |
 |
- |
| 2025-10 |
wang2025agenticrlguide |
A Practitioner’s Guide to Multi-turn Agentic Reinforcement Learning |
 |
- |
| 2025-09 |
Kimi-Dev |
Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents |
 |
- |
| 2025-09 |
jin2025rluserconversations |
The Era of Real-World Human Interaction: RL from User Conversations |
 |
- |
| 2025-09 |
zhang2026agenticrlsurvey |
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey |
 |
- |
| 2025-08 |
Agent Lightning |
Agent Lightning: Train ANY AI Agents with Reinforcement Learning |
 |
- |
| 2025-08 |
ComputerRL |
ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents |
 |
- |
| 2025-05 |
Absolute Zero |
Absolute Zero: Reinforced Self-play Reasoning with Zero Data |
 |
- |
| 2025-04 |
ReTool |
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs |
 |
- |
| 2025-04 |
SWE-smith |
SWE-smith: Scaling Data for Software Engineering Agents |
 |
- |
| 2025-04 |
ToolRL |
ToolRL: Reward is All Tool Learning Needs |
 |
- |
| 2025-02 |
SWE-RL |
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution |
 |
- |
| 2021-12 |
WebGPT |
WebGPT: Browser-Assisted Question-Answering with Human Feedback |
 |
- |
| 2021-01 |
kairouz2021advances |
Advances and Open Problems in Federated Learning |
 |
- |
| 2020-01 |
li2020federated |
Federated Optimization in Heterogeneous Networks |
 |
- |
| 2019-01 |
Towards Federated Learning at Scale |
Towards Federated Learning at Scale: System Design |
 |
- |
| 2017-06 |
lopezpaz2017gem |
Gradient Episodic Memory for Continual Learning |
 |
- |
| 2017-01 |
mcmahan2017communication |
Communication-Efficient Learning of Deep Networks from Decentralized Data |
 |
- |
| Date |
Name |
Title |
Paper |
Github |
| 2026-06 |
Agon |
Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy |
 |
- |
| 2026-05 |
Ace-Skill |
Ace-Skill: Bootstrapping Multimodal Agents with Prioritized and Clustered Evolution |
 |
- |
| 2026-05 |
Continual Harness |
Continual Harness: Online Adaptation for Self-Improving Foundation Agents |
 |
- |
| 2026-05 |
Skill1 |
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning |
 |
- |
| 2026-05 |
SkillOS |
SkillOS: Learning Skill Curation for Self-Evolving Agents |
 |
- |
| 2026-04 |
Agentic Harness Engineering |
Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses |
 |
- |
| 2026-04 |
Autogenesis |
Autogenesis: A Self-Evolving Agent Protocol |
 |
- |
| 2026-04 |
CORAL |
CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery |
 |
 |
| 2026-04 |
Experience as a Compass |
Experience as a Compass: Multi-agent RAG with Evolving Orchestration and Agent Prompts |
 |
- |
| 2026-04 |
Meta-TTL |
Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies |
 |
 |
| 2026-04 |
pan2026mstar |
M^: Every Task Deserves Its Own Memory Harness |
 |
- |
| 2026-04 |
cheng2026mem2evolve |
Mem^2Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation |
 |
- |
| 2026-04 |
qiao2026mia |
Memory Intelligence Agent |
 |
- |
| 2026-04 |
PRIME |
PRIME: Training Free Proactive Reasoning via Iterative Memory Evolution for User-Centric Agent |
 |
- |
| 2026-04 |
RoboPhD |
RoboPhD: Evolving Diverse Complex Agents Under Tight Evaluation Budgets |
 |
- |
| 2026-04 |
yang2026memoryextraction |
Self-Evolving LLM Memory Extraction Across Heterogeneous Tasks |
 |
- |
| 2026-04 |
The World Leaks the Future |
The World Leaks the Future: Harness Evolution for Future Prediction Agents |
 |
- |
| 2026-03 |
AgentFactory |
AgentFactory: A Self-Evolving Framework Through Executable Subagent Accumulation and Reuse |
 |
- |
| 2026-03 |
AI-Supervisor |
AI-Supervisor: Autonomous AI Research Supervision via a Persistent Research World Model |
 |
- |
| 2026-03 |
ARISE |
ARISE: Agent Reasoning with Intrinsic Skill Evolution in Hierarchical Reinforcement Learning |
 |
- |
| 2026-03 |
AutoAgent |
AutoAgent: Evolving Cognition and Elastic Memory Orchestration for Adaptive Agents |
 |
- |
| 2026-03 |
zhang2026hyperagents |
Hyperagents |
 |
- |
| 2026-03 |
Memento-Skills |
Memento-Skills: Let Agents Design Agents |
 |
- |
| 2026-03 |
Meta-Harness |
Meta-Harness: End-to-End Optimization of Model Harnesses |
 |
- |
| 2026-03 |
Mimosa Framework |
Mimosa Framework: Toward Evolving Multi-Agent Systems for Scientific Research |
 |
- |
| 2026-03 |
Nurture-First Agent Development |
Nurture-First Agent Development: Building Domain-Expert AI Agents Through Conversational Knowledge Crystallization |
 |
- |
| 2026-03 |
RetroAgent |
RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback |
 |
- |
| 2026-03 |
SAGE |
SAGE: Multi-Agent Self-Evolution for LLM Reasoning |
 |
- |
| 2026-02 |
AOrchestra |
AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration |
 |
- |
| 2026-02 |
Group-Evolving Agents |
Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing |
 |
- |
| 2026-02 |
KernelBlaster |
KernelBlaster: Continual Cross-Task CUDA Optimization via Memory-Augmented In-Context Reinforcement Learning |
 |
- |
| 2026-02 |
xiong2026alma |
Learning to Continually Learn via Meta-learning Agentic Memory Designs |
 |
- |
| 2026-02 |
MemSkill |
MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents |
 |
- |
| 2026-02 |
MetaMem |
MetaMem: Evolving Meta-Memory for Knowledge Utilization through Self-Reflective Symbolic Optimization |
 |
- |
| 2026-02 |
Position |
Position: Agentic Evolution is the Path to Evolving LLMs |
 |
- |
| 2026-02 |
ROMA |
ROMA: Recursive Open Meta-Agent Framework for Long-Horizon Multi-Agent Systems |
 |
- |
| 2026-02 |
SkillOrchestra |
SkillOrchestra: Learning to Route Agents via Skill Transfer |
 |
- |
| 2026-02 |
Tool-R0 |
Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data |
 |
- |
| 2026-01 |
GraphPlanner |
GraphPlanner: Graph Memory-Augmented Agentic Routing for Multi-Agent LLMs |
 |
 |
| 2026-01 |
ye2026mce |
Meta Context Engineering via Agentic Skill Evolution |
 |
- |
| 2026-01 |
MetaGen |
MetaGen: Self-Evolving Roles and Topologies for Multi-Agent LLM Reasoning |
 |
- |
| 2025-11 |
Agent0 |
Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning |
 |
- |
| 2025-10 |
MLE-Smith |
MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline |
 |
- |
| 2025-09 |
MetaEvo |
MetaEvo: A Meta-Optimization Framework for Experience-Driven Agent Evolution |
 |
- |
| 2025-08 |
MetaAgent |
MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning |
 |
- |
| 2025-05 |
AlphaEvolve |
AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms |
 |
- |
| 2025-05 |
dang2025evolvingorchestration |
Multi-Agent Collaboration via Evolving Orchestration |
 |
- |
| 2025-04 |
FlowReasoner |
FlowReasoner: Reinforcing Query-Level Meta-Agents |
 |
- |
| 2025-04 |
TTRL |
TTRL: Test-Time Reinforcement Learning |
 |
- |
| 2025-02 |
A-MEM |
A-MEM: Agentic Memory for LLM Agents |
 |
- |
| 2025-02 |
zhang2025maas |
Multi-agent Architecture Search via Agentic Supernet |
 |
- |
| 2024-10 |
AFlow |
AFlow: Automating Agentic Workflow Generation |
 |
- |
| 2024-10 |
AgentSquare |
AgentSquare: Automatic LLM Agent Search in Modular Design Space |
 |
- |
| 2024-08 |
hu2024adas |
Automated Design of Agentic Systems |
 |
- |
| 2023-05 |
Voyager |
Voyager: An Open-Ended Embodied Agent with Large Language Models |
 |
- |
Evaluation and Benchmarks
| Date |
Name |
Title |
Paper |
Github |
| 2026-05 |
ABRA |
ABRA: Agent Benchmark for Radiology Applications |
 |
- |
| 2026-05 |
Agent-BRACE |
Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertainty |
 |
- |
| 2026-05 |
wang2026dora |
Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations |
 |
- |
| 2026-05 |
raj2026consistency |
Consistency as a Testable Property |
 |
- |
| 2026-05 |
lee2026ctfusion |
CTFusion |
 |
- |
| 2026-05 |
DataClaw |
DataClaw: A Process-Oriented Agent Benchmark for Exploratory Real-World Data Analysis |
 |
- |
| 2026-05 |
gao2026evidencesupportedbounds |
Evidence-Supported Score Bounds |
 |
- |
| 2026-05 |
Evolving-RL |
Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents |
 |
- |
| 2026-05 |
surana2026gfcr |
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning |
 |
- |
| 2026-05 |
wu2026longmemevalv2 |
LongMemEval-V2 |
 |
- |
| 2026-05 |
MMTB |
MMTB: Evaluating Terminal Agents on Multimedia-File Tasks |
 |
- |
| 2026-05 |
zhao2026rethinkingexperience |
Rethinking Experience Utilization in Self-Evolving Language Model Agents |
 |
- |
| 2026-04 |
agentbeats_registry2026 |
AgentBeats Dashboard / Agent Registry |
 |
- |
| 2026-04 |
agentbeats_docs_aaa2026 |
Agentified Agent Assessment (AAA) & AgentBeats |
 |
- |
| 2026-04 |
gurram2026agentpropbench |
AgentProp-Bench |
 |
- |
| 2026-04 |
ClawArena |
ClawArena: Benchmarking AI Agents in Evolving Information Environments |
 |
- |
| 2026-04 |
ClawBench |
ClawBench: Can AI Agents Complete Everyday Online Tasks? |
 |
 |
| 2026-04 |
EvoAgentBench |
EvoAgentBench: A Multi-Domain Benchmark for Self-Evolving Agents |
 |
- |
| 2026-04 |
chi2026frontiereng |
Frontier-Eng |
 |
- |
| 2026-04 |
SEARL |
SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents |
 |
- |
| 2026-04 |
SkillLearnBench |
SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks |
 |
- |
| 2026-03 |
agentmemorybench2026 |
Benchmarking Continual Agent Memory for Online Learning, Transfer, and Forgetting |
 |
- |
| 2026-03 |
CUBE |
CUBE: A Standard for Unifying Agent Benchmarks |
 |
- |
| 2026-03 |
DomusMind |
DomusMind: A Benchmark for Evaluating Lifelong Smart Home Agents Under Drift |
 |
- |
| 2026-02 |
Agent World Model |
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning |
 |
- |
| 2026-02 |
ResearchGym |
ResearchGym: Evaluating Language Model Agents on Real-World AI Research |
 |
- |
| 2026-02 |
SE-Bench |
SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization |
 |
- |
| 2026-02 |
When AI Benchmarks Plateau |
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation |
 |
- |
| 2026-01 |
DevOps-Gym |
DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle |
 |
- |
| 2026-01 |
SIP-Bench |
SIP-Bench: An Open Protocol for Longitudinal Self-Improvement Evaluation |
 |
 |
| 2025-07 |
mohammadi2025agentbenchmarking |
Evaluation and Benchmarking of LLM Agents: A Survey |
 |
- |
| 2025-07 |
SWE-MERA |
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks |
 |
- |
| 2025-06 |
liu2025cer |
Contextual Experience Replay for Self-Improvement of Language Agents |
 |
- |
| 2025-05 |
LifelongAgentBench |
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners |
 |
- |
| 2025-05 |
zhang2025swebenchlive |
SWE-bench Goes Live! |
 |
- |
| 2025-05 |
SWE-rebench |
SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents |
 |
- |
| 2025-03 |
yehudai2025agentevaluation |
Survey on Evaluation of LLM-based Agents |
 |
- |
| 2025-01 |
cai2025building |
Building self-evolving agents via experience-driven lifelong learning: A framework and benchmark |
 |
- |
| 2024-12 |
TheAgentCompany |
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks |
 |
- |
| 2024-10 |
MLE-bench |
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering |
 |
- |
| 2024-09 |
wang2025awm |
Agent Workflow Memory |
 |
- |
| 2024-06 |
tau-bench |
tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains |
 |
- |
| 2024-03 |
LiveCodeBench |
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code |
 |
- |
| 2024-02 |
maharana2024locomo |
Evaluating Very Long-Term Conversational Memory of LLM Agents |
 |
- |
| 2023-08 |
AgentBench |
AgentBench: Evaluating LLMs as Agents |
 |
- |
| 2023-07 |
ToolLLM |
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs |
 |
- |
| 2019-04 |
vandeven2019threescenarios |
Three Scenarios for Continual Learning |
 |
- |
Safety and Governance
| Date |
Name |
Title |
Paper |
Github |
| 2026-05 |
wu2026biv |
Behavioral Integrity Verification for AI Agent Skills |
 |
- |
| 2026-05 |
liang2026mobius |
Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection |
 |
- |
| 2026-05 |
LITMUS |
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments |
 |
- |
| 2026-05 |
yin2026fate |
On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment |
 |
- |
| 2026-05 |
Proteus |
Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems |
 |
- |
| 2026-05 |
SARC |
SARC: A Governance-by-Architecture Framework for Agentic AI Systems |
 |
- |
| 2026-05 |
ShadowMerge |
ShadowMerge: A Novel Poisoning Attack on Graph-Based Agent Memory via Relation-Channel Conflicts |
 |
- |
| 2026-05 |
SkillScope |
SkillScope: Toward Fine-Grained Least-Privilege Enforcement for Agent Skills |
 |
- |
| 2026-05 |
SkillsVote |
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution |
 |
- |
| 2026-05 |
STALE |
STALE: Can LLM Agents Know When Their Memories Are No Longer Valid? |
 |
- |
| 2026-05 |
Under the Hood of SKILL.md |
Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry |
 |
- |
| 2026-04 |
AgentWatcher |
AgentWatcher: A Rule-Based Prompt Injection Monitor |
 |
- |
| 2026-04 |
ATBench |
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis |
 |
- |
| 2026-04 |
Claw-Eval |
Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents |
 |
- |
| 2026-04 |
chang2026visualinjections |
If You’re Waiting for a Sign… That Might Not Be It! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems |
 |
- |
| 2026-04 |
MemEvoBench |
MemEvoBench: Benchmarking Memory MisEvolution in LLM Agents |
 |
- |
| 2026-04 |
SafeAgent |
SafeAgent: A Runtime Protection Architecture for Agentic Systems |
 |
- |
| 2026-04 |
SkillClaw |
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver |
 |
- |
| 2026-04 |
SkillForge |
SkillForge: Forging Domain-Specific, Self-Evolving Agent Skills in Cloud Technical Support |
 |
- |
| 2026-04 |
Sovereign Agentic Loops |
Sovereign Agentic Loops: Decoupling AI Reasoning from Execution in Real-World Systems |
 |
- |
| 2026-04 |
Spore |
Spore: Efficient and Training-Free Privacy Extraction Attack on LLMs via Inference-Time Hybrid Probing |
 |
- |
| 2026-03 |
From Storage to Steering |
From Storage to Steering: Memory Control Flow Attacks on LLM Agents |
 |
- |
| 2026-03 |
lam2026ssgm |
Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework |
 |
- |
| 2026-03 |
SkillProbe |
SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration |
 |
- |
| 2026-03 |
SkillTester |
SkillTester: Benchmarking Utility and Security of Agent Skills |
 |
- |
| 2026-02 |
agentskills2026architecture |
Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward |
 |
- |
| 2026-02 |
AgentSys |
AgentSys: Secure and Dynamic LLM Agents Through Explicit Hierarchical Memory Management |
 |
- |
| 2026-02 |
ClawHavoc |
ClawHavoc: 341 Malicious Clawed Skills Found by the Bot They Were Targeting |
 |
- |
| 2026-02 |
SoK |
SoK: Agentic Skills – Beyond Tool Use in LLM Agents |
 |
- |
| 2026-01 |
WildClawBench |
WildClawBench: An In-the-Wild Benchmark for AI Agents in the OpenClaw Environment |
 |
 |
| 2025-12 |
MemoryGraft |
MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval |
 |
- |
| 2025-10 |
A-MemGuard |
A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory |
 |
- |
| 2025-09 |
InjecMEM |
InjecMEM: Memory Injection Attack on LLM Agent Memory Systems |
 |
- |
| 2025-07 |
shanghai2025frontierrisk |
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report |
 |
- |
| 2025-07 |
SafeWork-R1 |
SafeWork-R1: Coevolving Safety and Intelligence under the AI-45^ Law |
 |
- |
| 2025-06 |
su2025autonomyrisk |
A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents |
 |
- |
| 2025-06 |
Context Manipulation Attacks |
Context Manipulation Attacks: Web Agents Are Susceptible to Corrupted Memory |
 |
- |
| 2025-06 |
DRIFT |
DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents |
 |
- |
| 2025-06 |
ferrag2025promptprotocol |
From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows |
 |
- |
| 2025-06 |
RedDebate |
RedDebate: Safer Responses through Multi-Agent Red Teaming Debates |
 |
- |
| 2025-06 |
fang2025safemcp |
We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems |
 |
- |
| 2025-03 |
AutoRedTeamer |
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration |
 |
- |
| 2024-12 |
Agent-SafetyBench |
Agent-SafetyBench: Evaluating the Safety of LLM Agents |
 |
- |
| 2024-12 |
yang2024ai45law |
Towards AI-45^ Law: A Roadmap to Trustworthy AGI |
 |
- |
| 2024-10 |
zhang2025asb |
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-Based Agents |
 |
- |
| 2024-03 |
IsolateGPT |
IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems |
 |
- |
| 2024-02 |
Agent Smith |
Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast |
 |
- |
| 2024-02 |
dong2024conversationsafety |
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey |
 |
- |
| 2022-12 |
Constitutional AI |
Constitutional AI: Harmlessness from AI Feedback |
 |
- |
| 2022-03 |
ouyang2022instructgpt |
Training Language Models to Follow Instructions with Human Feedback |
 |
- |
Acknowledgment
This repository is maintained by the FrontisAI and Tsinghua University survey team. Its README structure follows the public awesome-list style of TsinghuaC3I/Awesome-RL-for-LRMs.
Star History