FA

frontisai/awesome-self-improving-agents

Developer tools
223 stars 품질 71 트렌드 71

We welcome issues and pull requests for missing work on harness agents, agent skills, memory, self-improvement, agent RL, evaluation, and safety.

개요

We welcome issues and pull requests for missing work on harness agents, agent skills, memory, self-improvement, agent RL, evaluation, and safety.

README

We welcome issues and pull requests for missing work on harness agents, agent skills, memory, self-improvement, agent RL, evaluation, and safety.

News

  • [2026-06-25] 🎉 Release: Our survey is now available on OpenReview.

Citation

If you find this survey or paper list helpful, please cite our work:

@article{jiang2026selfimprovingagents,
  title={Self-Improving Agents in the Era of Experience: A Survey of Self- to Meta-Evolution},
  author={Che Jiang and Jincheng Zhong and Yu Fu and Kai Tian and Junlin Yang and Kaikai Zhao and Yuchong Wang and Tianwei Luo and Weizhi Wang and Yuxin Zuo and Guoli Jia and Xingtai Lv and Dianqiao Lei and Sihang Zeng and Yuru Wang and Zhenzhao Yuan and Xinwei Long and Ermo Hua and Can Ren and Xin Jiang and Shulei Xie and Yuanchun Zheng and Youbang Sun and Biqing Qi and Ning Ding and Kaiyan Zhang and Bowen Zhou},
  journal={OpenReview Archive},
  year={2026},
  url={https://openreview.net/pdf?id=IUltZSgLMm}
}

Contents

Overview

This repository collects papers, systems, benchmarks, and resources for studying how deployed agentic AI systems become more capable after deployment.

We organize the landscape around the harness agent: a deployed runtime system whose behavior is jointly shaped by a base model, a mutable harness, a user-facing interface, and an environment-facing interface. Under this view, self-improvement first appears as fast runtime adaptation over external surfaces such as skills, memory, context, tools, and execution environments. Repeated experience may later be consolidated into model parameters through reinforcement learning, fine-tuning, or continual learning.

We organize the survey into four parts:

  1. Paradigm shift: from task-bounded tool-use loops to deployed harness-centered runtime systems.
  2. External path: skills, memory, context, tools, and environments as editable runtime adaptation surfaces.
  3. Parameter path and meta-evolution: agent RL, continual learning, and meta-agents for durable learning and update orchestration.
  4. Conditions and limits: evaluation, safety, governance, and open problems for reliable post-deployment improvement.

Paper List

This paper list follows the references cited by the LaTeX manuscript. It currently includes 331 unique cited entries from 379 unique manuscript citation keys and 379 cited BibTeX records.

Every row includes a date, display name, title, and at least one public source badge. Cited entries whose public source URL still needs verification are omitted from this table until complete metadata is available (44 currently omitted; 4 duplicate cited records collapsed).

Foundations and Surveys

Date Name Title Paper Github
2026-06 OpenSkill OpenSkill: Open-World Self-Evolution for LLM Agents -
2026-05 huang2026rawexperience From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills -
2026-05 SkillOpt SkillOpt: Executive Strategy for Self-Evolving Agent Skills -
2026-04 cursor2026cursor3 Meet the New Cursor -
2026-04 neuralcomputers2026_paper Neural Computers -
2026-04 SWE-chat SWE-chat: Coding Agent Interactions From Real Users in the Wild -
2026-01 AI Agent Systems AI Agent Systems: Architectures, Applications, and Evaluation -
2025-08 fang2025selfevolvingagents A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems -
2025-07 A Survey of Self-Evolving Agents A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence -
2025-06 chen2025compoundaisystems From Standalone LLMs to Integrated Intelligence: A Survey of Compound AI Systems -
2025-02 anthropic2025claudecode Claude 3.7 Sonnet and Claude Code -
2025-01 silver2025eraexperience Welcome to the Era of Experience -
2024-05 WildChat WildChat: 1M ChatGPT Interaction Logs in the Wild -
2024-04 tao2024selfevolutionsurvey A Survey on Self-Evolution of Large Language Models -
2024-01 langchain2024langgraph LangGraph -
2023-09 xi2023riseagentsurvey The Rise and Potential of Large Language Model Based Agents: A Survey -
2023-08 wang2023llmagentsurvey A Survey on Large Language Model based Autonomous Agents -
2022-01 ahn2022can Do as i can, not as i say: Grounding language in robotic affordances -

Harness and Runtime Architecture

Date Name Title Paper Github
2026-08 Ouroboros Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
2026-06 recursive2026automatedresearch First Steps Toward Automated AI Research -
2026-06 osmani2026loopengineering Loop Engineering -
2026-06 Traj-Evolve Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection -
2026-05 Agent Harness Engineering Agent Harness Engineering: A Survey -
2026-04 meng2026agentharness Agent Harness for Large Language Model Agents: A Survey -
2026-04 boeckeler2026harnessengineering Harness Engineering for Coding Agent Users -
2026-04 martin2026harnessfailures Most AI Agent Failures Are Harness Failures -
2026-04 Scaling Managed Agents Scaling Managed Agents: Decoupling the Brain from the Hands -
2026-04 xu2026futureagentsopensource The Future of Agents is Open Source (part 1 of 2) -
2026-03 cursor2026composer2 Composer 2 Technical Report -
2026-03 yue2026workflowsurvey From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents -
2026-02 Unlocking the Codex harness Unlocking the Codex harness: how we built the App Server -
2026-01 openharness2026_repo OpenHarness
2025-11 young2025effectiveharnesses Effective Harnesses for Long-Running Agents -
2025-07 mei2025contextengineering A Survey of Context Engineering for Large Language Models -
2025-05 Darwin Godel Machine Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents -
2025-05 openai2025codex Introducing Codex -
2020-01 Artificial Intelligence Artificial Intelligence: A Modern Approach -
2003-01 Goedel Machines Goedel Machines: Self-Referential Universal Problem Solvers Making Provably Optimal Self-Improvements -

Skills and Skill Libraries

Date Name Title Paper Github
2026-07 Skill-SP Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills
2026-06 li2026agenticenvironmentengineering Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application -
2026-05 zhou2026cta Counterfactual Trace Auditing of LLM Agent Skills -
2026-05 Group of Skills Group of Skills: Group-Structured Skill Retrieval for Agent Skill Libraries -
2026-05 MIND-Skill MIND-Skill: Quality-Guaranteed Skill Generation via Multi-Agent Induction and Deduction -
2026-05 MUSE-Autoskill MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation -
2026-05 OpenClaw Research OpenClaw Research: A Systematic Survey of Large Language Model Agents in Open Deployment -
2026-05 SkillEvolver SkillEvolver: Skill Learning as a Meta-Skill -
2026-05 SkillRAE SkillRAE: Agent Skill-Based Context Compilation for Retrieval-Augmented Execution -
2026-05 SkillRet SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents -
2026-04 CoEvoSkills CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification -
2026-04 Corpus2Skill Corpus2Skill: Distilling Document Corpora into Hierarchical Skill Directories for Agent Navigation -
2026-04 Externalization in LLM Agents Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering -
2026-04 hivemind2026_repo Hivemind: Continual Learning Layer that Distills Coding-Agent Session Trajectories into Reusable Skills
2026-04 skillswild2026realistic How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings -
2026-04 SKILL0 SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization -
2026-04 SkillX SkillX: Automatically Constructing Skill Knowledge Bases for Agents -
2026-03 bi2026repositorymining Automating Skill Acquisition through Large-Scale Mining of Open-Source Agentic Repositories: A Framework for Multi-Agent Procedural Knowledge Extraction -
2026-03 AutoSkill AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution -
2026-03 d2skill2026dynamic Dynamic Dual-Granularity Skill Bank for Agentic RL -
2026-03 EvoSkill EvoSkill: Automated Skill Discovery for Multi-Agent Systems -
2026-03 From Model to Agent From Model to Agent: Equipping the Responses API with a Computer Environment -
2026-03 MetaClaw MetaClaw: Just Talk – An Agent That Meta-Learns and Evolves in the Wild -
2026-03 SkillNet SkillNet: Create, Evaluate, and Connect AI Skills -
2026-03 SkillRouter SkillRouter: Retrieve-and-Rerank Skill Selection for LLM Agents at Scale -
2026-03 SWE-Skills-Bench SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering? -
2026-03 Trace2Skill Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills -
2026-03 XSkill XSkill: Continual Learning from Experience and Skills in Multimodal Agents -
2026-02 Skill-Pro Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents -
2026-02 SkillRL SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning -
2026-02 SkillsBench SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks -
2026-01 AutoRefine AutoRefine: From Trajectories to Reusable Expertise for Continual LLM Agent Refinement -
2026-01 hermesagent2026_repo Hermes Agent
2026-01 li2026singleagentskills When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail -
2025-12 sage2025selfimproving Reinforcement Learning for Self-Improving Agent with Skill Library -
2025-10 composeincontext2025skills Can Language Models Compose Skills In-Context? -

Memory and Context Management

Date Name Title Paper Github
2026-05 Auto-Dreamer Auto-Dreamer: Learning Offline Memory Consolidation for Language Agents -
2026-05 MemORAI MemORAI: Memory Organization and Retrieval via Adaptive Graph Intelligence for LLM Conversational Agents -
2026-05 MEMOREPAIR MEMOREPAIR: Barrier-First Cascade Repair in Agentic Memory -
2026-05 zou2026demem Remember the Decision, Not the Description: A Rate-Distortion Framework for Agent Memory -
2026-05 SAGE SAGE: A Self-Evolving Agentic Graph-Memory Engine for Structure-Aware Associative Memory -
2026-04 APEX-MEM APEX-MEM: Agentic Semi-Structured Memory with Temporal Reasoning for Long-Term Conversational AI -
2026-04 Cognis Cognis: Context-Aware Memory for Conversational AI Agents -
2026-04 DeltaMem DeltaMem: Towards Agentic Memory Management via Reinforcement Learning -
2026-04 zhang2026lightmem Lightweight LLM Agent Memory with Small Language Models -
2026-03 zhang2026amac Adaptive Memory Admission Control for LLM Agents -
2026-03 AriadneMem AriadneMem: Threading the Maze of Lifelong Memory for LLM Agents -
2026-03 Chronos Chronos: Temporal-Aware Conversational Agents with Structured Event Retrieval for Long-Term Memory -
2026-03 CLAG CLAG: Adaptive Memory Organization via Agent-Driven Clustering for Small Language Model Agents -
2026-03 GAAMA GAAMA: Graph Augmented Associative Memory for Agents -
2026-03 Memori Memori: A Persistent Memory Layer for Efficient, Context-Aware LLM Agents -
2026-03 Memory for Autonomous LLM Agents Memory for Autonomous LLM Agents: Mechanisms, Evaluation, and Emerging Frontiers -
2026-03 PlugMem PlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents -
2026-02 CAST CAST: Character-and-Scene Episodic Memory for Agents -
2026-02 From Lossy to Verified From Lossy to Verified: A Provenance-Aware Tiered Memory for Agents -
2026-02 Live-Evo Live-Evo: Online Evolution of Agentic Memory from Continuous Feedback -
2026-02 MemoryArena MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks -
2026-02 xinlewu2026umem Towards Autonomous Memory Agents -
2026-02 UI-Mem UI-Mem: Self-Evolving Experience Memory for Online Reinforcement Learning in Mobile GUI Agents -
2026-02 xMemory xMemory: Beyond RAG for Agent Memory – Retrieval by Decoupling and Aggregation -
2026-01 Active Context Compression Active Context Compression: Autonomous Memory Management in LLM Agents -
2026-01 Agentic Memory Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents -
2026-01 EMemBench EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents -
2026-01 H-Mem H-Mem: A Novel Memory Mechanism for Evolving and Retrieving Agent Memory via a Hybrid Structure -
2026-01 baldelli2026hangman LLMs Can’t Play Hangman: On the Necessity of a Private Working Memory for Language Agents -
2026-01 Mem2ActBench Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents -
2026-01 MemRL MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory -
2026-01 SimpleMem SimpleMem: Efficient Lifelong Memory for LLM Agents -
2025-11 WebCoach WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance -
2025-08 Nemori Nemori: Self-Organizing Agent Memory Inspired by Cognitive Science -
2025-03 Search-R1 Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning -
2024-10 LongMemEval LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory -
2024-04 zhang2024memorymechanism A Survey on the Memory Mechanism of Large Language Model Based Agents -
2023-12 empoweringworkingmemory2023 Empowering Working Memory for Large Language Model Agents -
2023-10 MemGPT MemGPT: Towards LLMs as Operating Systems -
2023-03 Reflexion Reflexion: Language Agents with Verbal Reinforcement Learning -
2020-05 lewis2020rag Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks -
2016-12 kirkpatrick2017ewc Overcoming catastrophic forgetting in neural networks -

Environments, Tools, and Runtime Feedback

Date Name Title Paper Github
2026-05 zhong2026executablebenchmark An Executable Benchmarking Suite for Tool-Using Agents -
2026-05 wu2026chemcost Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning -
2026-05 CUA-Gym CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents -
2026-05 MANTRA MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents -
2026-05 PhysicianBench PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments -
2026-05 SkillSmith SkillSmith: Compiling Agent Skills into Boundary-Guided Runtime Interfaces -
2026-05 When Simulation Lies When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents -
2026-04 Agent-World Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence -
2026-04 Agentic World Modeling Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond -
2026-04 Gym-Anything Gym-Anything: Turn any Software into an Agent Environment -
2026-04 zhou2026sandmle Synthetic Sandbox for Training Machine Learning Engineering Agents -
2026-04 ToolMisuseBench ToolMisuseBench: An Offline Deterministic Benchmark for Tool Misuse and Recovery in Agentic Systems -
2026-02 CLI-Gym CLI-Gym: Scalable CLI Task Generation via Agentic Environment Inversion -
2026-02 shen2026wac World-Model-Augmented Web Agents with Action Correction -
2026-01 xiang2026selfevolvingcoevolution A Systematic Survey of Self-Evolving Agents: From Model-Centric to Environment-Driven Co-Evolution -
2026-01 AG-UI AG-UI: The Agent-User Interaction Protocol
2026-01 a2a_spec_2026 Agent2Agent (A2A) Protocol
2026-01 CLI-Anything CLI-Anything: Making ALL Software Agent-Native
2026-01 Harbor Harbor: A Framework for Running Agent Evaluations and Creating and Using RL Environments
2026-01 LiteCoder-Terminal LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents -
2026-01 MEnvAgent MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering -
2026-01 openclaw2026_repo OpenClaw
2026-01 SearchGym SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation -
2026-01 Terminal-Bench Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces -
2026-01 WebGym WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks -
2025-08 SEAgent SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience -
2025-06 proceduraltooluse2025_paper Procedural Environment Generation for Tool-Use Agents -
2025-05 DeepResearchGym DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research -
2025-05 MLE-Dojo MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering -
2025-03 chezelles2025browsergym The BrowserGym Ecosystem for Web Agent Research -
2025-01 mcp_spec_2025 Model Context Protocol Specification
2024-12 pan2024swegym Training Software Engineering Agents and Verifiers with SWE-Gym -
2024-05 AndroidWorld AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents -
2024-04 OSWorld OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments -
2024-03 CRADLE CRADLE: General Computer Agents with Tool Creation and Knowledge Discovery -
2024-03 DeepSeek-VL DeepSeek-VL: Towards Real-World Vision-Language Understanding -
2024-02 tang2024worldcoder WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment -
2024-01 AppWorld AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents -
2024-01 WorkArena WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks? -
2023-10 SWE-bench SWE-bench: Can Language Models Resolve Real-World GitHub Issues? -
2023-07 WebArena WebArena: A Realistic Web Environment for Building Autonomous Agents -
2023-04 Generative Agents Generative Agents: Interactive Simulacra of Human Behavior -
2023-03 MM-ReAct MM-ReAct: Prompting ChatGPT for Multimodal Reasoning and Action -
2022-10 ReAct ReAct: Synergizing Reasoning and Acting in Language Models -
2020-10 ALFWorld ALFWorld: Aligning Text and Embodied Environments for Interactive Learning -

Agent RL and Continual Learning

Date Name Title Paper Github
2026-04 Skill-SD Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents -
2026-03 cursor2026realtimerl Improving Composer through real-time RL -
2026-03 OpenClaw-RL OpenClaw-RL: Train Any Agent Simply by Talking -
2026-02 xue2026acurl Autonomous Continual Learning of Computer-Use Agents for Environment Adaptation -
2026-02 liu2026empo2 Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization -
2026-02 ye2026opcd On-Policy Context Distillation for Language Models -
2025-12 wei2025selfplayswerl Toward Training Superintelligent Software Agents through Self-Play SWE-RL -
2025-11 MemSearcher MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning -
2025-10 wang2025agenticrlguide A Practitioner’s Guide to Multi-turn Agentic Reinforcement Learning -
2025-09 Kimi-Dev Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents -
2025-09 jin2025rluserconversations The Era of Real-World Human Interaction: RL from User Conversations -
2025-09 zhang2026agenticrlsurvey The Landscape of Agentic Reinforcement Learning for LLMs: A Survey -
2025-08 Agent Lightning Agent Lightning: Train ANY AI Agents with Reinforcement Learning -
2025-08 ComputerRL ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents -
2025-05 Absolute Zero Absolute Zero: Reinforced Self-play Reasoning with Zero Data -
2025-04 ReTool ReTool: Reinforcement Learning for Strategic Tool Use in LLMs -
2025-04 SWE-smith SWE-smith: Scaling Data for Software Engineering Agents -
2025-04 ToolRL ToolRL: Reward is All Tool Learning Needs -
2025-02 SWE-RL SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution -
2021-12 WebGPT WebGPT: Browser-Assisted Question-Answering with Human Feedback -
2021-01 kairouz2021advances Advances and Open Problems in Federated Learning -
2020-01 li2020federated Federated Optimization in Heterogeneous Networks -
2019-01 Towards Federated Learning at Scale Towards Federated Learning at Scale: System Design -
2017-06 lopezpaz2017gem Gradient Episodic Memory for Continual Learning -
2017-01 mcmahan2017communication Communication-Efficient Learning of Deep Networks from Decentralized Data -

Meta-Agents and Evolution Orchestration

Date Name Title Paper Github
2026-06 Agon Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy -
2026-05 Ace-Skill Ace-Skill: Bootstrapping Multimodal Agents with Prioritized and Clustered Evolution -
2026-05 Continual Harness Continual Harness: Online Adaptation for Self-Improving Foundation Agents -
2026-05 Skill1 Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning -
2026-05 SkillOS SkillOS: Learning Skill Curation for Self-Evolving Agents -
2026-04 Agentic Harness Engineering Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses -
2026-04 Autogenesis Autogenesis: A Self-Evolving Agent Protocol -
2026-04 CORAL CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery
2026-04 Experience as a Compass Experience as a Compass: Multi-agent RAG with Evolving Orchestration and Agent Prompts -
2026-04 Meta-TTL Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies
2026-04 pan2026mstar M^: Every Task Deserves Its Own Memory Harness -
2026-04 cheng2026mem2evolve Mem^2Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation -
2026-04 qiao2026mia Memory Intelligence Agent -
2026-04 PRIME PRIME: Training Free Proactive Reasoning via Iterative Memory Evolution for User-Centric Agent -
2026-04 RoboPhD RoboPhD: Evolving Diverse Complex Agents Under Tight Evaluation Budgets -
2026-04 yang2026memoryextraction Self-Evolving LLM Memory Extraction Across Heterogeneous Tasks -
2026-04 The World Leaks the Future The World Leaks the Future: Harness Evolution for Future Prediction Agents -
2026-03 AgentFactory AgentFactory: A Self-Evolving Framework Through Executable Subagent Accumulation and Reuse -
2026-03 AI-Supervisor AI-Supervisor: Autonomous AI Research Supervision via a Persistent Research World Model -
2026-03 ARISE ARISE: Agent Reasoning with Intrinsic Skill Evolution in Hierarchical Reinforcement Learning -
2026-03 AutoAgent AutoAgent: Evolving Cognition and Elastic Memory Orchestration for Adaptive Agents -
2026-03 zhang2026hyperagents Hyperagents -
2026-03 Memento-Skills Memento-Skills: Let Agents Design Agents -
2026-03 Meta-Harness Meta-Harness: End-to-End Optimization of Model Harnesses -
2026-03 Mimosa Framework Mimosa Framework: Toward Evolving Multi-Agent Systems for Scientific Research -
2026-03 Nurture-First Agent Development Nurture-First Agent Development: Building Domain-Expert AI Agents Through Conversational Knowledge Crystallization -
2026-03 RetroAgent RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback -
2026-03 SAGE SAGE: Multi-Agent Self-Evolution for LLM Reasoning -
2026-02 AOrchestra AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration -
2026-02 Group-Evolving Agents Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing -
2026-02 KernelBlaster KernelBlaster: Continual Cross-Task CUDA Optimization via Memory-Augmented In-Context Reinforcement Learning -
2026-02 xiong2026alma Learning to Continually Learn via Meta-learning Agentic Memory Designs -
2026-02 MemSkill MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents -
2026-02 MetaMem MetaMem: Evolving Meta-Memory for Knowledge Utilization through Self-Reflective Symbolic Optimization -
2026-02 Position Position: Agentic Evolution is the Path to Evolving LLMs -
2026-02 ROMA ROMA: Recursive Open Meta-Agent Framework for Long-Horizon Multi-Agent Systems -
2026-02 SkillOrchestra SkillOrchestra: Learning to Route Agents via Skill Transfer -
2026-02 Tool-R0 Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data -
2026-01 GraphPlanner GraphPlanner: Graph Memory-Augmented Agentic Routing for Multi-Agent LLMs
2026-01 ye2026mce Meta Context Engineering via Agentic Skill Evolution -
2026-01 MetaGen MetaGen: Self-Evolving Roles and Topologies for Multi-Agent LLM Reasoning -
2025-11 Agent0 Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning -
2025-10 MLE-Smith MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline -
2025-09 MetaEvo MetaEvo: A Meta-Optimization Framework for Experience-Driven Agent Evolution -
2025-08 MetaAgent MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning -
2025-05 AlphaEvolve AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms -
2025-05 dang2025evolvingorchestration Multi-Agent Collaboration via Evolving Orchestration -
2025-04 FlowReasoner FlowReasoner: Reinforcing Query-Level Meta-Agents -
2025-04 TTRL TTRL: Test-Time Reinforcement Learning -
2025-02 A-MEM A-MEM: Agentic Memory for LLM Agents -
2025-02 zhang2025maas Multi-agent Architecture Search via Agentic Supernet -
2024-10 AFlow AFlow: Automating Agentic Workflow Generation -
2024-10 AgentSquare AgentSquare: Automatic LLM Agent Search in Modular Design Space -
2024-08 hu2024adas Automated Design of Agentic Systems -
2023-05 Voyager Voyager: An Open-Ended Embodied Agent with Large Language Models -

Evaluation and Benchmarks

Date Name Title Paper Github
2026-05 ABRA ABRA: Agent Benchmark for Radiology Applications -
2026-05 Agent-BRACE Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertainty -
2026-05 wang2026dora Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations -
2026-05 raj2026consistency Consistency as a Testable Property -
2026-05 lee2026ctfusion CTFusion -
2026-05 DataClaw DataClaw: A Process-Oriented Agent Benchmark for Exploratory Real-World Data Analysis -
2026-05 gao2026evidencesupportedbounds Evidence-Supported Score Bounds -
2026-05 Evolving-RL Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents -
2026-05 surana2026gfcr Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning -
2026-05 wu2026longmemevalv2 LongMemEval-V2 -
2026-05 MMTB MMTB: Evaluating Terminal Agents on Multimedia-File Tasks -
2026-05 zhao2026rethinkingexperience Rethinking Experience Utilization in Self-Evolving Language Model Agents -
2026-04 agentbeats_registry2026 AgentBeats Dashboard / Agent Registry -
2026-04 agentbeats_docs_aaa2026 Agentified Agent Assessment (AAA) & AgentBeats -
2026-04 gurram2026agentpropbench AgentProp-Bench -
2026-04 ClawArena ClawArena: Benchmarking AI Agents in Evolving Information Environments -
2026-04 ClawBench ClawBench: Can AI Agents Complete Everyday Online Tasks?
2026-04 EvoAgentBench EvoAgentBench: A Multi-Domain Benchmark for Self-Evolving Agents -
2026-04 chi2026frontiereng Frontier-Eng -
2026-04 SEARL SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents -
2026-04 SkillLearnBench SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks -
2026-03 agentmemorybench2026 Benchmarking Continual Agent Memory for Online Learning, Transfer, and Forgetting -
2026-03 CUBE CUBE: A Standard for Unifying Agent Benchmarks -
2026-03 DomusMind DomusMind: A Benchmark for Evaluating Lifelong Smart Home Agents Under Drift -
2026-02 Agent World Model Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning -
2026-02 ResearchGym ResearchGym: Evaluating Language Model Agents on Real-World AI Research -
2026-02 SE-Bench SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization -
2026-02 When AI Benchmarks Plateau When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation -
2026-01 DevOps-Gym DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle -
2026-01 SIP-Bench SIP-Bench: An Open Protocol for Longitudinal Self-Improvement Evaluation
2025-07 mohammadi2025agentbenchmarking Evaluation and Benchmarking of LLM Agents: A Survey -
2025-07 SWE-MERA SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks -
2025-06 liu2025cer Contextual Experience Replay for Self-Improvement of Language Agents -
2025-05 LifelongAgentBench LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners -
2025-05 zhang2025swebenchlive SWE-bench Goes Live! -
2025-05 SWE-rebench SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents -
2025-03 yehudai2025agentevaluation Survey on Evaluation of LLM-based Agents -
2025-01 cai2025building Building self-evolving agents via experience-driven lifelong learning: A framework and benchmark -
2024-12 TheAgentCompany TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks -
2024-10 MLE-bench MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering -
2024-09 wang2025awm Agent Workflow Memory -
2024-06 tau-bench tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains -
2024-03 LiveCodeBench LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code -
2024-02 maharana2024locomo Evaluating Very Long-Term Conversational Memory of LLM Agents -
2023-08 AgentBench AgentBench: Evaluating LLMs as Agents -
2023-07 ToolLLM ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs -
2019-04 vandeven2019threescenarios Three Scenarios for Continual Learning -

Safety and Governance

Date Name Title Paper Github
2026-05 wu2026biv Behavioral Integrity Verification for AI Agent Skills -
2026-05 liang2026mobius Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection -
2026-05 LITMUS LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments -
2026-05 yin2026fate On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment -
2026-05 Proteus Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems -
2026-05 SARC SARC: A Governance-by-Architecture Framework for Agentic AI Systems -
2026-05 ShadowMerge ShadowMerge: A Novel Poisoning Attack on Graph-Based Agent Memory via Relation-Channel Conflicts -
2026-05 SkillScope SkillScope: Toward Fine-Grained Least-Privilege Enforcement for Agent Skills -
2026-05 SkillsVote SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution -
2026-05 STALE STALE: Can LLM Agents Know When Their Memories Are No Longer Valid? -
2026-05 Under the Hood of SKILL.md Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry -
2026-04 AgentWatcher AgentWatcher: A Rule-Based Prompt Injection Monitor -
2026-04 ATBench ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis -
2026-04 Claw-Eval Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents -
2026-04 chang2026visualinjections If You’re Waiting for a Sign… That Might Not Be It! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems -
2026-04 MemEvoBench MemEvoBench: Benchmarking Memory MisEvolution in LLM Agents -
2026-04 SafeAgent SafeAgent: A Runtime Protection Architecture for Agentic Systems -
2026-04 SkillClaw SkillClaw: Let Skills Evolve Collectively with Agentic Evolver -
2026-04 SkillForge SkillForge: Forging Domain-Specific, Self-Evolving Agent Skills in Cloud Technical Support -
2026-04 Sovereign Agentic Loops Sovereign Agentic Loops: Decoupling AI Reasoning from Execution in Real-World Systems -
2026-04 Spore Spore: Efficient and Training-Free Privacy Extraction Attack on LLMs via Inference-Time Hybrid Probing -
2026-03 From Storage to Steering From Storage to Steering: Memory Control Flow Attacks on LLM Agents -
2026-03 lam2026ssgm Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework -
2026-03 SkillProbe SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration -
2026-03 SkillTester SkillTester: Benchmarking Utility and Security of Agent Skills -
2026-02 agentskills2026architecture Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward -
2026-02 AgentSys AgentSys: Secure and Dynamic LLM Agents Through Explicit Hierarchical Memory Management -
2026-02 ClawHavoc ClawHavoc: 341 Malicious Clawed Skills Found by the Bot They Were Targeting -
2026-02 SoK SoK: Agentic Skills – Beyond Tool Use in LLM Agents -
2026-01 WildClawBench WildClawBench: An In-the-Wild Benchmark for AI Agents in the OpenClaw Environment
2025-12 MemoryGraft MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval -
2025-10 A-MemGuard A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory -
2025-09 InjecMEM InjecMEM: Memory Injection Attack on LLM Agent Memory Systems -
2025-07 shanghai2025frontierrisk Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report -
2025-07 SafeWork-R1 SafeWork-R1: Coevolving Safety and Intelligence under the AI-45^ Law -
2025-06 su2025autonomyrisk A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents -
2025-06 Context Manipulation Attacks Context Manipulation Attacks: Web Agents Are Susceptible to Corrupted Memory -
2025-06 DRIFT DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents -
2025-06 ferrag2025promptprotocol From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows -
2025-06 RedDebate RedDebate: Safer Responses through Multi-Agent Red Teaming Debates -
2025-06 fang2025safemcp We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems -
2025-03 AutoRedTeamer AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration -
2024-12 Agent-SafetyBench Agent-SafetyBench: Evaluating the Safety of LLM Agents -
2024-12 yang2024ai45law Towards AI-45^ Law: A Roadmap to Trustworthy AGI -
2024-10 zhang2025asb Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-Based Agents -
2024-03 IsolateGPT IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems -
2024-02 Agent Smith Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast -
2024-02 dong2024conversationsafety Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey -
2022-12 Constitutional AI Constitutional AI: Harmlessness from AI Feedback -
2022-03 ouyang2022instructgpt Training Language Models to Follow Instructions with Human Feedback -

Acknowledgment

This repository is maintained by the FrontisAI and Tsinghua University survey team. Its README structure follows the public awesome-list style of TsinghuaC3I/Awesome-RL-for-LRMs.

Star History

View this README on GitHub

추천 도구

다른 키워드를 입력하거나 필터를 제거해 보세요.

설치

npx skillfish add frontisai/awesome-self-improving-agents