
exponential-view-co/autoresearch-autobeta
Developer toolsThis repository contains two for autonomous iterative research — a pattern inspired by Andrej Karpathy's observation (March 2026; URL correct at time of writing) that any domain with a scalar oracle...
Overview
This repository contains two for autonomous iterative research — a pattern inspired by Andrej Karpathy's observation (March 2026; URL correct at time of writing) that any domain with a scalar oracle and fast feedback can be optimised. The core loop is simple: a human states a problem, the AI auto-designs a (judge panel + scoring dimensions), then iterates one atomic change at a time — keeping improvements, discarding regressions — until convergence. The oracle quality determines everything. canonical invocation and machine-readable names are autoresearch and autobeta; display titles in headings may use AutoResearch, Autoresearch, and AutoBeta. Read the skill definitions to understand the pattern. Adapt the templates, oracle design guide, and persona selection logic for your own agent setup. - skills-autoresearch/SKILL.
README
AutoResearch & AutoBeta: AI Research Skill Definitions
What is this?
This repository contains two AI skill definitions for autonomous iterative research — a pattern inspired by Andrej Karpathy’s observation (March 2026; URL correct at time of writing) that any domain with a scalar oracle and fast feedback can be optimised.
The core loop is simple: a human states a problem, the AI auto-designs a scoring oracle (judge panel + scoring dimensions), then iterates one atomic change at a time — keeping improvements, discarding regressions — until convergence. The oracle quality determines everything.
Naming convention: canonical invocation and machine-readable names are autoresearch and autobeta; display titles in headings may use AutoResearch, Autoresearch, and AutoBeta.
Two flavors:
| AutoResearch | AutoBeta | |
|---|---|---|
| Speed | 5–15 minutes | 45–90 minutes |
| Cost | ~$2–5 per run | ~$15–25 per run |
| Strategy | Single-thread hill climbing | Multi-thread landscape search with escape mechanisms |
| Best for | Quick refinement tasks | Complex strategy where the solution space has many distinct regions |
What you can do with this repo: Read the skill definitions to understand the pattern. Adapt the templates, oracle design guide, and persona selection logic for your own agent setup.
Contents
1. skills-autoresearch/ — AutoResearch Skill
Files:
skills-autoresearch/SKILL.md— Complete skill definition, usage guide, triggers, examplesskills-autoresearch/references/— Background materials (oracle design, personas)skills-autoresearch/templates/— Execution templates for the iteration loop
What it does:
- Autonomous iterative improvement of any writing or thinking task
- Auto-designs a scoring oracle (dimensions + judge panel)
- Iterates one change at a time, evaluates, keeps or discards
- Converges when 10 consecutive iterations are discarded
Use case: Thesis refinement, book titles, essay angles, product positioning, strategy arguments
2. skills-autobeta/ — AutoBeta Skill
Files:
skills-autobeta/SKILL.md— Complete skill definition, invocation syntax, configurationskills-autobeta/references/— Background materials (oracle design, personas)- Uses shared templates from
skills-autoresearch/templates/
What it does:
- Enhanced autoresearch with escape harness to avoid local minima
- Phase 0: Conceptual landscape survey (3-cluster mapping)
- Phase 1: Oracle design (persona + scoring dimensions)
- Phase 2–5: 3 parallel threads with tournament selection + escape detection
- Converges after 10 consecutive non-improvements across all threads
- Returns best candidate + score history + thread convergence snapshot + iteration history
Use case: Deep strategy research, thesis refinement, complex problem-solving
Key Concept: Skill vs. Instance
| Skill (This Repo) | Instance (A Specific Run) | |
|---|---|---|
| Purpose | Reusable tool definition | Specific output from one run |
| What it contains | SKILL.md + templates + reference docs | Results, experiment iterations, run logs |
| Example | skills-autoresearch/SKILL.md |
A book research run that produced 18 iterations |
| Analogy | The recipe | One meal cooked from the recipe |
When you invoke a skill, it creates an instance — a project folder containing the oracle, all experiment versions, and score progression for that specific run.
How to Use These Skills
These skills are designed to be integrated into an agent harness that supports sub-agent spawning and tool use. The invocation pattern:
AutoResearch
/autoresearch "problem statement" [--oracle-hint "what matters"] [--iterations N] [--auto]
The skill executes:
- Auto-designs a scoring oracle (dimensions + judge panel)
- Shows oracle for approval (unless
--auto) - Iterates: generate candidate → evaluate → keep or discard
- Reports convergence with best version + score history
AutoBeta
/autobeta "strategy question" [--oracle-hint "context"] [--iterations N] [--auto]
The skill executes:
- Landscape survey (maps the conceptual solution space)
- Oracle design (same pattern as autoresearch)
- 3 parallel search threads with tournament selection
- Escape mechanisms when threads get stuck
- Reports best strategy + score progression across all threads
Architecture
AutoResearch (Fast Loop)
Problem statement
↓
Auto-design oracle (3–4 scoring dimensions + 3 judge personas)
↓
User approves oracle
↓
Iteration loop:
Generate ONE changed version (Sonnet) → Score via oracle (Opus)
→ KEEP if improved, DISCARD if not
↓
Converge (10 consecutive discards = done)
↓
Return best version + score history (5–15 min)
AutoBeta (Deep Loop)
Strategy question + oracle hint
↓
Phase 0: Landscape survey (3-cluster conceptual map)
↓
Phase 1: Oracle design (personas + scoring dimensions)
↓
Phase 2–5: 3 parallel search threads
Thread 1: Dimension A exploration
Thread 2: Dimension B exploration
Thread 3: Dimension C exploration
Per thread:
- Generate successive candidates (v0, v1, v2, ...)
- Score via oracle (Opus)
- Escape if stalled (axis-based or basin crossover)
- Tournament every 15 iterations
↓
Converge on best (10 consecutive non-improvements across ALL threads = stop)
↓
Return best candidate + score progression + all iterations (45–90 min)
Configuration
AutoResearch
iterations: max 30 (default: run to convergence)convergence_threshold: 10 consecutive non-improvements
AutoBeta
iterations: typical 20–50 (user-configurable)tournament_interval: every 15 iterations (tunable)convergence_threshold: 10 consecutive non-improvements across all threads
LLM recommendations
- Oracle/evaluator: strong model (e.g. Claude Opus)
- Candidate generation: fast model (e.g. Claude Sonnet)
- Gate/cluster calls: lightweight model (e.g. Claude Haiku)
Example Run
Skill used: AutoBeta Question: “How do we build defensible AI capabilities that create lasting competitive advantage?” Iterations: 27 Result: Score 9.74/10, converged on regulatory monopoly thesis Time: 43m 55s Cost: ~$15–25
The skill (this repo) is the reusable tool definition. The run above is one instance — a specific output produced by invoking the skill.
Files in This Repo
autoresearch-autobeta/
├── README.md (this file)
├── skills-autoresearch/
│ ├── SKILL.md (full skill definition)
│ ├── references/
│ │ ├── oracle-design.md (oracle design patterns)
│ │ └── personas.md (persona library reference)
│ └── templates/
│ ├── evaluate.md.tmpl (scoring template)
│ ├── experiment_v0.md.tmpl (candidate template)
│ └── program.md.tmpl (agent instruction template)
└── skills-autobeta/
├── SKILL.md (full skill definition)
└── references/
├── oracle-design.md (oracle design patterns)
└── personas.md (persona library reference)
Quick Start
Start by reading the skill definitions:
# AutoResearch — the fast, single-thread version:
cat skills-autoresearch/SKILL.md
# AutoBeta — the enhanced multi-thread version:
cat skills-autobeta/SKILL.md
Then explore the oracle design guide (skills-autoresearch/references/oracle-design.md) — it contains scoring dimension patterns for different problem types (writing, strategy, title, essay, abstract) that are useful beyond this specific tool.
Technical Notes
- LLM: Any capable LLM (recommended: Opus-class for evaluation, Sonnet-class for generation)
- Execution: Sub-agent harness (isolated sessions, multi-threaded for AutoBeta)
- Output: Markdown reports with iteration history + scoring
- Cost: ~$2–25 per run depending on skill, iterations, and oracle calls
Maintained by
Exponential View
Repository created: 26 March 2026 Last updated: 1 April 2026 Status: autoresearch — stable; autobeta — beta
Recommended Tools
Try a different keyword or remove a filter.
Install
npx skillfish add exponential-view-co/autoresearch-autobeta