Built on PocketFlow · Inspired by karpathy/autoresearch
概览
Built on PocketFlow · Inspired by karpathy/autoresearch
README
🔥 News
- [2026. 05] 🎉 Major Release! 87 modular skills, CrossRef citation recovery, multi-model quality gate via
REVIEWER_MODEL, and MCP server injection. Full BibTeX stub fallback ensures zero undefined citations. - [2026. 01] 🚀 Nano-scientist public launch — budget-first autonomous research pipeline producing LaTeX + BibTeX + PDF in a single
python main.pycall.
✨ Why Nano-scientist
| Budget-first control | Fix a dollar limit; the agent adapts depth and report type automatically. Loops exit the moment estimated remaining calls fall too low — no wasted spend. |
| Four autonomous loops | Literature → Experimentation → Writing → Compiling — each self-terminating on quality gate or budget exhaustion. No central planner. |
| 87 modular skills | From paper search and code generation to grant proposals, patent drafting, and adversarial review. Lazy-loaded; each skill is one SKILL.md file. |
| Research-to-PDF pipeline | Produces LaTeX source, deduplicated BibTeX (CrossRef-verified), per-skill artifact files, figures, scripts, and a compiled PDF — all in one run. |
| Zero-drop citations | Entries failing CrossRef verification are recovered via title lookup or kept as @misc stubs — never silently dropped. |
🧪 Showcases
Sample reports generated by Nano-scientist at --budget 1:
- What techniques bridge the performance gap between small language models and LLMs in automated bug fixing?
- Challenge taxonomy for the Lean 4 theorem prover: clustering, labeling, and trend analysis across GitHub issues.
- Comparative analysis of five AI coding agents across 933k pull requests (AIDev).
- Survey of on-policy distillation techniques.
🧠 How it works
flowchart TD
I([Initializer\nzero LLM calls]) -->|literature| LIT
subgraph LIT_LOOP[" Literature Loop "]
LIT[LiteratureReviewLoop\ndecide → skill → quality gate]
LIT -->|"next iter"| LIT
end
LIT -->|"goal met · budget low"| EXP
subgraph EXP_LOOP[" Experiment Loop "]
EXP[ExperimentationLoop\ndecide → skill → quality gate]
EXP -->|"next iter"| EXP
end
EXP -->|"goal met · budget low"| WR
subgraph WRITE_LOOP[" Writing Loop "]
WR[WritingLoop\nwrite sections → review pass → fix]
end
WR -->|compile| CT
subgraph COMPILE[" Compiling Loop "]
CT[CompileTeX\npdflatex + bibtex]
FT[FixTeX\npatch errors]
CT -->|fix| FT
FT -->|compile| CT
end
CT -->|done| F([Finisher\ncost_log · summary])
FT -->|done| F
Stage breakdown
| Stage | What happens |
|---|---|
| Initializer | Creates outputs//, classifies topic as survey vs. experimental (is_survey) — zero LLM calls |
| LiteratureReviewLoop | Each iter: LLM picks skill|done → executes skill → quality gate checks goal; exits on goal met or budget low |
| ExperimentationLoop | Survey: synthesis skills (tables, figures from literature). Experimental: experiment-pipeline, experiment-craft, etc. |
| WritingLoop | Writes all required sections, runs a peer-review pass, addresses major comments, assembles .tex |
| CompilingLoop | pdflatex + bibtex; on error or undefined citations, FixTeX patches and recompiles (up to 2 attempts) |
| Finisher | Writes cost_log.json + summary.json, prints total cost |
🚀 Quickstart
# 1) Clone
git clone https://github.com/AI4Scientist/nano-scientist
cd nano-scientist
# 2) Install dependencies
pip install -r requirements.txt
# 3) Add API keys
cp .env.example .env
# edit .env — minimum: OPENROUTER_API_KEY
# 4) Run
python main.py "CRISPR off-target effects in primary T cells" --budget 2.00
# Or pass a research proposal .md file
python main.py proposal.md --budget 0.50
Output lands in outputs//:
outputs/
└── /
├── report.tex # assembled LaTeX source
├── report.pdf # final PDF (if pdflatex installed)
├── references.bib # deduplicated BibTeX
├── artifacts/ # per-skill markdown outputs
├── figures/ # generated plots / images
├── data/ # collected CSV / JSON data
├── scripts/ # executed code blocks
├── traj.txt # full stdout trace
├── history.json # step-by-step execution log
├── cost_log.json # per-step token costs
└── summary.json # final run summary
🖥️ CLI reference
python main.py [topic] [options]
Arguments:
topic Research topic — a plain string or path to a .md file.
Optional when using --list-skills.
Options:
-b, --budget FLOAT Spend limit in USD (default: $1.00)
-o, --output DIR Output directory (default: outputs/)
-e, --env FILE Path to .env file (default: .env)
--list-skills Print available skills and exit
python main.py "CRISPR off-target effects in primary T cells" --budget 1.00
python main.py proposal.md --budget 1.00
python main.py --list-skills
Budget
Every run targets a full 8-section paper. Budget controls depth, not report type — more budget means more skill calls, more citations, and more revision rounds. Loops terminate when estimated remaining LLM calls drop below a threshold, so the agent always spends as much as it can usefully spend.
🧩 Skills
Each skill is a folder under skills/ with a single SKILL.md (lazy-loaded at runtime). Skills with allowed-tools: Bash get a real tool-calling loop with bash execution and error feedback.
Add a skill
- Create
skills/my-skill/SKILL.mdwith YAML frontmatter:
---
id: my-skill
description: One-line description shown in the planner.
allowed-tools: Bash # grants bash tool-calling with error feedback
required-keys: [HF_TOKEN] # optional; skill is filtered out if key missing
---
Your skill instructions here.
- Register in
skills/skills.json:
{ "id": "my-skill", "description": "One-line description shown in the planner." }
🔐 Environment variables
Required
| Variable | Used for |
|---|---|
OPENROUTER_API_KEY |
Core LLM inference (all nodes) |
Skill-gated (optional)
| Variable | Skills that use it |
|---|---|
HF_TOKEN |
Skills accessing Hugging Face Hub |
GITHUB_TOKEN |
Skills querying GitHub repos/issues |
S2_API_KEY |
Semantic Scholar API |
OPENAI_API_KEY |
Skills using OpenAI-compatible endpoints |
Missing skill keys automatically filter out dependent skills at startup.
Tuning (all optional)
| Variable | Default | Purpose |
|---|---|---|
MODEL_NAME |
— | Override the inference model |
INFERENCE_BASE_URL |
— | Custom OpenAI-compatible endpoint |
REVIEWER_MODEL |
— | Second model for quality gate (e.g. openai/gpt-4o); falls back to MODEL_NAME if unset |
INPUT_TOKEN_COST_PER_MILLION |
— | Estimate remaining LLM calls |
OUTPUT_TOKEN_COST_PER_MILLION |
— | Estimate remaining LLM calls |
LOOKBACK |
3 |
History steps visible per LLM call |
MAX_REVIEW_ROUNDS |
1 |
Writing review/revision passes |
MAX_TOOL_ROUNDS |
16 |
Max bash tool-calling rounds per skill |
MAX_LOOP_ITERATIONS |
20 |
Max iterations per research loop |
MIN_CALLS_TO_CONTINUE |
3 |
Stop loop when estimated remaining calls falls below this |
OUTPUT_LANGUAGE |
auto-detect | Force output language (e.g. "French"); ASCII-only topics default to English |
🗂️ Project layout
nano-scientist/
├── main.py # CLI entry point
├── src/
│ ├── flow.py # PocketFlow wiring (3 loops + compile/fix)
│ ├── nodes.py # 7 nodes + helpers
│ └── utils.py # LLM client, cost tracking, BibTeX utils
├── skills/ # 87 modular research skills
│ ├── skills.json # skill index (id + description)
│ └── /
│ └── SKILL.md # instructions + YAML frontmatter
├── outputs/ # generated reports (git-ignored)
└── .env # API keys (git-ignored)
🤝 Join the Community
- Open an issue — bug reports, feature requests, skill ideas
- Submit a PR — new skills are one
SKILL.mdfile
📌 Citation
If you use Nano-scientist in your research, please cite:
@software{nano_scientist2026,
title = {Nano-scientist: Autonomous Research Agent for Budget-Constrained Scientific Reports},
author = {{AI4Scientist Team}},
year = {2026},
url = {https://github.com/AI4Scientist/nano-scientist}
}
推荐工具
换一个关键词,或者移除筛选条件。
安装
npx skillfish add ai4scientist/nano-scientist