An AI that actually does the work. Not just talks about it.
개요
An AI that actually does the work. Not just talks about it.
README
graff
An AI that actually does the work. Not just talks about it.
Install it on your Mac, Linux, or Windows machine, sign in with the AI subscription you already have, and hand it real tasks. graff writes and runs code, automates the boring stuff, digs through your files, researches the web, and runs its own experiments, on its own, until the job is done. You don’t chat with it. You give it work.
curl -fsSL https://github.com/justrach/codegraff/releases/latest/download/install.sh | sh
Prefer a window? Grab the desktop app. Then just run graff and tell it what you need.
What can I ask it?
If you could do it at a computer, you can ask graff to do it for you:
- “Build me a little app to track my workouts.” It writes it, runs it, and shows you.
- “Turn this folder of messy CSVs into one clean spreadsheet.”
- “Figure out why my site is slow, then fix it.”
- “Scrape these five pages and summarize them.”
- “Run an experiment: try three versions of this and tell me which scores best.”
It works in your real terminal, on your real files, with the real internet, and it can spin up a whole team of sub-agents to work in parallel. It even keeps score of which approaches work and gets better over time.
Don’t write code? You don’t have to. Say what you want in plain English; graff figures out the steps and does them.
How it compares
Run the same job on graff, Claude Code, and Codex (three read-only questions about this repo, plus an 8-trial latency test), and here is what it means for you:
Your AI bill is a fraction. graff runs the same task on whatever model fits your budget. On deepseek-v4-pro it averaged $0.022 per task, against Claude Code’s $0.51 (Opus 4.8) and Codex’s $0.42 (gpt-5.5). That is roughly 20× cheaper, because Claude Code only runs Claude and Codex only runs GPT, while graff runs deepseek, kimi, glm, grok, minimax, gpt, claude, and more. On the same model the token usage is comparable, so the win is the freedom to pick a cheaper one, not a token trick.
It stays out of your way. graff is one 2.7 MB Zig binary. In these runs it used about 25 MB of memory for focused work (more when it reads a lot of code), against Claude Code’s steady ~410 MB (Node) and Codex’s ~206 MB (Rust). Leave it running next to everything else and your laptop won’t notice.
Scripts and CI finish in half the time. For one-shot runs (graff -p, the SDKs, a CI step), graff completed a gpt-5.5 turn in 4.4 s versus Codex’s 8.9 s on the identical ChatGPT endpoint, on every single trial. That is graff’s near-instant startup beating a heavier per-call launch. In a long interactive session the startup amortizes and both settle to model latency, so this is a one-shot and automation win, not a blanket “graff is faster.”
Method: macOS, same machine, read-only code questions on this repo. Cost is each tool’s own reported usage at codegraff gateway prices; memory is peak RSS via /usr/bin/time -l; latency is 8 concurrent graff/Codex pairs on a tool-free prompt with reasoning effort matched. Your numbers will vary with the task, the model, and the network. Reproduce it yourself: benchmarks/.
Under the hood
the super simple harness. A minimal agentic coding harness in Zig 0.17 dev. One 2.7 MB binary, zero runtime dependencies. Talks to Anthropic, any OpenAI-compatible endpoint (DeepSeek, OpenAI, …), or your ChatGPT subscription.
user text ─→ POST ─→ model asks for tools?
↑ │
│ results ▼
└─ io.async ×N (parallel: bash/files/subagents)
… repeat until the model stops
A REPL that talks to the model directly over HTTPS (std.http.Client),
hand-rolls every JSON wire format (std.json), and runs tool calls (subagents
included) in parallel on the std.Io thread pool. It also compacts its own
context when the conversation gets long.
Contents
- Install · give it a key · run it
- Why
- Code intelligence
- An evolutionary harness
- Providers & models
- CLI reference
- REPL commands
- Permission modes
- SDKs: TypeScript & Python
- Reference
- Coming soon
Install
Desktop app: macOS (Apple Silicon)
Prefer a window over a terminal? Download the latest signed, notarized build, drag it to Applications, and open it. The desktop app is fully self-contained: it bundles the graff agent, so there’s nothing else to install to start coding, and it keeps itself up to date automatically. On first launch it drops two commands on your PATH: codegraff (opens that folder in the app, code-style) and graff itself (the agent CLI, in your terminal), so the one install covers both the window and the command line. The terminal graff is symlinked into the app, so it auto-updates along with it. Not on Apple Silicon, or want a standalone CLI? Use the command-line install below.
or browse all releases
Command line: macOS · Linux · Windows
Grab the latest prebuilt release binary: macOS builds are Developer ID signed
and Apple notarized; on any other platform the installer builds from source with
Zig 0.17.0-dev.813+2153f8143 (select it with zigup):
curl -fsSL https://github.com/justrach/codegraff/releases/latest/download/install.sh | sh
From a checkout, just run ./install.sh. The binary lands in ~/bin by default
(override with HARNESS_DIR). The installer appends that directory to your
~/.zshrc (and ~/.bashrc / fish config when those are your shell or already
exist), so a new terminal finds graff — the usual “it installed but
graff: command not found” miss. Skip with HARNESS_NO_PATH=1. Open a new tab
or source ~/.zshrc once after install:
| tool | purpose |
|---|---|
graff |
the agent CLI + REPL: the one binary this script installs |
codedb |
optional code-intelligence companion (structural search/outline/callers). graff auto-detects it and points at the one-line install if it’s missing; everything else works without it |
kuri |
optional browser companion backing webfetch’s markdown path, web crawling, and the kuri skill. Installed by default alongside graff; opt out with HARNESS_NO_KURI=1, never fatal if it fails |
On Windows, grab graff-x86_64-windows.tar.gz (or aarch64) from the
latest release, unpack
it, and put graff.exe on your PATH; the shell installer itself is
Unix-only and points Windows at WSL.
Give it a key
Three ways, pick whichever is easiest:
graff login # free codegraff key (device-code OAuth, no signup forms)
graff login kimi # Kimi Code subscription OAuth (device-code)
graff key set deepseek sk-... # store ANY provider's key (macOS Keychain, else 0600 file)
export DEEPSEEK_API_KEY=sk-... # or just an env var (env always wins)
Already logged into the Codex CLI? Skip this step. Your ChatGPT subscription is
picked up automatically from ${CODEX_HOME:-~/.codex}/auth.json. Or run
graff login codex.
Note on
login:graff loginis the free codegraff key,graff login codexis the ChatGPT-subscription OAuth, andgraff login kimiis Kimi Code’s device-code OAuth. Every other provider (deepseek, openai, anthropic, xai, zai, minimax, xiaomi) is a key: set it withgraff key setor its_API_KEYenv var, then select a model with--model//model. See Providers & models.
Run it
graff # starts on the first provider you have a key for
graff --model deepseek-reasoner # or pin one explicitly
First things to try once you’re at the › prompt:
› what's in this directory? summarize the build setup.
› /model sonnet # fuzzy-switches to claude-sonnet-4-6
› spawn three subagents to summarize src/, count TODOs, and check git status, in parallel
› ultracode audit this repo for error-handling gaps # codeword → multi-agent workflow mode
› /help # everything else
Zed (External Agents / ACP)
graff acp speaks the Agent Client Protocol, so Zed can drive it as an
External Agent. Register it in ~/.config/zed/settings.json:
{
"agent_servers": {
"graff": {
"type": "custom",
"command": "/path/to/graff",
"args": ["acp", "--model", ""],
"env": {}
}
}
}
Then fully quit and reopen Zed (agent_servers is read at startup), and start
a thread via agent: new external agent thread in the command palette — or the
Agent Panel’s new-thread menu → graff under External Agents.
Gotchas worth knowing:
- Pin a default model (
--model) — otherwise the ACP turn fails withno language model configured. Verify the route first withgraff route. - Zed’s own model picker is empty for external agents — it shows
“no match / configure a provider”, which looks like an error but isn’t.
Switch models inside the thread with
/model. - Old threads stay broken: a thread created before auth/model was configured keeps failing even after restart; start a fresh one.
- Provider logins are shared with terminal graff (codegraff OAuth etc.), but
e.g. Codex models need their own
graff login codex. - Debugging:
dev: open acp logsin Zed shows the raw ACP traffic.
Why
Measured, not vibes: arm64 macOS, ReleaseFast; methodology and the budgets each change is held to live in architecture.md:
| metric | measured |
|---|---|
| binary | 2.74 MB, self-contained with zero runtime dependencies |
| cold start | ~1.8 ms |
| full agentic turn | 12 MB peak RSS, ~4% CPU (network-bound) |
| 8 parallel subagents | +0.4 MB each (15 MB total) |
| tool output into history | one 4 KB handle, whatever the result’s size: a 500 MB python child process never touches the harness’s footprint |
Benchmarked against the Rust codegraff (justrach/codegraff, 39 MB binary, 934 crates) on the same model through the same endpoint, interleaved 3×: turn speed was a dead tie (2.94 s vs 2.93 s; the network and the model dominate the turn), but the Zig harness ran in 4.3× less memory (11.3 MB vs 48.5 MB), starts roughly 3× faster, and is about 14× smaller on disk. An agent CLI rarely wins on turn speed; it can win on the cost of being there.
Code intelligence: token-efficient by default
The fastest way to blow a context window is to read whole files into it. graff
ships with a built-in codedb tool: read-only, structural code
intelligence over a local index of the repo
(github.com/justrach/codedb), and the
system prompt steers the model to reach for it before grep or
whole-file reads. Instead of paying for a 2,000-line file to find one function,
the model asks for exactly the shape it needs:
codedb outline src/main.zig # just the symbol map, functions/types, no bodies
codedb symbol switchProvider --body # one function, by name
codedb callers recordUsage # who calls it (call sites, not files)
codedb search "parse SSE" # indexed search, ranked hits, not a grep dump
codedb context "add a new provider" # task-shaped orientation across the codebase
Why this keeps token cost low:
- Structural slices, not files.
outline/symbol/callers/depsreturn a function map or a single definition, tens of lines where aread_filewould spend thousands. The index is queried, not the raw bytes streamed into history. - It’s free and indexed; the metered tools come second. The system prompt
encodes an explicit search order: try the free, indexed
codedbfirst; fall to (metered)muonry/raw search only for literal/regex or non-indexed files. The cheap path is the default path. - Hard output cap. A query is truncated at 64 KB with a marker that nudges
the model back toward targeted queries (
outline,symbol --body) rather than whole-file reads, so even a broad search can’t balloon the context. - Same index powers the
@file picker (codedb glob), so attaching a file by name never shells out to a directory walk.
Pure-Zig client to a pure-Zig server, zero dependencies on either side. Allowed
subcommands: search · symbol · callers · find · outline · read · tree · context · word · deps · glob · ls · file · hot. Not installed? The tool says so
and points at the one-line install; everything else keeps working without it.
An evolutionary harness
graff doesn’t just run an agent; it records every run as a node in a
Darwin Gödel Machine-style archive tree (arXiv:2505.22954),
so the harness itself is the substrate for agent self-improvement. Each run
writes a unique .graff/trajectories/.jsonl; archive readers aggregate
the directory, so concurrent processes never share a truncate/append cursor:
- A lineage tree, not a flat log. Interactive root turns form a spine (each
turn’s parent is the previous one); every subagent and workflow task hangs off
the turn that spawned it. Each node carries a fingerprint of the system prompt
it ran with (
prompt_sha= first 8 bytes of SHA-256), so prompt mutations (set_system_prompton the spine, per-childsystem_promptoverrides on the fan-out) show up as hash changes along edges. A lineage can be replayed or scored offline. - Personas are variants. Subagents pick a persona with
agent(built-ins:reviewer · researcher · implementer · skeptic, plus anything in.harness/agents/) or take a customsystem_prompt; either way the trajectory records the lineage, so you can mine which agent variant actually worked. - A fitness ledger with integrity. The
scorechannel appends evaluation records (prompt_sha,score,parent_sha; the lineage edge DGM parent selection counts children with). Because the archive lives in the working directory, a forgedscorerow could manufacture fitness, so every score the harness writes is HMAC-signed (keyed byGRAFF_SCORE_KEY_FILE, a secret outside the cwd that the evolving agent’s confined tools can’t reach). Readers recompute the HMAC and reject unsigned or forged rows. Signing is opt-in and backward-compatible (no key → unsigned, accepted as before). - Tool-use is mined too. Each agent logs its tool calls (name + error flag,
in order): the process signal behind “which tool combinations work”,
joinable to scores via
prompt_sha. - Consent-scoped fleet loop. Learning contributes prompt-free aggregate
fitness by default, announced once per machine;
/privacy localopts out entirely, and individually reviewed reusable templates need an exact per-artifact approval on top./trajectoryrenders the current session’s agent tree; see docs/hyperagents.md for the full design.
For controlled local hill-climbing, graff learn adds a separate
parent → mutate → paired-evaluate → select loop with immutable evidence, manual
promotion by default, explicitly gated automatic promotion, atomic activation,
and rollback. It never treats trajectories or best-effort telemetry as promotion
authority. See Local prompt-policy learning, including
the no-sandbox trust boundary and the collective-learning design that is not
yet an implemented remote authority.
Providers & models
Direct API-key and OAuth providers across three wire formats. A
ProviderSpec table holds each built-in provider’s
endpoint, auth style, env var, and default model; base URLs and key names come
from models.dev’s api.json (snapshot 2026-06-10).
| Provider | Wire format / auth | Key env var |
|---|---|---|
anthropic |
Anthropic Messages, x-api-key | ANTHROPIC_API_KEY |
codegraff |
OpenAI chat, bearer | CODEGRAFF_API_KEY (cg_sk_...) |
deepseek |
OpenAI chat, bearer | DEEPSEEK_API_KEY |
openai |
OpenAI chat, bearer | OPENAI_API_KEY |
minimax |
Anthropic Messages, bearer | MINIMAX_API_KEY |
xiaomi (MiMo) |
OpenAI chat, bearer | XIAOMI_API_KEY |
kilo |
OpenAI chat, bearer | KILO_API_KEY |
groq |
OpenAI chat, bearer | GROQ_API_KEY |
cerebras |
OpenAI chat, bearer | CEREBRAS_API_KEY |
vercel |
OpenAI chat, bearer (AI Gateway coding-agent) | AI_GATEWAY_API_KEY |
openrouter |
OpenAI chat, bearer | OPENROUTER_API_KEY |
mistral |
OpenAI chat, bearer | MISTRAL_API_KEY |
kimi |
Live catalog-selected: native Kimi chat + bearer, or Anthropic beta Messages + x-api-key when declared | graff login kimi or KIMI_API_KEY |
xai (grok) / zai (GLM) |
OpenAI chat, bearer | XAI_API_KEY / ZAI_API_KEY (via graff key set) |
codex |
Responses API, ChatGPT login | ${CODEX_HOME:-~/.codex}/auth.json (no API key) |
Using a specific provider directly is always the same two steps: give it
the key, then name a model. For example, DeepSeek straight to api.deepseek.com:
graff key set deepseek sk-... # or: export DEEPSEEK_API_KEY=sk-...
graff --model deepseek-reasoner # models: deepseek-v4-pro · deepseek-v4-flash · deepseek-chat · deepseek-reasoner
The same pattern works for every API-key row above: swap in the provider id
and one of its models (graff key set openai sk-... → --model gpt-...,
graff key set anthropic sk-ant-... → --model sonnet, and so on).
To add one workspace-local OpenAI-compatible router without changing Graff,
create .graff/.config.router:
{
"id": "myrouter",
"name": "My Router",
"base_url": "https://router.example.com/v1",
"env_key": "MYROUTER_API_KEY",
"default_model": "example/model"
}
name is optional. takes_effort: true is also available for routers that
accept OpenAI-style reasoning-effort requests. The file contains no secret;
use export MYROUTER_API_KEY=... or graff key set myrouter ....
Select it with graff --model myrouter or /model myrouter;
graff models refresh pulls its full catalog. Graff derives
/chat/completions and /models from base_url, caches that router’s model
catalog in .graff/.models.router, and exposes it to the CLI and GUI schema.
Only one additional router is configured per workspace. .graff/ is ignored
by Git in this repository.
A model is routed to the first provider (in the table order above) that both has
a key set and lists the model in the active catalog. Codex names, rollout
visibility, ordering, and context windows come from its account-scoped /models
endpoint, cached for five minutes with Graff’s supported Codex protocol version
(a separately installed older Codex CLI cannot hide newer models). Baked Codex
rows are only the logged-out/offline fallback, currently including gpt-5.6-sol,
Terra, and Luna. Unknown claude* models
fall back to Anthropic; any other unknown model falls back to the codegraff
gateway, and /model prints a warning when that fallback fires, since a typo’d
name will be rejected by the API on the first request. The startup default is
the first provider with a key, on its default model. /models prints the full
table: context window, compaction point, provider, and which providers you have
keys for; /model switches (a bare /model opens an interactive fuzzy
picker). graff models refresh forces a fresh Codex catalog request and refreshes
the Codegraff/workspace-router catalogs plus the independent models.dev
price/context metadata cache.
CLI reference
usage:
graff [flags] start the REPL
graff [-p] "prompt" one-shot: run the prompt, print the answer, exit
graff login get a codegraff key (device-code OAuth)
graff login codex [--refresh] ChatGPT/Codex OAuth login (PKCE)
graff key set store a key (macOS Keychain, else 0600 file)
graff key list show which providers have keys
graff mcp add -- add an MCP server to .mcp.json
graff mcp list configured MCP servers
graff plugins list Cursor/Claude/Grok/Codex plugin trees (in place)
graff learn local prompt-policy learning and rollback
graff --schema print the machine-readable interface (SDK codegen)
flags:
--model start on this model (same fuzzy resolution as /model)
--subagent-model pin children/workflows/judges to this model on the root provider
--subagent-provider route pinned workers through this explicit provider
--allow-cross-provider-subagents
consent to worker prompts/code going to another provider
--yolo skip all permission prompts for the session
--no-local-tools embedder mode: hard-disable the built-in bash/file/codedb
tools process-wide (see "Embedder mode" below)
-p, --print one-shot print mode (answer on stdout, tool progress on stderr)
--timing show per-tool wall-clock on result lines (✓ (312ms) …)
--cost show running session spend in the prompt ([model · 12k tok · $0.0042])
--json structured stdio protocol (JSON in, JSONL events out, SDK transport)
--max-model-calls N cap provider calls across root, children, retries, titles, compaction, and judges (default 0 = unlimited)
-h, --help usage
-V, --version version
Unknown flags are an error (with a pointer to --help), missing model-flag
values are errors, and --help/--version are handled before subcommand
dispatch, so graff login --help prints usage instead of starting an OAuth
flow. With no key configured at all, startup fails with the three quickest fixes
spelled out rather than a bare env-var list.
graff learn help lists the local learning commands. Configuration, adapter
protocols, statistical gates, activation semantics, and security limitations are
specified in docs/local-learning.md.
One-shot mode makes the harness scriptable without the SDK: graff -p "how many TODOs in src/?" runs a full agentic turn (tools included), prints only the
final answer on stdout (progress lines go to stderr), and exits non-zero on
failure. There’s no human to ask, so the permission gate denies anything not
already allowed. Pre-approve commands in .harness/settings.json or pass
--yolo.
REPL commands
A bare / opens the whole list as a filterable full-screen menu (type to
narrow, Enter runs it); Esc during a response interrupts the turn:
generation stops (it works from the moment the request is sent, including a slow
provider connect), what already streamed stays in history with an
[interrupted] marker, and you’re back at the prompt. A bare Esc at the prompt
clears the input line.
While a response streams you stay in control: besides Esc to interrupt,
Ctrl-T (^T) folds/unfolds the live “Thinking” block in place, and the mouse
wheel scrolls your terminal’s own scrollback: the REPL doesn’t grab the mouse, so
scrolling up to re-read earlier output works like any normal terminal (parity with
Claude Code). Folding the Thinking block is keyboard-only (^T). There is no
click-to-fold.
Streaming Markdown is rendered for terminal readability: heading levels get a clear colored hierarchy; bullets, numbered items, nested lists, task checkboxes, and blockquotes use terminal-native markers; bold and inline code drop their raw delimiters; tables align; and fenced code stays copyable without decorative prefixes on body lines.
The interactive UI uses Codegraff’s accent-only Ensō palette: vermilion
coral marks the model, prompt, active selections, tools, and primary Markdown
structure; ordinary text and supporting metadata stay neutral. Success,
warning, and error colors remain semantic, the terminal background is never
overridden, and NO_COLOR is respected. The quiet enso thinking animation is
the stable default; /animation random restores per-request variety and the
other animation names remain available.
The full catalog, straight from the / menu (a bare / opens it as a
filterable full-screen picker; /help prints this same list in the REPL).
/models pings every keyed provider catalog live on each listing, so new
gateway rollouts appear without a restart.
/model switch model/provider, fuzzy match (e.g. "sonnet", "opus")
/models [health] list known models, context windows, compaction points; health shows live state
/clear wipe the conversation and start fresh
/new start a fresh autosaved session
/rename set the current session title
/goal [30m] [text|pause|resume|status|clear]
set a standing objective and work it autonomously; an optional 30s/30m/2h budget paces the run; pause/resume steering, status shows state, clear removes it
/loop [30m] the same autonomous run as /goal, without adopting a standing objective
/review
run one isolated read-only review pass; no edits, delegation, or workflows
/never [|rm ] standing constraints that ride every subagent brief and survive compaction; bare lists them, rm retires one (alias /constraint)
/tell
message a running graff: is a DM (only it hears, any folder); all broadcasts to every graff on this device; /sessions lists who's around
/peek see what a live co-resident session is doing right now (its transcript tail)
/routes [|add ]
your own priced model lanes across providers: view the set and which seat wins each lane now
/plan toggle plan mode: read-only explore + propose; writes/edits denied
/ultracode toggle persistent workflow mode; bare opens an on/off picker, or /ultracode on|off
/fallback [allow|remove|off]
opt-in cross-provider fallback for this workspace (same-provider rollout stays on)
/key [provider secret] show API-key status; /key adds one live (+ Keychain)
/login [codegraff|codex|kimi]
OAuth sign-in (no key to paste); bare opens a picker (codex alias: oai)
/keepcontext toggle keeping the conversation when /model switches wire format (default on)
/effort reasoning depth: low|medium|high|... (codex, deepseek, codegraff; persists)
/reasoning alias for /effort
/fast codex only: priority service tier for lower latency (toggle, persists)
/thinking stream reasoning live vs spinner only (toggle, persists)
/title name the tab from your first prompt (AI session title; toggle, persists)
/strict toggle "every message is a tool" mode
/yolo toggle bash auto-approval (skip permission prompts)
/trace toggle this run's JSONL event trace (and show its path)
/privacy [local|aggregate|templates|examples]
control prompt-learning data egress for this session
/trajectory show this session's agent tree: turns + spawned subagents
/agents list agent types: builtin personas + .harness/agents/*.md
/skills [add|remove ]
list SKILL.md playbooks + companion tools; add/remove enables or disables one
/plugins list Cursor/Claude/Grok/Codex plugin trees graff is reading in place
/hooks list lifecycle hooks and the built-in codedb guard
/doctor read-only health check: goal/todo invariants, and why steering will or will not be appended
/btw ask one side question about this conversation: billed, never added, rides the parent cache prefix
/compact compact history into a fresh context (OpenAI server-side when available)
/rewind [n] list past prompts; /rewind drops prompt n+after & reverts its file edits
/image attach an image to your next message (vision models only)
/images open image URLs from the last response (e.g. issue attachments) in your browser
/paste attach the clipboard image: macOS; also Ctrl-V (⌘V can't be captured)
/bash run a shell command directly
/save [name] write the conversation to .session.json (default: current)
/resume [name] restore a saved conversation (no arg → interactive picker)
/sessions list saved sessions in the cwd
/todo show the current task list
/jobs list background jobs
/cost session token usage and cost
/usage alias for /cost
/debug live content-free observability HUD (turns, tokens, tools, last events)
/tools session tool balance: codedb-pro vs zigrep vs native usage, gate refusals, skew
/animation pick the thinking animation; persists to settings
/theme [name] pick a color theme; /theme off resets to your terminal default; persists
/fleet [on|off] federated DGM contribution (propose/submit/elite_pull)
/mcp [add …] list MCP servers/tools; /mcp add [args...] connects one live
/import-claude copy Claude/Cursor MCP servers and skills into ~/.codegraff and this repo
/help list every command
exit | ctrl-d quit (also /exit; ctrl-c on an empty line)
/plan, /yolo, and /strict change how the permission gate behaves for the
session. See Permission modes.
/goal sets a standing objective that steers every turn as a live
checklist, and starts working it right away: it runs turn after turn (plan, act,
verify) instead of pausing for confirmation between routine steps. /goal pause
stops the steering without losing the objective, /goal resume turns it back on,
and /goal status shows the objective and its current state.
/loop is the same autonomous run for a one-off task, without adopting
a standing objective.
Either way the run stops on its own with a named outcome: accepted once the work
is done, idle when the model stops making tool progress without claiming
completion, cancelled or blocked when you step in or it needs you, exhausted when
a safety limit is hit, and expired when a time budget runs out.
Start with a duration to give the run one: /goal 30m fix the flaky test (also
45s, 2h). Each continuation turn then tells the model where it stands: which
continuation it is on, how long the run has taken, how much of the budget is
left, and one phase hint (explore, implement, finish, wrap up).
Nothing is enforced except the stop itself, and the model is never cut off
mid-turn. Subagents spawned during a timed run are told the parent’s remaining
time, minus a margin for the parent to integrate their results.
/review is the deliberately narrower path for code
review: it suppresses goal/eval/ultracode steering, admits only local
read/search tools and read-only shell inspection, and runs with fresh
model-visible history. There are no implicit review-specific tool or model-call
limits; the ordinary invocation budget is unlimited by default, while explicit
--max-tool-calls and --max-model-calls settings still apply. Only the request
and final report join the parent transcript. Use a later, explicit turn to fix
accepted findings.
Skills
A skill is a markdown playbook graff loads only when a task calls for it. Drop
one in .harness/skills//SKILL.md (or ~/.harness/skills/ for every
project), give it name and description frontmatter, and write the
instructions in the body. Skills already written for Claude Code, Cursor,
Grok, or Codex work as they are: .claude/skills/, ~/.cursor/skills-cursor/,
~/.grok/skills/, plugin skills/ and Claude commands/*.md trees, and
~/.agents/skills/ are read in place (not copied). A Claude plugin works the
same way it does in Claude Code: commands/, a root SKILL.md, inline
mcpServers, and ${CLAUDE_PLUGIN_ROOT} are honored. /plugins and
graff plugins list the trees; /plugins load shows one.
GRAFF_NO_PLUGINS=1 skips them.
Only the name and description enter the system prompt, so a large skill library
costs one line each. The model calls the skill tool to pull a body in when it
needs it, and a skill written mid-session is loadable straight away.
Two skills ship inside the binary: skill-creator (how to author and install
new skills) and mcp-config (how to inspect and change the MCP servers below).
/skills lists everything with its source, /skills remove hides one,
and /skills add brings it back. See
docs/skills.md for the full reference.
MCP servers
Graff speaks both MCP transports directly: local stdio servers and remote
Streamable HTTP servers. Servers already configured for Claude, Cursor, or
Grok are read in place (plugin mcp.json / .mcp.json, ~/.claude.json,
~/.cursor/mcp.json, .cursor/mcp.json) and fill names graff does not
already define; they still need /mcp trust or --yolo. Plugin manifests
may also declare mcpServers inline or as a path (Claude’s shape).
/plugins and graff plugins show which plugin trees contributed. Smolify (https://app.smol.ly/mcp) is available as a
core documentation service; it needs no Node bridge or project configuration.
Its public-read schemas are bundled locally, so startup makes no Smolify
request. The anonymous transport initializes only after an approved tool call,
and recognizable credentials in arguments are blocked locally. Set
GRAFF_NO_SMOLIFY=1 to remove its tool surface entirely. Authenticated and
write-capable tools are hidden unless the session explicitly opts in with
GRAFF_SMOLIFY_ACCESS=full. Other servers can be added from the shell or during
a session:
graff mcp add context7 -- npx -y @upstash/context7-mcp
graff mcp add mobbin --url https://api.mobbin.com/mcp
graff mcp login mobbin # OAuth discovery + browser PKCE flow
graff mcp login smolify # optional access to authenticated Smolify tools
GRAFF_SMOLIFY_ACCESS=full graff # expose the authenticated/full catalog
# In the REPL: /mcp add mobbin --url https://api.mobbin.com/mcp
The equivalent .mcp.json URL entry is
{"mcpServers":{"mobbin":{"url":"https://api.mobbin.com/mcp"}}}. Remote
responses may use either application/json or text/event-stream; Graff keeps
Mcp-Session-Id state and sends MCP-Protocol-Version on requests.
For OAuth-protected endpoints, graff mcp login performs protected
resource and authorization-server discovery, dynamic client registration, and
a browser PKCE flow. Tokens are stored outside the repository under
~/.simple-harness-mcp with user-only permissions and refreshed automatically.
Static HTTP headers can alternatively be added with
--header 'Authorization=Bearer TOKEN' (they are stored in .mcp.json, so
prefer a restricted token and do not commit that file).
The line editor supports ↑/↓ history (persisted to ~/.simple-harness-history),
Tab completion (commands, and model names after /model ), and emacs-style
editing (Ctrl-A/E/W/U/K, Option+Delete, word moves). The selected model is
remembered in ~/.simple-harness-model and resumed next launch
(--model overrides). If the remembered provider/model is absent from
the current catalog or its credentials are missing, graff falls back for that
session with a note, without overwriting the preference. If the preferred
provider later returns a clear authentication, access, removed-model, quota, or
credit failure before producing text or running tools, graff tries the next
configured provider and keeps the saved preference for a future launch. The
prompt is a small statusline:
[model · Fast · Extra high · Plan · cwd /repo · 12345/800k tok (1%) · ⚡cached].
Fast stays immediately beside the model. Active reasoning/workflow modes are
visible at a glance: Low is green, Medium/Extra high/Ultra/Ultracode use the
Codegraff coral accent, High/Plan are yellow, and Max/Strict are red. YOLO is
reported
as an explicit warning when enabled instead of occupying the compact prompt.
Badges for unsupported settings are hidden instead of implying they apply. The
tail shows context used vs the compaction budget, last cache hit, and (for
metered providers) session spend.
Errors aim to be actionable: /resume nope says the session file wasn’t found
and points at /sessions; an unknown /foo points at /help.
Permission modes
By default graff asks before doing anything that can change your machine.
File writes (write_file/edit_file), MCP tool calls, and any bash command
that isn’t read-only stop at a permission gate:
⚠ rm -rf build/
[y]es once · [a]lways allow "rm" (saved to .harness/settings.json) · [n]o ›
- y runs it once · a runs it and remembers the rule · n denies it (the model is told and picks another path).
- Always appends a prefix rule to
.harness/settings.jsonunder"allow", so that command never prompts again, this session or a future one. Pre-seed that file by hand to allow commands up front (it lives next to your hooks; the harness preserves the rest of the file). - Read-only commands are auto-allowed and never prompt:
ls cat head tail wc grep rg pwd which file,git status|diff|log|show,zig build|fmt, but only while every path stays inside the working directory (cat /etc/passwdstill asks), and only as a plain command. A pipe, redirect,&&, or$(…)always prompts, so a second command can’t be smuggled past a prefix match.
Three session-wide modes change the gate. Set on the CLI, or flip them live in the REPL:
| mode | turn on | what it does |
|---|---|---|
| yolo | --yolo · /yolo |
Skip every prompt: bash, edits, and MCP all run without asking. For sandboxes, CI, and -p/--json runs where there’s no human to answer. --yolo starts the session in it; /yolo toggles mid-session. |
| plan | /plan |
Read-only: the model explores and proposes a plan; the gate hard-denies writes, edits, MCP, and any bash beyond the read-only seed (even your saved allow-list) until you /plan again to execute. The prompt shows a yellow Plan badge. |
| strict | /strict |
“Every message is a tool”: the model must call exactly one tool per message and finish with attempt_completion. Useful for deterministic, scriptable agent loops. |
One-shot mode (graff -p "…" or --json) has no human to answer the
prompt, so the gate denies anything not already allowed. Pre-approve commands
in .harness/settings.json or pass --yolo.
Hooks
Shell hooks in .harness/settings.json run at tool-call boundaries — your
own policy layer next to the built-in gate:
{
"hooks": {
"pre_tool": [{ "match": "write_file|edit_file",
"command": "./scripts/guard.sh",
"suggest": "the mcp edit tool",
"timeout_ms": 5000 }],
"post_tool": [{ "match": "write_file", "command": "zig fmt ." }],
"turn_end": [{ "command": "./scripts/notify.sh" }]
}
}
match— tool name,|-separated list, or*(default).commandruns via/bin/sh -cwith the event JSON on stdin:{"event","tool","input"}(post_tool adds"is_error","output").pre_tool: exit 2 blocks the call — the hook’s stderr becomes the tool error the model sees. Addsuggestto name the sanctioned replacement; the denial becomesblocked by pre_tool hook: — use instead:, turning a blocked call into a one-call recovery instead of a guess-and-retry spiral. Any other exit code, timeout, or spawn failure allows — a broken hook never bricks the loop.post_toolruns sequentially after the call (a formatter finishes before the next tool runs); exit codes are ignored.turn_endfires once per completed turn. Default timeout 10 s (timeout_msto change); stderr is capped at 4 KiB.
SDKs: TypeScript & Python
graff is scriptable from your own code. graff --json is a structured stdio
protocol (JSON requests in, JSONL events out; ask_user is answered with a
structured {"type":"answer","text":"...","cancelled":false} line) and graff --schema prints the
machine-readable interface, and the TypeScript and Python SDKs in
sdk/ are auto-generated from that schema, so they never drift from
the binary. On every release tag a GitHub Action rebuilds, regenerates, fails if
the committed SDKs are stale, and publishes to npm (@graff-new/sdk) and PyPI
(simple-harness-sdk).
# Python
from harness_sdk import Harness
with Harness(yolo=True, model="gpt-5.5") as h:
print(h.ask("what is 2+2?"))
for ev in h.chat("read foo.txt"):
print(ev["type"], ev)
// TypeScript
import { Harness, runAgent } from "@graff-new/sdk";
// one-shot, streamed
for await (const ev of runAgent({ prompt: "summarize README.md", model: "gpt-5.5", yolo: true })) {
if (ev.type === "text") process.stdout.write(ev.text);
if (ev.type === "turn") console.log("\ncost $", ev.cost_usd);
}
// long-lived, multi-turn
const session = Harness.init({ model: "claude-opus-4-8", yolo: true }).session();
console.log(await session.ask("what files are here?"));
session.close();
Can’t spawn a local process (edge runtimes, browsers, other machines)? Run
graff serve and both SDKs ship matching remote clients that drive it over
HTTP: @graff-new/sdk/remote (fetch-only: Workers/Deno/Bun/browsers) and
Python’s RemoteHarness (stdlib only). Same method surface, same event stream.
See sdk/README.md.
Embedder mode: run the harness outside the sandbox
If you are embedding graff in a product, the safe shape is to run the agent loop
on your trusted backend and let it reach an isolated sandbox only through tool
calls. --no-local-tools is what makes that shape enforceable:
graff --json --no-local-tools --model gpt-5.5
# same thing, for a process whose argv you don't control:
GRAFF_NO_LOCAL_TOOLS=1 graff --json
With the gate on, bash, bash_output, bash_kill, read_file, edit_file,
write_file and codedb are hard-disabled for the whole process. It is a gate
in the binary, not a permission rule the model can talk its way past, and it
works in two layers because either one alone would be a promise rather than a
guarantee:
- those tools are never advertised, so no provider is told they exist;
- if a provider hallucinates one anyway, dispatch refuses it with a tool error naming the flag, before anything runs.
Subagents and workflow workers inherit the gate, since they run in the same
process. --yolo does not lift it.
Where the coding tools come from. Point graff at your sandbox as an MCP server. MCP tools are untouched by the gate, which is the entire point: you stand up a thin proxy that maps exec/read/write onto your microVM provider, and the model gets those in place of the local ones.
graff mcp add sandbox --url https://sandbox-proxy.example.com/mcp
graff --json --no-local-tools
webfetch stays available (plain HTTP from your own host, with no sandbox to
escape), and so do the orchestration tools: subagent, workflow, the todo
list, and eval.
Why it’s worth the wiring. The tenant provider key stays on your backend
instead of sitting inside the same VM where prompt-injected commands run, so a
hijacked agent can’t read it. Run state lives with your supervisor rather than
in a disposable machine. And because the sandbox is now something a tool call
reaches rather than something the harness lives inside, your MCP proxy can
create the VM lazily, on the first call that actually needs one. Runs that
answer from model knowledge plus webfetch never boot a machine at all: they
get web-request economics, and no boot-and-provision tax on time to first token.
Reference
Coming soon
Active directions. See CHANGELOG.md for what already landed, the Status & roadmap details above for the full list, and the GitHub issues for what’s in flight:
- Sandboxes. Run the agent’s
bash/file tools inside an isolated sandbox (ephemeral container / microVM) so untrusted or destructive steps can’t touch the host. It’s the natural next layer above today’s cwd-confinement and permission gate, and the safe substrate for hands-off evolutionary runs. Embedder mode is the first half of this today:--no-local-toolsplus a sandbox MCP server. A first-class sandbox backend (create/exec/read/write/destroy with provider adapters) would remove the proxy you have to write yourself. - Scaling the evolution loop. The local half shipped in v0.0.219:
graff learn initruns privacy-bounded background trials and promotes the winning genome into the root prompt. What remains is fleet scale: grounded judging and trajectory sync-back across many installs. - Shell completions + man page, and a config file for default flags/model.
Working on codegraff
Install the tracked git hooks once:
scripts/install-hooks.sh
From then on, every push runs tier 1 of the internal eval set. It is deterministic and offline (no provider calls, no network, no spend) and takes about 20 seconds warm:
| check | what it holds |
|---|---|
fmt |
zig fmt --check src build.zig |
lines |
the 600-line ceiling on hand-written Zig |
reach |
every file that declares tests is reachable from the test root |
build |
zig build |
tests |
zig build test, and a suite count that may grow but never shrink |
invariants |
the named goal/loop/todo tests actually ran, not just compiled |
sdk |
the committed SDKs still match graff --schema |
A push that only touches docs skips the whole thing. When a check fails it names the invariant, says which regression it guards, and prints the one-liner that reruns only it:
scripts/eval-tier1.sh # everything
scripts/eval-tier1.sh --only sdk # one check
scripts/eval-tier1.sh --list # the check names
If you need to push past it, git push --no-verify (or GRAFF_SKIP_PREPUSH=1 git push). CI runs the same checks and more, so the hook is a fast local
opinion, not the last word.
Tier 2 is the model-backed half, and it is not in the hook because it runs
the harness for real. The cases live in evals/harness_behavior.jsonl, one line
each, carrying the regression it guards: a goal writes a checklist and finishes
it, an empty todo_write changes nothing, a compaction keeps the items already
completed, --max-tool-calls actually stops the run. The model is scripted, so
the default run is offline and free:
python3 scripts/eval-tier2.py # every case, scripted model
python3 scripts/eval-tier2.py --list # the cases and what they guard
python3 scripts/eval-tier2.py --dump # everything one case did
python3 scripts/eval-tier2.py --provider anthropic --model claude-sonnet-4-6
Two checks exist because the unit suite could not see the failure. reach walks
the @import graph: Zig only runs tests in files the root pulls in, so a
split-out module nobody references compiles to nothing and the suite still
reports green. invariants goes further, rerunning the suite under
-Dtest-filter to prove the named tests executed, because a test can sit in the
source and never run. Both are configured as data in
scripts/eval/tier1-manifest.json.
Both read the count off the compiled test binary rather than off the build
summary, because a fully cached zig build test prints no N/N tests passed
line, and they pick that binary by the test names it carries rather than by
mtime, so a -Dtest-filter build left in .zig-cache/o is never mistaken for
the whole suite. test_count_baseline is a floor nothing raises on its own:
bump it at each release cut, and tier 1 warns once the suite runs more than
test_count_slack tests ahead of it.
License
codegraff is licensed under a modified GNU AGPL-3.0 (see LICENSE).
The public receives it under the AGPL-3.0, so network use triggers the Section 13
obligation to make Corresponding Source available to remote users. The authors
Rach Pradhan (justrach) and Yu Xi Lim (yxlyx) reserve full rights to use,
distribute, and offer it (and modified versions) as a private, proprietary, or
hosted/cloud product, free of those obligations.
A recipient’s AGPL-3.0 licence is perpetual and irrevocable unless they breach it: it can’t be withdrawn at will, which is what makes open use safe to rely on. Any proprietary or commercial permission to use codegraff without the AGPL’s copyleft is a separate thing, and exists only if both authors grant it jointly in writing. Such a permission is revocable at the authors’ discretion at any time, and neither the provision of consultancy or other services nor any side agreement grants it or makes it irrevocable. If it is revoked, the user falls back to full AGPL compliance or must stop using codegraff. For commercial or proprietary licensing, contact the authors.
Built in Zig 0.17 dev · AGPL-3.0 (modified) · architecture.md · uxlog.md
설치
npx -y @modelcontextprotocol/server-everything설정
{
"mcpServers": {
"codedb": { "command": "codedb", "args": ["mcp", "."] },
"everything": { "command": "npx", "args": ["-y", "@modelcontextprotocol/server-everything"] }
}
}