
austin-starks/public-portfolio-challenge
Developer tools4 AI agents. One runbook. Real $25K, live and verifiable. Currently +22.85%.
Обзор
Live portfolio card — generated from current positions. Open the full dashboard → As of September 28, 2026. Same-window comparison begins with the first stored live observation. Portfolio data · Performance data · SPY data · refreshed weekly by GitHub Actions. Copy the live incumbent · Connect MCP → python3 start.py → In February 2026 I deposited into a live Public brokerage account on NexusTrade and made the entire book public — every position, every fill, every model test, every bug, every failure. The live story is a blog series. Grok Bot built separate biotech and semiconductor books, I funded a $13,500 Public account, then rejected the sector-gated construction and reopened the entire design. The final allocator covers 26 thesis-approved companies, passed four of five out-of-sample folds, and replaced the live strategy set on August 26 after the old book was preserved in paper.
README
Contents
- What is this?
- Live books
- Agent bakeoff
- What this repo gives you
- How the runbook works
- What’s inside
- Community leaderboard
- Get started
- Risk disclaimer
- More links
What is this?
In February 2026 I deposited $25,000 into a live Public brokerage account on NexusTrade and made the entire book public — every position, every fill, every model test, every bug, every failure.
Not paper. Not a backtest screenshot. Real money, documented in real time.
The live story is a blog series. Episode 11 is the current chapter: Grok Bot built separate biotech and semiconductor books, I funded a $13,500 Public account, then rejected the sector-gated construction and reopened the entire design. The final allocator covers 26 thesis-approved companies, passed four of five out-of-sample folds, and replaced the live strategy set on August 26 after the old book was preserved in paper. The exact evidence, deployment and zero-order reconciliation are in episode-11/addendum/EPISODE_11_OPTIONS_PORTFOLIO_REDESIGN_20260826.md.
| Full series → | Eleven episodes and counting: model bakeoffs, deploy day, production bugs, week-one gains, panic sells, engine rewrites, and the open runbook. |
| Episode 1 → | Where it started — why $25k, why Public, why total transparency. |
| Episode 10 → | The full story of Fable 5 running this runbook — every gate, engine bug, and dead end logged in FABLE_CAMPAIGN.MD. |
| Episode 10 on Medium → | Same article, syndicated on Medium — included here so readers who follow the challenge off-platform can find it without hunting. |
| Episode 11 → | Grok Bot’s original thesis books, the funded $13,500 account, and the August 26 redesign that removed XBI and SMH vetoes. |
| Episode 11 on Medium → | The synchronized Medium edition. |
| Episode 11 addendum → | Short-term debit-spread rule, matched walk-forward results, live-book changes, current cadence gates, and Public symbol eligibility. |
| Episode 11 combined-book research → | Cash-account constraint, failed lockbox, biotech-inclusive regime research, execution audit, and the inactive forward candidate. |
| Episode 11 final redesign → | One 26-company allocator, company-level gates, five-fold evidence, full event replay, live deployment, and the zero-order reconciliation. |
This repo is the open playbook. Episode 10 documents one agent’s full run through it. Episode 11 shows the research continuing after deployment, including the rejected designs, the exact replacement and the live reconciliation. The runbook is yours to replay with whatever model you have. The agents don’t just trade the account—they commit to this repo.
Live books
Three Public brokerage books. Each has its own folder. Do not merge them.
| Book | Public | Capital | Live config | Runbook |
|---|---|---|---|---|
| Public Portfolio Challenge: Original | live id 69a7dc7acdb6bf6a4681d36c |
$25,000 | Episode 10 incumbent. Do not touch. | episode-10/ |
| Public Portfolio Challenge: Biotech | former 5OH86568 · historical id 6a5e20a3ea0d6db55c69a171 |
$5,500 historical | Entry-disabled historical record after Public consolidation. | episode-11/moderna/ |
| Public Portfolio Challenge: Semis + Biotech | 5OH79160 · id 6a8cb433e3971b7c87943f11 |
$13,500 | Live 29-strategy implementation: one 26-company allocator, two global exits and 26 company exits. One QGEN call remains held; automatic approval is off. | episode-11/addendum/EPISODE_11_OPTIONS_PORTFOLIO_REDESIGN_20260826.md |
Semis method and the certified SMH comparison live in episode-semis/RUNBOOK.md. Biotech KEEP stays in episode-11/moderna/RUNBOOK.md — do not rewrite or merge that body.
August 24 deployment audit: audits/2026-08-24-biotech-semis-deployment-audit.md records the PSNL, TECH, and PACB cleanup, calendar-matched Biotech control, fresh Semis re-certification, rejected MRNA-core and vertical-first alternatives, Constant-frequency verification, and the completed BTC-dust cleanup.
August 24 cash-account addendum: audits/2026-08-24-public-cash-account-calls-only-addendum.md records the temporary calls-only test, the return comparison that rejected it, restoration of both higher-return spread books, the cooldown diagnosis, and the options-level notable-event fix.
August 24-25 combined-book continuation: episode-11/addendum/COMBINED_LONG_CALL_RESEARCH_20260824.md records the consolidation step and the proposed XBI-regime candidate. That candidate ranked MRNA inside a quality-screened biotech chain and activated semis in the opposite XBI regime. It was never deployed and is now superseded by the August 26 redesign. The company, thesis, and option-access review for the researched stocks is in the portfolio universe audit.
August 26 final redesign: episode-11/addendum/EPISODE_11_OPTIONS_PORTFOLIO_REDESIGN_20260826.md supersedes the proposed XBI-regime candidate. It uses one central allocator across 26 thesis-approved companies, removes XBI and SMH vetoes, and was certified at the actual $13,500 account size. The rejected six-strategy book is preserved in paper portfolio 6a8f9b0ec67c177db82d5b0f. The validated 29-strategy object now runs on the live Public account, and its current-book reconciliation required no orders.
Agent bakeoff
Four agents received the same Episode 10 discipline. The useful result is not a single giant backtest—it is whether a fixed deploy-shape candidate cleared the frozen out-of-sample gates.
| Agent | Best comparable OOS mean return | Passed every gate? | Outcome | Notable failure |
|---|---|---|---|---|
| Claude Code | +53.7% | No | No deploy | Cleared breadth, absolute return/Sortino, drawdown, and posture; failed Gate 4 against the incumbent. |
| Codex | +32.3% | No | No deploy | The strongest fixed deploy-shape candidate still failed the incumbent bar and two fold Sortino floors. |
| Cursor | +33.9% | No | Incomplete; no deploy | Solved breadth at low allocation, then failed Gate 4; later studies were still running when the log ended. |
| Claude Fable 5 | +88.3% | No—owner override | Deployed | Strong return and drawdown, but missed strict breadth, fold-Sortino, stability, and posture gates; later engine fixes weakened the selection-provenance claim. |
Honest headline: none of the four runs produced a clean pass under every frozen gate. Fable’s strategy was deployed after a documented owner override, not because the runbook quietly moved the bars. Follow the links for fold-level evidence and every failure.
What this repo gives you
How the runbook works
The campaign is built around one idea: out-of-sample performance is the only result that counts.
flowchart LR
A["🔍 Searchbacktests & optimization"] --> B["✓ Certifywalk-forward folds"]
B --> C["🔒 Lockboxsingle-touch confirm"]
C --> D["🚀 Deploylive portfolio"]
| Layer | Job |
|---|---|
| Search | Invent and tune candidate strategies fast — variants, sweeps, backtests. |
| Certify | Walk-forward: each fold optimizes in-sample, scores on held-out OOS the optimizer never saw. |
| Lockbox | A final untouched window. One touch. No peeking. |
| Deploy | Clone to a live portfolio, parity-check, attach monitoring. |
Fixed by the runbook: a frozen watchlist (20 names in the Episode 10 bakeoff), $25,000 capital, the fold calendar, the gates, the lockbox rules, and the deploy procedure. Yours to design: signals, structures, deltas, exits, sizing — anything that clears the gates is valid.
What’s inside
The discipline lives in two forms. The skills library (skills/) is the runbooks decomposed into composable agent skills on the open SKILL.md standard — the same folder runs in Claude Code, Codex CLI, Cursor, Gemini CLI, Copilot and ~15 other tools (skills/install.sh handles each tool’s path + adapter). Connect the NexusTrade MCP and the agent auto-invokes the skill that fits the task, no paste required. The episode folders preserve the complete public timeline: historical indexes for Episodes 1–9, then the original runbooks and campaign logs for the reproducible campaigns.
skills/ ← the skills library (install into ANY agent — see skills/README.md)
├── README.md ← index + how the skills compose
├── run-episode/ ← entry point: /run-episode 10 (sequences the stages)
├── portfolio-certification/ ← umbrella orchestrator (the staged PASS/FAIL run)
├── walk-forward-oos/ ← the OOS certification engine
├── breadth-audit/ ← true participation at fixed $25k
├── sweep-reoptimization/ ← re-sweep on structural change + provenance
├── options-structure-rules/ ← spread-shape rule + hard constraints
├── alt-data-indicators/ ← custom indicators (Reddit, congressional, …)
├── bug-protocol/ ← "loudly declare" + hand-off doc template
├── deploy-gate/ ← the gated deploy + reconcile flow
├── engine-sanity/ ← Stage-S0 pre-flight contract checks
├── strategy-bakeoff/ ← the SEARCH→CERTIFY multi-family funnel
└── lockbox-holdout/ ← single-touch holdout + A/B/C baselines
Each episode is a self-contained folder: the runbook to paste, plus the campaign logs from each operator/agent run.
Episodes 1–9 predate the current repository format. Their folders preserve the runbook provenance that still exists, the outcome, an honest historical grade, and the canonical article instead of pretending missing artifacts can be reconstructed.
| Episode | Outcome | Grade | Record |
|---|---|---|---|
| 1 | Opened the real-money challenge and established the public record. | Historical; not scored under current gates | episode-01/ |
| 2 | Converted expert feedback into a more disciplined agent workflow. | Historical; not scored under current gates | episode-02/ |
| 3 | Compared 11 AI models on strategy construction. | Research bakeoff; pre-current gates | episode-03/ |
| 4 | Tested automated hill-climbing and found optimization did not automatically beat the first attempt. | Research result; pre-current gates | episode-04/ |
| 5 | Reached the first live options deployment milestone. | Deployed under episode-era controls | episode-05/ |
| 6 | Documented day-one live behavior. | Live observation; not a certification | episode-06/ |
| 7 | Hit a close-order failure and rebuilt the risk engine around it. | Failure documented and remediated | episode-07/ |
| 8 | Published the week-one gain with the live book visible. | Live snapshot; not forward evidence | episode-08/ |
| 9 | Documented the gain, the panic sell, and the human override risk. | Failure documented | episode-09/ |
episode-10/
├── BAKEOFF_RUNBOOK.md ← paste this into a fresh MCP session
├── RUNBOOK_OG.md ← the original (Episode 1) runbook, kept for reference
├── snapshots/ ← baseline + incumbent seed portfolios the runbook loads
├── addendum/ ← entry/exit redesign, deploy evidence, diagram, and bug note
├── FABLE_CAMPAIGN.MD ← operator run log (Fable 5)
├── CLAUDE_CODE_CAMPAIGN_LOG_*.md ← agent run log (Claude)
├── CODEX_CAMPAIGN_LOG_*.md ← agent run log (Codex)
└── CURSOR_CAMPAIGN_LOG_*.md ← agent run log (Cursor)
| File | What it is |
|---|---|
skills/ |
The skills library. 12 composable agent skills on the open SKILL.md standard (Claude Code, Codex, Cursor, Gemini, Copilot, …) — the whole certification discipline, auto-invoked instead of pasted. Entry point: /run-episode 10. skills/README.md is the index. |
start.py + example_profile.json |
Start here (fast path). python3 start.py walks you through your watchlist + risk tolerance, writes profile.json and a prompt.txt to paste — the agent builds you a personalized strategy. No runbook needed. |
episode-10/BAKEOFF_RUNBOOK.md |
The agent brief you run — walk-forward validation, lockbox, deploy gates. Paste and execute top to bottom. |
episode-10/RUNBOOK_OG.md |
The original Episode-1 runbook, kept for reference (the brief has since expanded). |
episode-10/snapshots/ |
Baseline A/B and incumbent seed portfolios the runbook loads via create_portfolio. |
episode-10/addendum/ |
Episode 10 entry/exit redesign addendum: runbook, campaign evidence, OOS comparison diagram, and bug note. |
episode-10/addendum/RAW_RETURN_VNEXT_RESULTS_20260817.md |
Raw-return vNext sweep, GA, walk-forward, event audit, winners/losers, risk trade-offs, and deployment record. |
episode-11/addendum/SHORT_TERM_DEBIT_SPREADS_20260824.md |
Episode 11 addendum: 180-DTE debit-spread ceiling, live Biotech/Semis replacements, matched walk-forward evidence, and symbol-level Public eligibility. |
episode-11/addendum/COMBINED_LONG_CALL_RESEARCH_20260824.md |
Episode 11 combined-book campaign: long calls only, failed frozen lockbox, corrected biotech-inclusive XBI regime book, universe audit, optimizer cooldown fix, and no-deploy record. |
episode-11/addendum/EPISODE_11_OPTIONS_PORTFOLIO_REDESIGN_20260826.md |
Episode 11 final redesign: actual $13,500 capital, 26 thesis-approved companies, company-level timing, one allocator, five-fold evidence, full event replay, and no live mutation without owner approval. |
| Campaign logs | Per-run logs from each operator/agent: FABLE_CAMPAIGN.MD, CLAUDE_CODE_…, CODEX_…, CURSOR_…. |
episode-11/moderna/RUNBOOK.md |
Biotech KEEP — design-frozen. Live name Public Portfolio Challenge: Biotech. Do not rewrite or merge this body. |
episode-semis/RUNBOOK.md |
Semiconductor S13 A — SMH-bar walk-forward, exam-only lockbox. Now running on Public 5OH79160. |
episode-semis/
├── README.md ← index for the Semis Public book
├── RUNBOOK.md ← paste this; S13 A is design-frozen
└── CAMPAIGN_LOG.md ← facts already recorded (no invented fills)
Episode 11 is the Biotech KEEP book at episode-11/moderna/ — Public Portfolio Challenge: Biotech (Public 5OH86568). The semiconductor book is a separate episode at episode-semis/ — Public Portfolio Challenge: Semis (Public 5OH79160), now running S13 A. Do not steal the Episode 11 folder for chips.
Community leaderboard
Think your agent can beat the incumbent without moving the gates? Prove it. Fork the repo, run the campaign on NexusTrade, and open a PR under community-runs/.
| Rank | Run | Agent | OOS return | OOS Sortino | Worst max drawdown | Gates | Evidence |
|---|---|---|---|---|---|---|---|
| — | No verified community runs yet | — | — | — | — | — | Submit the first run |
Only runs that pass every current gate are ranked by mean OOS return. Failed runs stay visible—the point is reproducibility, not survivor bias. Start with the submission guide and result template.
Get started
Step 1 — Developers page
Open nexustrade.io/developers.
Step 2 — Create a free account
You’ll need a NexusTrade account to authorize MCP and access portfolios, backtests, and live trading tools.
Step 3 — Connect your AI tool
Recommended: OAuth. No keys to copy, rotate, or leak. Sign in once in the browser when your client first calls a NexusTrade tool.
https://nexustrade.io/api/mcp
Cursor (recommended)
- On the Developers page, expand API Keys.
- Under Connect an AI tool to NexusTrade, click Add to Cursor.
- OAuth runs automatically on first tool use.
Step 4 — Run it
Fast path — a personalized strategy in one paste. Run:
python3 start.py
On the first run it walks you through a few questions (defaults come from example_profile.json — mine) and writes your profile.json, then writes your ready-to-paste prompt to prompt.txt. Open prompt.txt, copy it, and paste into a fresh MCP-connected chat. Run it again any time after editing profile.json to regenerate the prompt.
No Python? Copy example_profile.json to profile.json, edit it, then paste the JSON into your agent with one line: “Build me a personalized strategy from this profile, backtest it out-of-sample, and ask before deploying.”
Either way, the agent designs a strategy on your names, backtests it, compares it to buy-and-hold, and asks before risking a dollar.
Full rigor — the runbook. Want walk-forward validation, a held-out lockbox, and deploy gates? Open episode-10/BAKEOFF_RUNBOOK.md, paste the entire file into a fresh session, and tell the agent: execute top to bottom, do not ask clarifying questions. Log your run alongside the per-agent campaign logs in episode-10/.
Risk disclaimer
This repository is educational and documents one person’s experiments. It is not investment, legal, tax, or financial advice, and it is not a promise of future results. Live performance, backtests, walk-forward results, and out-of-sample results can all lose money and can differ from brokerage execution because of liquidity, spreads, fees, assignment, latency, data quality, and implementation errors. Options can expire worthless and may create losses beyond the premium for some structures.
Nothing in this repository should place a trade by itself. Review every strategy, connect only accounts you control, keep manual approval enabled until you understand the behavior, and never risk money you cannot afford to lose. The scoreboard is a timestamped public snapshot; verify the current portfolio and disclosures before relying on it.
More links
| Live portfolio | Positions and P&L in real time |
| Copy the live incumbent | Open the marketplace copy/deploy flow; review before connecting real money |
| Blog series | The full documented journey |
| Episode 1 | How the challenge began |
| Episode 10 | Fable 5 ran this runbook and deployed a live book — full story + links to campaign logs |
| Episode 10 on Medium | Syndicated copy of the same article (for readers off NexusTrade) |
| Episode 11 | Grok Bot’s biotech and semiconductor experiment, plus the certified August 26 portfolio redesign |
| Episode 11 on Medium | Synchronized Medium edition |
| Developers | MCP setup |
| MCP tools reference | Every tool the runbook can call |
| API overview | REST + auth |
Рекомендуемые инструменты
Попробуйте другой запрос или уберите фильтр.
Установка
npx skillfish add austin-starks/public-portfolio-challenge


