An AI product analyst built on Claude Code. You ask a business question, it runs a pipeline of 18 agents that frame the question, explore your data, find the root cause, build a narrative, and hand...
Overview
An AI product analyst built on Claude Code. You ask a business question, it runs a pipeline of 18 agents that frame the question, explore your data, find the root cause, build a narrative, and hand...
README
AI Analyst v2
An AI product analyst built on Claude Code. You ask a business question, it runs a pipeline of 18 agents that frame the question, explore your data, find the root cause, build a narrative, and hand you a validated slide deck with speaker notes. Minutes, not weeks.
18 specialized agents | 39 auto-applied skills | 20 slash commands | DAG-based parallel execution | PDF + HTML export
Before You Start
This is a tool for analysts, not a replacement for them. It handles about 80% of what a human analyst does. The 80% that takes all the time. But it only works if you’re the expert.
You are the eval. Run this on data you know like the back of your hand. Run it on the reports you were already going to run this week. When it picks the wrong column or misinterprets a metric, you’ll catch it immediately because you’ve written that query before. You correct it, it saves the correction, and it doesn’t make that mistake again. That’s the whole loop. Look, know, correct, move on.
Don’t hand this to someone who can’t validate the output. Don’t run it on data you’ve never seen. The analyses it produces need your judgment before they go anywhere near a stakeholder. If you skip the validation, you’ll get confident-sounding numbers that might be wrong. If you do the validation, you’ll move faster than you ever have.
The byproduct of building this is the work itself. You’re not taking time off from your job to set up an AI tool. You’re doing your actual work through it. The first analysis takes a bit longer because you’re connecting data and teaching it your context. By the third one, you’re faster than doing it by hand. By next week, you’re doing 15 analyses instead of 5.
This doesn’t work out of the box. It’s a starting point, not a finished product. The model capability is there with Opus 4.6, but you need to teach it your data, your metrics, your business context. Correct it when it’s wrong. Grow it into something that works for your specific use case, or tear it apart and rebuild it how you want. The agents, skills, and pipeline are all markdown files you can read and modify. Nothing is hidden.
Bring your own data. No bundled datasets. Connect your CSVs, DuckDB, Postgres, BigQuery, or Snowflake with /connect-data and start analyzing.
What’s New in V2
V2 is a ground-up rebuild of the intelligence layer. The pipeline and agents from V1 still work the same way — you won’t notice a difference in how you use it. What changed is everything underneath.
| Area | V1 | V2 |
|---|---|---|
| Data | Bundled NovaMart e-commerce dataset | Bring your own — CSV, DuckDB, Postgres, BigQuery, Snowflake |
| Onboarding | Manual setup, read the docs | /setup interview learns your role, data, and business context |
| Memory | Stateless across sessions | Knowledge system persists corrections, learnings, query patterns, business glossary |
| Self-learning | None | Captures feedback, logs corrections, retrieves proven SQL patterns — never repeats the same mistake |
| Theming | Hardcoded chart style | YAML-based theme system with brand colors, WCAG-compliant palettes |
| Business context | None | Organization knowledge base — glossary, metrics, products, teams. Notion ingest. |
| Pipeline | Single run, restart on failure | Run tracking (/runs), reliable resume, comms drafter for Slack/email output |
| Testing | Minimal | 606 tests with synthetic fixtures, no external data dependencies |
| Dataset coupling | NovaMart table names hardcoded in agents | Fully dataset-agnostic — agents resolve from active manifest and schema |
Don’t Know What to Do? Just Ask.
Claude knows the entire system — every agent, skill, command, and dataset. If you’re stuck, ask it:
What can I do with this data?
What should I run to refresh the deck?
How do I connect my own CSV files?
Which agents handle root cause analysis?
Re-run just the chart maker and deck creator.
Claude will tell you the exact command. You don’t need to memorize anything in this README. Think of it as a reference — Claude is the guide.
Quick Start
1. Install Claude Code (requires a Claude Pro subscription)
npm install -g @anthropic-ai/claude-code
2. Clone and set up
git clone https://github.com/ai-analyst-lab/ai-analyst.git
cd ai-analyst
pip install -e ".[dev]"
3. Start Claude Code
claude
4. Connect your data and go
/connect-data
Or skip the wizard and just ask a question with your data in a directory:
/run-pipeline data_path=data/my_csvs/ question="Why is conversion dropping?"
For full setup details: docs/setup-guide.md
Five Things You Can Do
1. Ask a quick question
What's our conversion rate by device?
Claude queries the data and returns an answer with a chart. Simple questions get answered in under 2 minutes without running the full pipeline.
2. Run a full analysis
/run-pipeline data_path=data/your_dataset/ question="What's driving the decline in conversion?"
The pipeline runs 18 agents across 4 phases: Frame the question, Analyze the data, Build the story, Create the deck. You get a validated analysis, branded charts, a narrative, and a slide deck with speaker notes. Exports to PDF and HTML.
3. Explore a dataset
/explore
Interactive data browsing without committing to a full analysis. Preview tables, check distributions, spot patterns, form hypotheses. Use /data users to inspect a specific table’s schema.
4. Connect your own data
/connect-data
Guided wizard that walks you through connecting CSV files, local DuckDB, Postgres, BigQuery, or Snowflake. Auto-profiles your data, creates schema docs, and remembers your dataset context across sessions.
5. Make a single chart
Make a funnel chart of the checkout flow, highlighting the biggest drop-off step.
Claude generates a chart following Storytelling with Data methodology: warm off-white background, decluttered axes, action title, direct labels instead of legends.
How It Works: The Pipeline
When you run /run-pipeline, Claude orchestrates 18 agents across 4 phases:
1. FRAME 2. ANALYZE 3. STORY 4. DECK
+-----------------+ +-----------------------------+ +--------------------+ +------------------+
| Question | | Data Explorer | | Story Architect | | Storytelling |
| Framing | | > Source Tie-Out | | > Coherence | | > Deck Creator |
| > Hypothesis | | > Descriptive Analytics | | Reviewer | | > Slide Review |
| Generation | | > Root Cause Investigator | | > Chart Maker | | > Close the |
| |-->| > Validation |-->| > Design Critic |-->| Loop |
+-----------------+ | > Opportunity Sizer | +--------------------+ +------------------+
+-----------------------------+
Phase 1 — Frame: Structures your business question into analytical questions with testable hypotheses. Checkpoint: review the framing before analysis begins.
Phase 2 — Analyze: Explores the data, verifies loading integrity, runs segmentation/funnel/drivers analysis, drills down to root cause, validates findings, and sizes the opportunity. Checkpoint: automated quality gate.
Phase 3 — Story: Designs a storyboard (Context-Tension-Resolution arc), generates charts with collision detection, and reviews visual quality against a 16-point checklist.
Phase 4 — Deck: Writes a stakeholder narrative, builds a branded Marp slide deck with HTML components, reviews slide design, and ensures every recommendation has a follow-up plan. Exports to PDF and HTML.
You don’t have to run the whole thing. Five execution plans let you run just the part you need:
| Plan | Use When | What Runs |
|---|---|---|
full_presentation |
Complete analysis to slide deck | All 18 agents |
deep_dive |
Analysis without presentation | Phases 1-2 only |
quick_chart |
Just need one chart | Chart Maker + Design Critic |
refresh_deck |
Re-do the presentation layer | Phases 3-4 (reuses analysis) |
validate_only |
Check existing work | Validation + Source Tie-Out |
/run-pipeline data_path=data/your_dataset/ question="..." plan=deep_dive
If the pipeline gets interrupted, resume where you left off:
/resume-pipeline
Preview what would run without executing:
/run-pipeline data_path=data/your_dataset/ question="..." dry-run=true
How It Works: The DAG Engine
The pipeline doesn’t run agents one at a time. It resolves dependencies automatically and runs independent agents in parallel:
Tier 0 (parallel) Question Framing -----> Hypothesis
Data Explorer --------> Source Tie-Out
|
Tier 2 (parallel) Descriptive Analytics / Overtime Trend / Cohort Analysis
|
Tier 3 (sequential) Root Cause --> Validation --> Opportunity Sizer
|
Tier 4 (sequential) Story Architect --> Coherence Review
|
Tier 5 (parallel fan-out) Chart Maker (per beat) --> Design Critic
|
Tier 6 (sequential) Storytelling --> Deck Creator --> Slide Review --> Close the Loop
- Parallel execution: Agents in the same tier run concurrently (up to 3 at once). Tier 0 starts Question Framing and Data Explorer simultaneously.
- Automatic dependency resolution: The engine reads
agents/registry.yamland computes execution tiers using topological sort. - Circuit breaker: If 3 agents fail in the same tier, the pipeline halts with a diagnostic report.
- Timeouts: Each agent gets 5 minutes. One retry on timeout. Critical agents (source tie-out, validation) halt the pipeline; non-critical agents (design critic) degrade gracefully.
- Checkpoints: Quality gates between phases. Two are automated (analysis verification, final deck lint). Two are user-facing (frame review, storyboard review). Say “just do it” to skip the user-facing ones.
All Commands
| Command | What It Does | Example |
|---|---|---|
/run-pipeline |
Full analysis to slide deck | /run-pipeline data_path=data/your_dataset/ question="Why is conversion dropping?" |
/resume-pipeline |
Resume interrupted pipeline | /resume-pipeline |
/explore |
Interactive data exploration | /explore events |
/data |
Show active dataset schema | /data users |
/datasets |
List all connected datasets | /datasets |
/switch-dataset |
Change the active dataset | /switch-dataset my_dataset |
/connect-data |
Add a new data source | /connect-data |
/setup |
Interactive onboarding interview | /setup |
/metrics |
Browse the metric dictionary | /metrics conversion_rate |
/history |
View past analyses | /history |
/patterns |
View recurring patterns | /patterns --global |
/export |
Export results in various formats | /export slides or /export email or /export slack |
/forecast |
Generate a time-series forecast | /forecast |
/runs |
List, inspect, compare pipeline runs | /runs |
/business |
Browse organization knowledge | /business glossary |
/log-correction |
Log a data or methodology correction | /log-correction |
/architect |
Multi-persona planning methodology | /architect |
/notion-ingest |
Import business context from Notion | /notion-ingest |
/compare-datasets |
Compare metrics across datasets | /compare-datasets |
/setup-dev-context |
Add codebase context for dev teams | /setup-dev-context |
Or just ask in plain English. “Show me conversion by device” works as well as any command.
Charts and Visualization
Every chart follows the Storytelling with Data methodology:
Your Data --> chart_helpers.py --> Base Chart (150 DPI)
|
Collision Check
(3 fix strategies)
|
Marp Deck (HTML components)
|
marp_linter.py (8 check categories)
|
marp_export.py --> PDF + HTML
What happens automatically:
swd_style()applies warm off-white background (#F7F6F2), removes chart clutter (gridlines, borders, redundant legends), sets consistent typography- Every chart gets an action title (takeaway statement, not a label) and a subtitle (data source, time range)
- Direct labels replace legends wherever possible
- Collision detection checks for overlapping text with 3 auto-fix strategies: offset the label, reduce font size, or drop the least important label. Charts with unresolved collisions halt the pipeline.
- The deck uses branded HTML components: KPI cards, finding cards, recommendation rows, so-what callouts, before/after panels, timelines, and more
- A lint gate validates every deck before export: checks frontmatter completeness, HTML component usage (minimum 3 types), valid slide classes, slide count, and pacing
- YAML-based theming with brand color overrides and WCAG-compliant palettes (see docs/theming.md)
Your Data
This repo ships clean — no bundled datasets. Connect your own data and the system builds context around it.
Connect your own
Run /connect-data for a guided setup wizard, or /setup for a full onboarding interview. Supported sources:
- CSV files — drop them in a directory, point Claude at it
- DuckDB — local or MotherDuck
- Postgres — any Postgres-compatible database
- BigQuery — Google BigQuery with service account
- Snowflake — Snowflake with user/password or key pair
The system auto-profiles your data, creates schema documentation, notes data quirks, and remembers context across sessions in .knowledge/datasets/.
Example datasets
Curated public datasets with README guides are available in data/examples/.
Fallback chain
If your primary connection fails, the system falls back automatically:
- Primary connection (e.g., MotherDuck via MCP)
- Local DuckDB (from
manifest.local_data.duckdb) - CSV files via pandas (from
manifest.local_data.path)
You’re always told which source is active.
What Just Happened? (Output Guide)
After running a pipeline, here’s what you’ll find:
outputs/
question_brief_YYYY-MM-DD.md # Your question, structured
hypothesis_doc_YYYY-MM-DD.md # Testable hypotheses
data_inventory_YYYY-MM-DD.md # What data exists
analysis_report_YYYY-MM-DD.md # Full analysis with findings
validation__YYYY-MM-DD.md # Independent validation of findings
narrative__YYYY-MM-DD.md # Stakeholder-ready story
deck__YYYY-MM-DD.marp.md # Slide deck (Marp source)
deck__YYYY-MM-DD.pdf # PDF export
deck__YYYY-MM-DD.html # HTML export (self-contained)
close_the_loop_YYYY-MM-DD.md # Follow-up plan for recommendations
charts/ # All generated charts
working/ # Intermediate files (safe to delete)
pipeline_state.json # Pipeline progress (for /resume-pipeline)
pipeline_metrics.json # Execution timing and parallel efficiency
storyboard_.md # Story beats + visual mapping
design_review_.md # Chart quality review (16-point checklist)
investigation_.md # Root cause drill-down log
sizing_*.md # Opportunity sizing with sensitivity analysis
outputs/ contains your deliverables. working/ contains intermediate artifacts that support resumability and debugging.
Customization
| Want to… | Do this |
|---|---|
| Change how Claude thinks | Edit CLAUDE.md (the AI’s persona, rules, workflow) |
| Add a new skill | Create .claude/skills/my-skill/skill.md, reference it in CLAUDE.md |
| Add a new agent | Create agents/my-agent.md using agents/CONTRACT_TEMPLATE.md as a starting point |
| Change the slide theme | Create a YAML theme in themes/brands/ (see docs/theming.md) |
| Add deck components | Edit templates/marp_components.md (snippet library) |
| Modify the pipeline | Edit .claude/skills/run-pipeline/skill.md (rules, checkpoints, execution) |
| Add to the agent DAG | Edit agents/registry.yaml (dependencies, execution order) |
Requirements
- Python 3.10+
- Node.js 18+ (for Claude Code)
- Claude Code with a Claude Pro subscription ($20/month)
- Internet connection (for Claude API and optional MotherDuck)
Getting Help
- Setup guide: docs/setup-guide.md
- Theming: docs/theming.md
- Questions or bugs: Open a GitHub Issue
License
MIT – use it however you want.
Recommended Tools
Try a different keyword or remove a filter.
Install
npx skillfish add ai-analyst-lab/ai-analyst