Skills that guide AI coding agents to help you build product-specific AI evals.
개요
Skills that guide AI coding agents to help you build product-specific AI evals (not foundation model benchmarks). These skills guard against common mistakes we've seen helping 50+ companies and teaching thousands of students in our AI Evals course. There are many easily avoidable footguns in evals. These skills help you avoid them. is the entry point. It looks at your situation and routes you to the right skill. Most of the time it will send you to one of these two: - , if you already have an eval pipeline. It inspects your setup and recommends next steps. - , if you have traces but haven't analyzed them yet. It builds an annotation interface and helps you sample traces. Watch Shreya's walkthrough. The error-discovery skill in the plugin is the most important. Error discovery involves qualitative and quantitative analysis of your traces to find failure modes. You should only write evals after doing this step. This skill makes an AI agent run error analysis on a dataset.
README
Eval Skills
Skills that guide AI coding agents to help you build product-specific AI evals (not foundation model benchmarks).
These skills guard against common mistakes we’ve seen helping 50+ companies and teaching thousands of students in our AI Evals course.
Why skills for evals
There are many easily avoidable footguns in evals. These skills help you avoid them.
evals-start is the entry point. It looks at your situation and routes you to the right skill. Most of the time it will send you to one of these two:
- eval-audit, if you already have an eval pipeline. It inspects your setup and recommends next steps.
- error-discovery, if you have traces but haven’t analyzed them yet. It builds an annotation interface and helps you sample traces. Watch Shreya’s walkthrough.
Installation
Install with npx skills:
npx skills add https://github.com/ai-evals-course/evals-skills
Install one skill only:
npx skills add https://github.com/ai-evals-course/evals-skills --skill error-discovery
Check for updates:
npx skills check
npx skills update
Available skills
| Skill | What it does |
|---|---|
| evals-start | Entry point. Routes to the skill that matches your situation |
| eval-audit | Audit an eval pipeline and surface problems with prioritized severity |
| error-discovery | Build a review app, select diverse samples, and organize your notes into failure modes |
| generate-synthetic-data | Create diverse synthetic test inputs using dimension-based tuple generation |
| write-code-eval | Write code checks for objective failure modes |
| write-judge-prompt | Design LLM-as-Judge evaluators for subjective quality criteria |
| validate-evaluator | Calibrate LLM judges against human labels using data splits, TPR/TNR, and bias correction |
| evaluate-rag | Evaluate retrieval and generation quality in RAG pipelines |
| build-review-interface | Build custom annotation interfaces for human trace review |
The error-discovery skill
The error-discovery skill in the plugin is the most important. Error discovery involves qualitative and quantitative analysis of your traces to find failure modes. You should only write evals after doing this step.
This skill makes an AI agent run error analysis on a dataset. Point it at a JSONL/CSV/JSON file of LLM outputs or traces, and it:
- Reads the dataset and figures out the content type (articles, agent traces, code, structured output, etc.).
- Designs visual encoding based on what varies in the data. Uses Gestalt principles (color for categories, spacing for hierarchy, opacity for importance).
- Builds a single-file HTML review app served by a Python stdlib server. No dependencies.
- Clusters the data and picks a diverse initial sample (cluster reps + random picks).
- Runs an interactive loop: monitors annotations, categorizes failure modes, proposes new samples to increase coverage.
You read and leave free-text notes. The agent sorts them into failure modes, tracks coverage, and picks new samples to fill gaps.
Once installed, point the agent at a dataset:
Can you help me do error analysis on traces.jsonl?
Write your own skills
These skills address common eval mistakes. Start here, then write skills for your own data and domain. Matt Pocock’s writing-for-agents skill explains how to write instructions for agents.
Beyond these skills
These skills cover eval work that applies across projects. The AI Evals course also covers production monitoring, regression suites, and cost.
추천 도구
다른 키워드를 입력하거나 필터를 제거해 보세요.
설치
npx skillfish add ai-evals-course/evals-skills