Adversarial, deterministic benchmarking for AI agent skills — measures whether a skill makes the agent measurably better. Includes biomedical eval cases.
概要
Adversarial benchmarking for AI agent skills. We measure whether a skill makes the AI measurably better than not having it. No vibes. No star counts. Deterministic, reproducible measurement. For each skill, two agents solve the same tasks: 1. — agent has the skill's SKILL.md loaded 2. — baseline agent with no skill guidance Both sides are graded by the — never by an LLM. If the skill doesn't consistently beat the baseline, it's not adding value. = Recommended | = Acceptable | = Needs Improvement Pre-built eval libraries for biomedical/bioinformatics skills: - evals/biomedical/blast-alignment.yaml — BLAST search and E-value interpretation - evals/biomedical/variant-calling.yaml — VCF interpretation and variant filtering - evals/biomedical/drug-interaction.yaml — Drug-drug interactions and pharmacokinetics - evals/biomedical/protein-folding.yaml — PDB structure and AlphaFold queries - evals/biomedical/differential-expression.
README
SkillBench
Adversarial benchmarking for AI agent skills. We measure whether a skill makes the AI measurably better than not having it.
No vibes. No star counts. Deterministic, reproducible measurement.
How it works
For each skill, two agents solve the same tasks:
- With-skill — agent has the skill’s SKILL.md loaded
- Without-skill — baseline agent with no skill guidance
Both sides are graded by the same deterministic script — never by an LLM. If the skill doesn’t consistently beat the baseline, it’s not adding value.
Scoring (0-100)
| Dimension | Points | What it measures |
|---|---|---|
| Correctness | /40 | % of deterministic assertions passing (with-skill side) |
| Security | /20 | Static analysis — hardcoded keys, injection, exfiltration |
| Completeness | /20 | % of eval cases where with-skill produced output |
| Robustness | /20 | % of eval cases where with-skill beats without-skill |
75+ = Recommended | 50-74 = Acceptable | <50 = Needs Improvement
Quick start
pip install -e .
skillbench scan # Security scan only
skillbench run # Set up benchmark round
skillbench grade # Grade completed round
skillbench report # Generate HTML + JSON report
With eval cases
skillbench run --evals evals/biomedical/blast-alignment.yaml
Biomedical eval cases
Pre-built eval libraries for biomedical/bioinformatics skills:
evals/biomedical/blast-alignment.yaml— BLAST search and E-value interpretationevals/biomedical/variant-calling.yaml— VCF interpretation and variant filteringevals/biomedical/drug-interaction.yaml— Drug-drug interactions and pharmacokineticsevals/biomedical/protein-folding.yaml— PDB structure and AlphaFold queriesevals/biomedical/differential-expression.yaml— DESeq2 workflow and result interpretation
Key discovery
LLM self-grading hallucinates. When we let LLM agents grade their own outputs, they fabricated quality assessments — inventing metrics they never computed and grading identical outputs differently depending on which “side” they were evaluating. This is why SkillBench uses deterministic grading only.
Target skill repos
| Repo | Skills | Domain |
|---|---|---|
| OpenClaw-Medical-Skills | 869 | Clinical, genomics, drug discovery |
| ClawBio | 39 | Bioinformatics, local-first |
| claude-scientific-skills | 170 | 250+ databases, PubMed/UniProt/PDB |
| LabClaw | 144 | Genomics, proteomics, clinical |
License
MIT
推奨ツール
別のキーワードを試すか、フィルタを外してください。
インストール
npx skillfish add boheling/skillbench