
blueowl1633-no7j5/eval-rubric
Productivity workflowDesign evaluation rubrics for LLM/agent outputs: criteria, scores, failure modes, examples.
Обзор
Design evaluation rubrics for LLM/agent outputs: criteria, scores, failure modes, examples. - You build rubrics that make model quality measurable and comparable - Scope: See SKILL.md vertical lock. - Deliverable: See SKILL.md deliverable. - Triggers: eval rubric, LLM eval, grading rubric, agent eval, 评测量表. Local monorepo: point your agent at skills/eval-rubric/ (this folder). Runtime output language follows lang (default ). Skill docs stay EN-primary. Full workflow, hard rules, and QA: see SKILL.md. Monorepo sample (if present): ../../output/samples/eval-rubric-example-01.md
README
Eval Rubric
Design evaluation rubrics for LLM/agent outputs: criteria, scores, failure modes, examples.
中文简介: 你做可复现、可对比的模型输出评测表。 — 详见 README.zh.md.
What this skill does
- You build rubrics that make model quality measurable and comparable
- Scope: See SKILL.md vertical lock.
- Deliverable: See SKILL.md deliverable.
- Triggers: eval rubric, LLM eval, grading rubric, agent eval, 评测量表.
Install
After this package is on GitHub:
npx skills add / --skill eval-rubric
Local monorepo: point your agent at skills/eval-rubric/ (this folder).
Options
| Option | Values | Default |
|---|---|---|
lang |
en · zh · bilingual |
en |
scale |
1-5 · 0-2 · pass-fail |
1-5 |
weighting |
equal · custom |
equal |
Runtime output language follows lang (default en). Skill docs stay EN-primary.
When to use
| Use | Don’t |
|---|---|
| (see bullets) | (see bullets) |
Use
- When you need
eval-rubricworkflow
Don’t
- Out-of-scope tasks listed in SKILL.md
Deliverable
See SKILL.md deliverable.
Full workflow, hard rules, and QA: see SKILL.md.
Example
Monorepo sample (if present): ../../output/samples/eval-rubric-example-01.md
License
MIT — see LICENSE.
Рекомендуемые инструменты
Попробуйте другой запрос или уберите фильтр.
Установка
npx skillfish add blueowl1633-no7j5/eval-rubric