BN

blueowl1633-no7j5/eval-rubric

效率工作流
55 stars 质量 40 趋势 40

Design evaluation rubrics for LLM/agent outputs: criteria, scores, failure modes, examples.

概览

Design evaluation rubrics for LLM/agent outputs: criteria, scores, failure modes, examples. - You build rubrics that make model quality measurable and comparable - Scope: See SKILL.md vertical lock. - Deliverable: See SKILL.md deliverable. - Triggers: eval rubric, LLM eval, grading rubric, agent eval, 评测量表. Local monorepo: point your agent at skills/eval-rubric/ (this folder). Runtime output language follows lang (default ). Skill docs stay EN-primary. Full workflow, hard rules, and QA: see SKILL.md. Monorepo sample (if present): ../../output/samples/eval-rubric-example-01.md

README

Eval Rubric

Design evaluation rubrics for LLM/agent outputs: criteria, scores, failure modes, examples.

中文简介: 你做可复现、可对比的模型输出评测表。 — 详见 README.zh.md.

What this skill does

  • You build rubrics that make model quality measurable and comparable
  • Scope: See SKILL.md vertical lock.
  • Deliverable: See SKILL.md deliverable.
  • Triggers: eval rubric, LLM eval, grading rubric, agent eval, 评测量表.

Install

After this package is on GitHub:

npx skills add / --skill eval-rubric

Local monorepo: point your agent at skills/eval-rubric/ (this folder).

Options

Option Values Default
lang en · zh · bilingual en
scale 1-5 · 0-2 · pass-fail 1-5
weighting equal · custom equal

Runtime output language follows lang (default en). Skill docs stay EN-primary.

When to use

Use Don’t
(see bullets) (see bullets)

Use

  • When you need eval-rubric workflow

Don’t

Deliverable

See SKILL.md deliverable.

Full workflow, hard rules, and QA: see SKILL.md.

Example

Monorepo sample (if present): ../../output/samples/eval-rubric-example-01.md

License

MIT — see LICENSE.

View this README on GitHub

推荐工具

换一个关键词,或者移除筛选条件。

安装

npx skillfish add blueowl1633-no7j5/eval-rubric