AR

amelnagdy/review-skills

Developer tools
131 stars 品質 40 トレンド 40

Two-model debate review of PRs and MRs, plus a babysitter that works the review rounds. Skills for any coding agent.

概要

debate-review has a main reviewer read the PR, a second reviewer try to knock its findings down, and the main reviewer make the final call. One review with inline comments lands on the PR, posted from your own gh, glab, or az account. babysit-pr then works the rounds: verifies each finding, fixes the blockers, replies in-thread, resolves, and re-triggers the next review. Both skills run through whatever coding agent you already drive (Claude Code, Codex, Cursor, OpenCode, Grok, and others) and the model subscriptions you already pay for. - Node 18 or newer. - Git 2.31 or newer for Azure DevOps (--config-env keeps the access token out of command arguments). - gh (GitHub), glab (GitLab, including self-hosted) or az (Azure DevOps) logged in to an account that can comment on the PR. Reviews and replies post as that account. Azure DevOps needs no az extension: the script uses az rest against the REST API.

README

review-skills

Two models argue over a pull request before anything is posted. You keep the merge.

debate-review has a main reviewer read the PR, a second reviewer try to knock its findings down, and the main reviewer make the final call. One review with inline comments lands on the PR, posted from your own gh, glab, or az account. babysit-pr then works the rounds: verifies each finding, fixes the blockers, replies in-thread, resolves, and re-triggers the next review. Both skills run through whatever coding agent you already drive (Claude Code, Codex, Cursor, OpenCode, Grok, and others) and the model subscriptions you already pay for.

npx skills add amElnagdy/review-skills

Then ask your agent:

Use $debate-review on https://github.com/owner/repo/pull/123
Use $debate-review --local on this repo before I open a PR.
Use $babysit-pr on PR 123 until it is ready to merge.
flowchart LR
  P["PR or MR"] --> M["Main reviewerlane review-main"]
  M -->|"findings"| D["Debate reviewerlane review-debate"]
  D -->|"confirm / refute / downgrade / add"| F["Main reviewerfinal call"]
  F --> R["One reviewinline P0 / P1 / P2"]
  R --> B["$babysit-prverify, fix, reply, resolve"]
  B -->|"push, re-run"| M

The skills

Skill Job Never does
debate-review Reviews a GitHub PR, GitLab MR or Azure DevOps PR with two models in sequence and posts non-approval inline comments plus a summary. Findings that survive the debate are posted as agreed; findings the second model refuted but the main reviewer kept are posted as contested, with both sides’ reasoning. Edit code, approve, request changes, post twice for the same head sha.
babysit-pr Harvests every reviewer thread on the PR (debate-review, Codex, Greptile, any bot), checks each finding against the code, fixes what is real, replies in-thread with evidence and attribution, resolves, and re-runs the review for the next round. Reports when the PR meets the merge gate. Merge, resolve a thread it did not answer, reply as anyone other than “model on behalf of user”.

Requirements

  • Node 18 or newer.

  • Git 2.31 or newer for Azure DevOps (--config-env keeps the access token out of command arguments).

  • gh (GitHub), glab (GitLab, including self-hosted) or az (Azure DevOps) logged in to an account that can comment on the PR. Reviews and replies post as that account. Azure DevOps needs no az extension: the script uses az rest against the REST API.

  • delegate-skills, which dispatches the reviewer models through the *-delegate relays, read-only.

  • Two delegate lanes named review-main and review-debate. Create them once with $delegate-setup:

    Use $delegate-setup to add two lanes: review-main on claude at high effort and review-debate on codex at high effort.
    

    Pick two different implementers. The debate is only worth something when the second model does not share the first one’s blind spots. Keep these lanes for the reviewer; if you already use a debate lane for plan debates, leave it alone. Only implementers whose relay supports --read-only are accepted, so the reviewers cannot touch your tree.

What a review looks like

Each finding is one inline comment, anchored to the lines it is about, rendered as a forge alert so the colour reads before the words:

Level Meaning Rendered as
P0 blocking, security [!CAUTION]
P1 blocking, anything else [!WARNING]
P2 non-blocking [!NOTE]

The review body carries a level count table, who reviewed, and a short summary. Each comment and the body start with an HTML marker (``) that babysit-pr uses to recognise the threads, since they are posted from your account rather than a bot’s. One review per head sha: a re-run on the same push exits with code 3 instead of posting again. Templates and the marker format are in comment-format.md; the JSON contract between the three passes is in schema.md.

The reviewers are prompted for precision, not volume: blocking bugs with a concrete trigger and wrong result, spec and standards violations quoted against the rule, and little else. The second reviewer can only refute a finding when it can point at the code that makes it impossible. Findings below the confidence floor are dropped before anything is posted.

GitHub, GitLab and Azure DevOps

GitHub GitLab Azure DevOps
Review posting gh api, PR review with inline comments glab api, MR discussions with diff positions az rest, one comment thread per finding plus a closed summary thread
Target URL /pull/ /-/merge_requests/ /_git//pullrequest/, on dev.azure.com or *.visualstudio.com
Spec source (#123) issue issue work item
Alert colours yes 17.10+ no, alerts fall back to plain quotes
Thread harvest (babysit-pr) GraphQL review threads + REST reviews discussions + notes not implemented yet
Reply and resolve (babysit-pr) verified live implemented, not yet verified on a live instance not implemented yet
Bot author detection reliable (Bot type) only when the instance exposes author.bot; debate-review threads are found by marker either way n/a

Azure DevOps has no single review object, so one debate-review is N inline threads plus one closed summary thread carrying the marker. --force and the “already reviewed” check read the same marker back off the PR’s threads, so a re-run on an unchanged head still exits 3.

Local preview

--local reviews the files on disk (committed, uncommitted, and untracked, honoring .gitignore) against a base branch. It never calls a forge CLI. --dry-run still needs a live PR; it only skips the post. The two flags do not combine. Local snapshots reject non-UTF-8 Git paths instead of silently changing their bytes.

Invocation Source Forge Post
--local Working tree snapshot No No
--dry-run Live PR Yes No
`` Live PR Yes Yes

Run it by hand

Your agent normally runs these for you. They are here for testing, CI, or when there is no agent in the loop. `` is the directory containing the skill’s SKILL.md.

node "/scripts/review-pr.mjs" --local                  # working tree; print, no forge
node "/scripts/review-pr.mjs"  --dry-run   # print, do not post
node "/scripts/review-pr.mjs"              # post
"/scripts/threads.sh"                       # harvest one round as JSON

review-pr.mjs --help lists the flags: --local to review the working tree with no forge, --main / --debate to override the lanes for one run, --contested post|drop, --min-confidence, --timeout (default 30 minutes per reviewer), --force to post again on the same head, --keep to leave the temporary worktree or snapshot clone. Every run leaves its briefs, raw model output, and the three JSON documents under ~/.cache/debate-review/__/// so a surprising review can be traced back to the pass that produced it.

How this relates to the sibling repos

  • delegate-skills is the transport. review-pr.mjs reads your lanes and calls the matching relay; this repo has no model code of its own.
  • guard-skills are review lenses for code an agent just wrote. The reviewers here cannot load them (the relays run the reviewer with skills disabled, on purpose), and they are tuned for a different job than Clean Code checks. Where a guard fits is the babysitter’s own fixes: babysit-pr runs the matching guard on a fix before pushing it, when one is installed.

Development

node --test test/*.test.mjs

The babysit tests run threads.sh end to end against fake gh and glab binaries in test/fixtures/babysit/. The debate-review tests drive the Azure forge functions against a fake az in test/fixtures/azure/. Both suites run without network.

License

MIT

View this README on GitHub

推奨ツール

別のキーワードを試すか、フィルタを外してください。

インストール

npx skillfish add amelnagdy/review-skills