Two-model debate review of PRs and MRs, plus a babysitter that works the review rounds. Skills for any coding agent.
Обзор
debate-review has a main reviewer read the PR, a second reviewer try to knock its findings down, and the main reviewer make the final call. One review with inline comments lands on the PR, posted from your own gh, glab, or az account. babysit-pr then works the rounds: verifies each finding, fixes the blockers, replies in-thread, resolves, and re-triggers the next review. Both skills run through whatever coding agent you already drive (Claude Code, Codex, Cursor, OpenCode, Grok, and others) and the model subscriptions you already pay for. - Node 18 or newer. - Git 2.31 or newer for Azure DevOps (--config-env keeps the access token out of command arguments). - gh (GitHub), glab (GitLab, including self-hosted) or az (Azure DevOps) logged in to an account that can comment on the PR. Reviews and replies post as that account. Azure DevOps needs no az extension: the script uses az rest against the REST API.
README
review-skills
Two models argue over a pull request before anything is posted. You keep the merge.
debate-review has a main reviewer read the PR, a second reviewer try to knock its findings down,
and the main reviewer make the final call. One review with inline comments lands on the PR, posted
from your own gh, glab, or az account. babysit-pr then works the rounds: verifies each finding, fixes
the blockers, replies in-thread, resolves, and re-triggers the next review. Both skills run through
whatever coding agent you already drive (Claude Code, Codex, Cursor, OpenCode, Grok, and others) and
the model subscriptions you already pay for.
npx skills add amElnagdy/review-skills
Then ask your agent:
Use $debate-review on https://github.com/owner/repo/pull/123
Use $debate-review --local on this repo before I open a PR.
Use $babysit-pr on PR 123 until it is ready to merge.
flowchart LR
P["PR or MR"] --> M["Main reviewerlane review-main"]
M -->|"findings"| D["Debate reviewerlane review-debate"]
D -->|"confirm / refute / downgrade / add"| F["Main reviewerfinal call"]
F --> R["One reviewinline P0 / P1 / P2"]
R --> B["$babysit-prverify, fix, reply, resolve"]
B -->|"push, re-run"| M
The skills
| Skill | Job | Never does |
|---|---|---|
debate-review |
Reviews a GitHub PR, GitLab MR or Azure DevOps PR with two models in sequence and posts non-approval inline comments plus a summary. Findings that survive the debate are posted as agreed; findings the second model refuted but the main reviewer kept are posted as contested, with both sides’ reasoning. | Edit code, approve, request changes, post twice for the same head sha. |
babysit-pr |
Harvests every reviewer thread on the PR (debate-review, Codex, Greptile, any bot), checks each finding against the code, fixes what is real, replies in-thread with evidence and attribution, resolves, and re-runs the review for the next round. Reports when the PR meets the merge gate. | Merge, resolve a thread it did not answer, reply as anyone other than “model on behalf of user”. |
Requirements
-
Node 18 or newer.
-
Git 2.31 or newer for Azure DevOps (
--config-envkeeps the access token out of command arguments). -
gh(GitHub),glab(GitLab, including self-hosted) oraz(Azure DevOps) logged in to an account that can comment on the PR. Reviews and replies post as that account. Azure DevOps needs noazextension: the script usesaz restagainst the REST API. -
delegate-skills, which dispatches the reviewer models through the
*-delegaterelays, read-only. -
Two delegate lanes named
review-mainandreview-debate. Create them once with$delegate-setup:Use $delegate-setup to add two lanes: review-main on claude at high effort and review-debate on codex at high effort.Pick two different implementers. The debate is only worth something when the second model does not share the first one’s blind spots. Keep these lanes for the reviewer; if you already use a
debatelane for plan debates, leave it alone. Only implementers whose relay supports--read-onlyare accepted, so the reviewers cannot touch your tree.
What a review looks like
Each finding is one inline comment, anchored to the lines it is about, rendered as a forge alert so the colour reads before the words:
| Level | Meaning | Rendered as |
|---|---|---|
| P0 | blocking, security | [!CAUTION] |
| P1 | blocking, anything else | [!WARNING] |
| P2 | non-blocking | [!NOTE] |
The review body carries a level count table, who reviewed, and a short summary. Each comment and the
body start with an HTML marker (``) that babysit-pr uses to recognise the
threads, since they are posted from your account rather than a bot’s. One review per head sha: a
re-run on the same push exits with code 3 instead of posting again. Templates and the marker format
are in comment-format.md; the JSON contract
between the three passes is in schema.md.
The reviewers are prompted for precision, not volume: blocking bugs with a concrete trigger and wrong result, spec and standards violations quoted against the rule, and little else. The second reviewer can only refute a finding when it can point at the code that makes it impossible. Findings below the confidence floor are dropped before anything is posted.
GitHub, GitLab and Azure DevOps
| GitHub | GitLab | Azure DevOps | |
|---|---|---|---|
| Review posting | gh api, PR review with inline comments |
glab api, MR discussions with diff positions |
az rest, one comment thread per finding plus a closed summary thread |
| Target URL | /pull/ |
/-/merge_requests/ |
/_git//pullrequest/, on dev.azure.com or *.visualstudio.com |
Spec source (#123) |
issue | issue | work item |
| Alert colours | yes | 17.10+ | no, alerts fall back to plain quotes |
| Thread harvest (babysit-pr) | GraphQL review threads + REST reviews | discussions + notes | not implemented yet |
| Reply and resolve (babysit-pr) | verified live | implemented, not yet verified on a live instance | not implemented yet |
| Bot author detection | reliable (Bot type) |
only when the instance exposes author.bot; debate-review threads are found by marker either way |
n/a |
Azure DevOps has no single review object, so one debate-review is N inline threads plus one closed summary
thread carrying the marker. --force and the “already reviewed” check read the same marker back off
the PR’s threads, so a re-run on an unchanged head still exits 3.
Local preview
--local reviews the files on disk (committed, uncommitted, and untracked, honoring .gitignore)
against a base branch. It never calls a forge CLI. --dry-run still needs a live PR; it only
skips the post. The two flags do not combine. Local snapshots reject non-UTF-8 Git paths instead of
silently changing their bytes.
| Invocation | Source | Forge | Post |
|---|---|---|---|
--local |
Working tree snapshot | No | No |
--dry-run |
Live PR | Yes | No |
| `` | Live PR | Yes | Yes |
Run it by hand
Your agent normally runs these for you. They are here for testing, CI, or when there is no agent in
the loop. `` is the directory containing the skill’s SKILL.md.
node "/scripts/review-pr.mjs" --local # working tree; print, no forge
node "/scripts/review-pr.mjs" --dry-run # print, do not post
node "/scripts/review-pr.mjs" # post
"/scripts/threads.sh" # harvest one round as JSON
review-pr.mjs --help lists the flags: --local to review the working tree with no forge, --main /
--debate to override the lanes for one run, --contested post|drop, --min-confidence, --timeout
(default 30 minutes per reviewer), --force to post again on the same head, --keep to leave the
temporary worktree or snapshot clone. Every run leaves its briefs, raw model output, and the three JSON
documents under ~/.cache/debate-review/__/// so a surprising review can be
traced back to the pass that produced it.
How this relates to the sibling repos
- delegate-skills is the transport.
review-pr.mjsreads your lanes and calls the matching relay; this repo has no model code of its own. - guard-skills are review lenses for code an agent just
wrote. The reviewers here cannot load them (the relays run the reviewer with skills disabled, on
purpose), and they are tuned for a different job than Clean Code checks. Where a guard fits is the
babysitter’s own fixes:
babysit-prruns the matching guard on a fix before pushing it, when one is installed.
Development
node --test test/*.test.mjs
The babysit tests run threads.sh end to end against fake gh and glab binaries in
test/fixtures/babysit/. The debate-review tests drive the Azure forge functions against a fake
az in test/fixtures/azure/. Both suites run without network.
License
Рекомендуемые инструменты
Попробуйте другой запрос или уберите фильтр.
Установка
npx skillfish add amelnagdy/review-skills