GN

geogeeklab/nature-reviewer-skills

开发工具
45 stars 质量 40 趋势 40

Nature-style AI peer reviewer for scientific manuscripts.

概览

Nature-style scientific review for claims, controls, validation, uncertainty, mechanism, and generalization. Architecture · Reviewers · Failure cases · Validation · Security · Changelog · Releases Your paper has bugs. RefFox tries to find them. Suspicious by default. Evidence first. Domain-aware reviewer skills that stress-test scientific claims and turn evidence failures into actionable revision paths. 7 domain reviewers · 1 polar orchestrator · 644 reasoning patterns · 48 controlled cases · 24 matched pairs · 108 blinded outputs ~~~text claim → evidence → domain gate → failure mode → major concern → revision path ~~~ the stronger the claim, the stronger and more discriminating the evidence must be. This is an independent open-source project. It is not affiliated with Nature Portfolio or Springer Nature.

README

Nature Reviewer Skills

Nature-style scientific review for claims, controls, validation, uncertainty, mechanism, and generalization.

Architecture · Reviewers · Failure cases · Validation · Security · Changelog · Releases

Your paper has bugs. RefFox tries to find them. Suspicious by default. Evidence first.

Domain-aware reviewer skills that stress-test scientific claims and turn evidence failures into actionable revision paths.

7 domain reviewers · 1 polar orchestrator · 644 reasoning patterns · 48 controlled cases · 24 matched pairs · 108 blinded outputs

claim → evidence → domain gate → failure mode → major concern → revision path

Core principle: the stronger the claim, the stronger and more discriminating the evidence must be.

This is an independent open-source project. It is not affiliated with Nature Portfolio or Springer Nature.

Meet RefFox →

See it in 60 seconds

Three synthetic examples showing how evidence chains can fail across domains.

Remote sensing Chemistry Engineering
Sensor transition can masquerade as a vegetation breakpoint. Raw HPLC-UV area is not automatically comparable quantitative yield. Human recovery and data-quality decisions break a “fully autonomous” claim.
REVIEW FOCUS: harmonization + independent validation REVIEW FOCUS: calibrated quantification + response factors REVIEW FOCUS: autonomy boundary + failure recovery
Open case → Open case → Open case →

Each case contains a synthetic manuscript excerpt, a curated reference review, and the reasoning behind the concern.

Browse all examples →

Validation

Layer Artifact Scale Status
Runtime CI + package matrix Python 3.10–3.13 + 8 reviewer packages PASS
Controlled diagnostic CRD-v1 48 cases / 24 matched pairs / 8 review groups PASS
Counterfactual controls Matched negative controls 24 repaired counterparts PASS
Uncertainty Pair-cluster bootstrap Matched-pair resampling PASS
Blinded pilot protocol 3 domains 18 cases / 9 pairs / 3 repeats / 2 conditions LOCKED
Formal inference Generic vs skill-assisted 108 outputs = 54 + 54 PASS
Automated annotation Condition-blinded AI annotation Pilot outputs PROVISIONAL
Generic-vs-skill effect Performance comparison — OPEN
Human expert gold Independent domain annotation — OPEN
Real submissions Prospective manuscripts — OPEN

Formal pilot domains: remote sensing · chemistry · engineering

Evidence trail: CRD-v1 · Pilot harness · Preregistration · Run protocol · Evaluation protocol

Quick start

1. Clone the repository

git clone https://github.com/GeoGeekLab/nature-reviewer-skills.git
cd nature-reviewer-skills

2. Load a reviewer skill

For basic agent use, you do not need to install the Python runtime first.

Point a SKILL.md-capable agent runtime at the full skill directory you want to use, or copy that directory into the skills location used by your runtime.

Example:

nature-earth-system-reviewer-skills/
  skills/
    nature-remote-sensing-reviewer-skill/
      SKILL.md
      reviewer_db/
      references/
      templates/

Then ask the agent to use that reviewer:

Use the remote-sensing reviewer skill to conduct a rigorous pre-submission review.

Identify the manuscript's central claims and test spatial independence, product validity,
uncertainty, validation design, transferability, alternative explanations, and whether
the conclusions exceed the evidence.

For each major concern, anchor the criticism to manuscript evidence and give a concrete
revision path.

Runtime note: installation paths and compatibility vary across agent systems.

3. Optional: install the shared Python runtime

Install this when you want the repository’s validation, retrieval, extraction, or evaluation CLI tooling.

Core runtime:

python -m pip install -e .

Add PDF/DOCX extraction support:

python -m pip install -e ".[documents]"

Development tools:

python -m pip install -e ".[dev]"

4. Optional: validate the repository

python scripts/sync_skill_assets.py
python scripts/validate_all.py
python -m pytest

Why not use a generic reviewer prompt?

Generic review prompt Nature Reviewer Skills
Broad criticism across many dimensions Claim-dependent, domain-specific evidence gates
Often treats all concerns similarly Distinguishes central scientific risks from local issues
May criticize without locating the evidence Requires evidence anchors when manuscript content is available
Generic reproducibility and statistics checks Discipline-specific failure modes such as spatial leakage, gauge representativeness, purity, benchmark fairness, or operating-envelope limits
Usually one undifferentiated voice Supports complementary reviewer roles and panel deduplication
Can stop at criticism Requires a revision direction for major concerns

The intended advantage is domain-specific evidence stress testing, not simply a longer review.

Included reviewer skills

Earth-system science

Reviewer Typical stress tests
Remote sensing spatial leakage, product validity, mixed pixels, QA screening, independent validation, uncertainty propagation, out-of-domain transfer
Atmospheric science observation-model consistency, attribution, scale mismatch, forcing assumptions, internal variability, extremes
Hydrology water balance, gauge representativeness, calibration/validation separation, equifinality, event sampling, basin transfer
Climate and ecology confounding, mechanism, representativeness, driver separation, ecological scale, impact and policy overreach

Chemistry, engineering, and materials

Reviewer Typical stress tests
Chemistry identity and purity, discriminating controls, substrate scope, selectivity, mechanism, quantitative comparison
Engineering requirement-design-validation coherence, benchmark fairness, robustness, failure modes, operating envelope, deployment claims
Materials science phase identity, characterization, structure-property relations, benchmark comparability, durability, processability, scalability

Polar Earth-System Review Orchestrator

The Polar Earth-System Review Orchestrator is an upper-layer coordinator, not an eighth foundational reviewer.

It routes Arctic, Antarctic, Southern Ocean, cryosphere, polar atmosphere/ocean, ecology, biogeochemistry, remote-sensing, paleoclimate, and instrumentation claims to the relevant domain reviewers; adds polar-specific evidence gates; assigns non-overlapping reviewer responsibilities; and consolidates the final review.

What a major concern should contain

A strong concern should make the scientific logic inspectable:

  1. Claim under review
  2. Evidence anchor — figure, table, section, method, result, or explicit missing evidence
  3. Why it matters
  4. Severity
  5. Alternative explanation or failure mode
  6. Actionable revision path

A concern should be major only when it affects the central contribution or the reader’s confidence in the evidence chain.

What the review stress-tests

Review target Core question
Central claim What is the strongest claim, and what evidence would be required to defend it?
Evidence chain Do data, preprocessing, methods, validation, interpretation, and conclusion form a defensible chain?
Controls and baselines Can the comparison distinguish the proposed explanation from plausible alternatives?
Validation Is testing independent, representative, and aligned with the claimed operating domain?
Uncertainty Are important measurement, model, sampling, and inferential uncertainties quantified and propagated?
Causality and mechanism Does the evidence support mechanism or causality, or only association and consistency?
Generalization Where does the evidence end, and where does extrapolation begin?
Novelty and significance Is the contribution positioned against the strongest relevant prior work?
Reproducibility Are methods, parameters, data, code, and reporting sufficient for verification?
Revision path What additional evidence, analysis, control, or claim narrowing would resolve the concern?

How it works

manuscript
   ↓
identify central claims
   ↓
route claims to discipline-specific evidence gates
   ↓
retrieve relevant reviewer-reasoning patterns
   ↓
stress-test controls, validation, uncertainty, mechanism, and scope
   ↓
generate complementary referee-style reports
   ↓
deduplicate concerns and propose revision paths

The reviewer-memory layer stores abstracted, non-verbatim scientific reasoning patterns, not copied referee reports. A pattern follows the logic:

claim type → evidence risk → stress-test gate → reviewer concern → revision direction

Evaluation

CRD-v1 is a source-backed public development benchmark. Every positive case has a matched negative control so a reviewer is rewarded for detecting a specific evidence failure and penalized for continuing to raise it after the failure has been repaired or the claim has been narrowed.

The benchmark contains 48 synthetic cases in 24 matched pairs across 8 review groups. Controlled gold labels are developer-authored.

The preregistered three-domain pilot covers remote sensing, chemistry, and engineering. Its formal blinded inference run completed 108 outputs: 54 generic and 54 skill-assisted, with 3 repeated runs per case.

The evaluation framework supports essential-issue recall, target-specific negative-control specificity, balanced accuracy, concern precision, severity agreement, evidence anchors, panel duplication, matched-pair bootstrap uncertainty, and run-to-run stability.

See:

Reliability and defensive behavior

The shared runtime includes:

  • typed Python models and validation;
  • deterministic field-weighted BM25 retrieval;
  • phrase boosts and query expansion;
  • confidence reporting and diversity-aware selection;
  • matrix CI and package validation;
  • defensive PDF/DOCX/archive/path/size/page-count limits;

Technical details:

Command-line tools

Search a skill’s reviewer-pattern database:

nature-reviewer-search \
  --root nature-chemistry-reviewer-skill \
  --query "mechanism claim without discriminating controls" \
  --limit 5

Validate a package:

nature-reviewer-validate nature-chemistry-reviewer-skill

Evaluate structured predictions:

nature-reviewer-evaluate \
  benchmarks/cases \
  benchmarks/predictions/example_predictions.jsonl

Extract manuscript text with anchors and defensive limits:

python nature-chemistry-reviewer-skill/scripts/extract_text_with_anchors.py \
  manuscript.pdf \
  --output extracted.jsonl

Repository layout

src/nature_reviewer_core/                 shared runtime
scripts/                                  repository-wide validation and sync commands
tests/                                    shared-runtime tests
examples/                                 60-second curated review demonstrations
benchmarks/                               benchmark schemas and fixtures

nature-earth-system-reviewer-skills/
  skills/
    nature-remote-sensing-reviewer-skill/
    nature-atmospheric-science-reviewer-skill/
    nature-hydrology-reviewer-skill/
    nature-climate-ecology-reviewer-skill/
    polar-earth-system-review-orchestrator/

nature-chemistry-reviewer-skill/
nature-engineering-reviewer-skill/
nature-materials-science-reviewer-skill/

The current layout reflects the project’s history. A future migration may normalize all domain reviewers under one top-level skills/ directory; paths should not be changed casually because agent integrations may already depend on them.

Who it is for

  • researchers preparing manuscripts for selective journals;
  • principal investigators running internal pre-submission review;
  • graduate students learning evidence-centered scientific criticism;
  • research teams building manuscript quality-control workflows;
  • editors and reviewers structuring an initial claim-evidence audit;
  • agent developers building domain-aware scientific review systems.

The reviewer-memory databases contain generalized, non-verbatim reasoning patterns distilled from public peer-review materials and domain evidence standards.

The repository does not redistribute raw referee reports, identifiable reviewer language, private review material, or full copyrighted manuscripts.

See the package-level provenance files and security documentation for additional details.

Contributing

Useful contributions include:

  • expert-annotated benchmark cases;
  • false-positive and negative-control cases;
  • new disciplinary reviewer skills;
  • stronger domain evidence gates;
  • retrieval and deduplication improvements;
  • documentation and reproducibility improvements.

Behavior-changing contributions should add or update benchmark cases.

Scientific benchmark gold labels require at least two domain experts, an adjudication record, and inter-rater agreement reporting.

See CONTRIBUTING.md.

License

Released under the MIT License.


Stress-test the evidence before reviewers stress-test the paper.

View this README on GitHub

推荐工具

换一个关键词,或者移除筛选条件。

安装

npx skillfish add geogeeklab/nature-reviewer-skills