Nature-style AI peer reviewer for scientific manuscripts.
概要
Nature-style scientific review for claims, controls, validation, uncertainty, mechanism, and generalization. Architecture · Reviewers · Failure cases · Validation · Security · Changelog · Releases Your paper has bugs. RefFox tries to find them. Suspicious by default. Evidence first. Domain-aware reviewer skills that stress-test scientific claims and turn evidence failures into actionable revision paths. 7 domain reviewers · 1 polar orchestrator · 644 reasoning patterns · 48 controlled cases · 24 matched pairs · 108 blinded outputs ~~~text claim → evidence → domain gate → failure mode → major concern → revision path ~~~ the stronger the claim, the stronger and more discriminating the evidence must be. This is an independent open-source project. It is not affiliated with Nature Portfolio or Springer Nature.
README
Nature Reviewer Skills
Nature-style scientific review for claims, controls, validation, uncertainty, mechanism, and generalization.
Architecture · Reviewers · Failure cases · Validation · Security · Changelog · Releases
Your paper has bugs. RefFox tries to find them. Suspicious by default. Evidence first.
Domain-aware reviewer skills that stress-test scientific claims and turn evidence failures into actionable revision paths.
7 domain reviewers · 1 polar orchestrator · 644 reasoning patterns · 48 controlled cases · 24 matched pairs · 108 blinded outputs
claim → evidence → domain gate → failure mode → major concern → revision path
Core principle: the stronger the claim, the stronger and more discriminating the evidence must be.
This is an independent open-source project. It is not affiliated with Nature Portfolio or Springer Nature.
See it in 60 seconds
Three synthetic examples showing how evidence chains can fail across domains.
| Remote sensing | Chemistry | Engineering |
|---|---|---|
| Sensor transition can masquerade as a vegetation breakpoint. | Raw HPLC-UV area is not automatically comparable quantitative yield. | Human recovery and data-quality decisions break a “fully autonomous” claim. |
| REVIEW FOCUS: harmonization + independent validation | REVIEW FOCUS: calibrated quantification + response factors | REVIEW FOCUS: autonomy boundary + failure recovery |
| Open case → | Open case → | Open case → |
Each case contains a synthetic manuscript excerpt, a curated reference review, and the reasoning behind the concern.
Validation
| Layer | Artifact | Scale | Status |
|---|---|---|---|
| Runtime | CI + package matrix | Python 3.10–3.13 + 8 reviewer packages | PASS |
| Controlled diagnostic | CRD-v1 | 48 cases / 24 matched pairs / 8 review groups | PASS |
| Counterfactual controls | Matched negative controls | 24 repaired counterparts | PASS |
| Uncertainty | Pair-cluster bootstrap | Matched-pair resampling | PASS |
| Blinded pilot protocol | 3 domains | 18 cases / 9 pairs / 3 repeats / 2 conditions | LOCKED |
| Formal inference | Generic vs skill-assisted | 108 outputs = 54 + 54 | PASS |
| Automated annotation | Condition-blinded AI annotation | Pilot outputs | PROVISIONAL |
| Generic-vs-skill effect | Performance comparison | — | OPEN |
| Human expert gold | Independent domain annotation | — | OPEN |
| Real submissions | Prospective manuscripts | — | OPEN |
Formal pilot domains: remote sensing · chemistry · engineering
Evidence trail: CRD-v1 · Pilot harness · Preregistration · Run protocol · Evaluation protocol
Quick start
1. Clone the repository
git clone https://github.com/GeoGeekLab/nature-reviewer-skills.git
cd nature-reviewer-skills
2. Load a reviewer skill
For basic agent use, you do not need to install the Python runtime first.
Point a SKILL.md-capable agent runtime at the full skill directory you want to use, or copy that directory into the skills location used by your runtime.
Example:
nature-earth-system-reviewer-skills/
skills/
nature-remote-sensing-reviewer-skill/
SKILL.md
reviewer_db/
references/
templates/
Then ask the agent to use that reviewer:
Use the remote-sensing reviewer skill to conduct a rigorous pre-submission review.
Identify the manuscript's central claims and test spatial independence, product validity,
uncertainty, validation design, transferability, alternative explanations, and whether
the conclusions exceed the evidence.
For each major concern, anchor the criticism to manuscript evidence and give a concrete
revision path.
Runtime note: installation paths and compatibility vary across agent systems.
3. Optional: install the shared Python runtime
Install this when you want the repository’s validation, retrieval, extraction, or evaluation CLI tooling.
Core runtime:
python -m pip install -e .
Add PDF/DOCX extraction support:
python -m pip install -e ".[documents]"
Development tools:
python -m pip install -e ".[dev]"
4. Optional: validate the repository
python scripts/sync_skill_assets.py
python scripts/validate_all.py
python -m pytest
Why not use a generic reviewer prompt?
| Generic review prompt | Nature Reviewer Skills |
|---|---|
| Broad criticism across many dimensions | Claim-dependent, domain-specific evidence gates |
| Often treats all concerns similarly | Distinguishes central scientific risks from local issues |
| May criticize without locating the evidence | Requires evidence anchors when manuscript content is available |
| Generic reproducibility and statistics checks | Discipline-specific failure modes such as spatial leakage, gauge representativeness, purity, benchmark fairness, or operating-envelope limits |
| Usually one undifferentiated voice | Supports complementary reviewer roles and panel deduplication |
| Can stop at criticism | Requires a revision direction for major concerns |
The intended advantage is domain-specific evidence stress testing, not simply a longer review.
Included reviewer skills
Earth-system science
| Reviewer | Typical stress tests |
|---|---|
| Remote sensing | spatial leakage, product validity, mixed pixels, QA screening, independent validation, uncertainty propagation, out-of-domain transfer |
| Atmospheric science | observation-model consistency, attribution, scale mismatch, forcing assumptions, internal variability, extremes |
| Hydrology | water balance, gauge representativeness, calibration/validation separation, equifinality, event sampling, basin transfer |
| Climate and ecology | confounding, mechanism, representativeness, driver separation, ecological scale, impact and policy overreach |
Chemistry, engineering, and materials
| Reviewer | Typical stress tests |
|---|---|
| Chemistry | identity and purity, discriminating controls, substrate scope, selectivity, mechanism, quantitative comparison |
| Engineering | requirement-design-validation coherence, benchmark fairness, robustness, failure modes, operating envelope, deployment claims |
| Materials science | phase identity, characterization, structure-property relations, benchmark comparability, durability, processability, scalability |
Polar Earth-System Review Orchestrator
The Polar Earth-System Review Orchestrator is an upper-layer coordinator, not an eighth foundational reviewer.
It routes Arctic, Antarctic, Southern Ocean, cryosphere, polar atmosphere/ocean, ecology, biogeochemistry, remote-sensing, paleoclimate, and instrumentation claims to the relevant domain reviewers; adds polar-specific evidence gates; assigns non-overlapping reviewer responsibilities; and consolidates the final review.
What a major concern should contain
A strong concern should make the scientific logic inspectable:
- Claim under review
- Evidence anchor — figure, table, section, method, result, or explicit missing evidence
- Why it matters
- Severity
- Alternative explanation or failure mode
- Actionable revision path
A concern should be major only when it affects the central contribution or the reader’s confidence in the evidence chain.
What the review stress-tests
| Review target | Core question |
|---|---|
| Central claim | What is the strongest claim, and what evidence would be required to defend it? |
| Evidence chain | Do data, preprocessing, methods, validation, interpretation, and conclusion form a defensible chain? |
| Controls and baselines | Can the comparison distinguish the proposed explanation from plausible alternatives? |
| Validation | Is testing independent, representative, and aligned with the claimed operating domain? |
| Uncertainty | Are important measurement, model, sampling, and inferential uncertainties quantified and propagated? |
| Causality and mechanism | Does the evidence support mechanism or causality, or only association and consistency? |
| Generalization | Where does the evidence end, and where does extrapolation begin? |
| Novelty and significance | Is the contribution positioned against the strongest relevant prior work? |
| Reproducibility | Are methods, parameters, data, code, and reporting sufficient for verification? |
| Revision path | What additional evidence, analysis, control, or claim narrowing would resolve the concern? |
How it works
manuscript
↓
identify central claims
↓
route claims to discipline-specific evidence gates
↓
retrieve relevant reviewer-reasoning patterns
↓
stress-test controls, validation, uncertainty, mechanism, and scope
↓
generate complementary referee-style reports
↓
deduplicate concerns and propose revision paths
The reviewer-memory layer stores abstracted, non-verbatim scientific reasoning patterns, not copied referee reports. A pattern follows the logic:
claim type → evidence risk → stress-test gate → reviewer concern → revision direction
Evaluation
CRD-v1 is a source-backed public development benchmark. Every positive case has a matched negative control so a reviewer is rewarded for detecting a specific evidence failure and penalized for continuing to raise it after the failure has been repaired or the claim has been narrowed.
The benchmark contains 48 synthetic cases in 24 matched pairs across 8 review groups. Controlled gold labels are developer-authored.
The preregistered three-domain pilot covers remote sensing, chemistry, and engineering. Its formal blinded inference run completed 108 outputs: 54 generic and 54 skill-assisted, with 3 repeated runs per case.
The evaluation framework supports essential-issue recall, target-specific negative-control specificity, balanced accuracy, concern precision, severity agreement, evidence anchors, panel duplication, matched-pair bootstrap uncertainty, and run-to-run stability.
See:
- CRD-v1 benchmark card
- Three-domain blinded pilot harness
- Pilot preregistration
- Blinded run protocol
- Annotation protocol
- Evaluation protocol
- Legacy synthetic benchmark report
- Test report
Reliability and defensive behavior
The shared runtime includes:
- typed Python models and validation;
- deterministic field-weighted BM25 retrieval;
- phrase boosts and query expansion;
- confidence reporting and diversity-aware selection;
- matrix CI and package validation;
- defensive PDF/DOCX/archive/path/size/page-count limits;
Technical details:
Command-line tools
Search a skill’s reviewer-pattern database:
nature-reviewer-search \
--root nature-chemistry-reviewer-skill \
--query "mechanism claim without discriminating controls" \
--limit 5
Validate a package:
nature-reviewer-validate nature-chemistry-reviewer-skill
Evaluate structured predictions:
nature-reviewer-evaluate \
benchmarks/cases \
benchmarks/predictions/example_predictions.jsonl
Extract manuscript text with anchors and defensive limits:
python nature-chemistry-reviewer-skill/scripts/extract_text_with_anchors.py \
manuscript.pdf \
--output extracted.jsonl
Repository layout
src/nature_reviewer_core/ shared runtime
scripts/ repository-wide validation and sync commands
tests/ shared-runtime tests
examples/ 60-second curated review demonstrations
benchmarks/ benchmark schemas and fixtures
nature-earth-system-reviewer-skills/
skills/
nature-remote-sensing-reviewer-skill/
nature-atmospheric-science-reviewer-skill/
nature-hydrology-reviewer-skill/
nature-climate-ecology-reviewer-skill/
polar-earth-system-review-orchestrator/
nature-chemistry-reviewer-skill/
nature-engineering-reviewer-skill/
nature-materials-science-reviewer-skill/
The current layout reflects the project’s history. A future migration may normalize all domain reviewers under one top-level skills/ directory; paths should not be changed casually because agent integrations may already depend on them.
Who it is for
- researchers preparing manuscripts for selective journals;
- principal investigators running internal pre-submission review;
- graduate students learning evidence-centered scientific criticism;
- research teams building manuscript quality-control workflows;
- editors and reviewers structuring an initial claim-evidence audit;
- agent developers building domain-aware scientific review systems.
Provenance and copyright
The reviewer-memory databases contain generalized, non-verbatim reasoning patterns distilled from public peer-review materials and domain evidence standards.
The repository does not redistribute raw referee reports, identifiable reviewer language, private review material, or full copyrighted manuscripts.
See the package-level provenance files and security documentation for additional details.
Contributing
Useful contributions include:
- expert-annotated benchmark cases;
- false-positive and negative-control cases;
- new disciplinary reviewer skills;
- stronger domain evidence gates;
- retrieval and deduplication improvements;
- documentation and reproducibility improvements.
Behavior-changing contributions should add or update benchmark cases.
Scientific benchmark gold labels require at least two domain experts, an adjudication record, and inter-rater agreement reporting.
See CONTRIBUTING.md.
License
Released under the MIT License.
Stress-test the evidence before reviewers stress-test the paper.
推奨ツール
別のキーワードを試すか、フィルタを外してください。
インストール
npx skillfish add geogeeklab/nature-reviewer-skills