Local-first, evidence-controlled academic writing workflows for AI agents, with bounded revision, clean-room review, and release governance.
개요
Academic Writing Toolkit (AWT) is an open-source, local-first system for evidence-controlled academic work. It gives AI agents repeatable skills, inspectable files, and deterministic checks for reading, literature review, argument design, bounded revision, citation auditing, clean-room review, and release governance. , and the last state of AWT as a DeepSeek Harness (dsh) distribution: the skills as they stood at that tag, plus a guard plugin with typed denials. That distribution is retired from main as of 2026-09-20; the decision record carries the evidence. In short: its enforcement claims were CI-proven (E0), its one paired-source pilot (E1) was negative, and the author's own chapter cycle (E2) was never run. Everything it contained is at the tag. What main ships now is the skill catalogue, the audits and the workspace scaffold, run inside the author's own agent host. Enforcement is that host's business.
README
Academic Writing Toolkit
Academic Writing Toolkit (AWT) is an open-source, local-first system for evidence-controlled academic work. It gives AI agents repeatable skills, inspectable files, and deterministic checks for reading, literature review, argument design, bounded revision, citation auditing, clean-room review, and release governance.
The core promise is simple: agents may help operate the workflow; the author keeps control of claims, boundaries, approvals, and the exact artifact that ships.
Latest tag: v0.6.0-rc.2, a pre-release, and the last state of AWT as a DeepSeek Harness (dsh) distribution: the skills as they stood at that tag, plus a guard plugin with typed denials. That distribution is retired from
mainas of 2026-09-20; the decision record carries the evidence. In short: its enforcement claims were CI-proven (E0), its one paired-source pilot (E1) was negative, and the author’s own chapter cycle (E2) was never run. Everything it contained is at the tag.What
mainships now is the skill catalogue, the audits and the workspace scaffold, run inside the author’s own agent host. Enforcement is that host’s business.v0.5.0 is the last release of the previous product — the Workbench wheel, Codex plugin package and ChatGPT App, all decommissioned since. It is still the place to get those, and nothing on
mainreplaces them.
AWT is not a hosted writing service and does not operate a manuscript-storage
backend. Its deterministic tools stay local and need no model key. Online
reference metadata checks run only when you explicitly add --online.
Why AWT exists
Fluent text hides three failures that a manuscript pays for at review. A source is remembered rather than read, and the sentence says more than the source does. A contribution is named as if nobody had named it before. And the prose carries the shape of the model that drafted it, in ways the author has stopped seeing. AWT is built around those three, in that order: the first two are the kind of error reviewers do not forgive, the third the kind they only notice.
| Question | What AWT gives the author |
|---|---|
| Does each cited claim stay inside what the source says? | one notes file per source with page anchors and an evidence-status field; a citation-fidelity audit that reads those notes; a claim ledger binding each claim to an archived snippet; offline BibTeX checks |
| Is the contribution placed against what already exists? | a claim-positioning audit on the manuscript’s own terms, and a prose baseline built from what the target venue actually published |
| Does the prose read as the author’s? | a fingerprint measured against the project’s own literature, by distribution rather than by count, with the numbers a manuscript prints bound to the artifacts they came from |
| Did a check that found nothing actually look? | every check exits non-zero on an empty target and reports what it read; a registry keeps that true for every new check |
The judgements stay with the author. A machine reading is labelled as the machine’s, and an approval an agent types into a file it can edit is a record, not a gate.
How it fits into writing
- Read and note.
/readand/noteproduce the notes files; the notes lint keeps their contract, so a source read only in abstract cannot be cited as if it had been read in full. - Map and integrate.
/mapshows which sources reach which chapters and the word counts against targets;/integrateweaves notes into chapters with attribution and refuses sources whose evidence status is partial. - Audit before you send.
/auditruns the citation-fidelity, claim ledger, number ledger, claim-positioning and prose-fingerprint checks;/verify-refschecks the bibliography offline, and online only when asked. - Review, then export.
/reviewwrites findings asfile:linerows the author can check against the source they name;/exportproduces the Word and ZIP files.
A writing loop that records, turn by turn, what the author asked for and what
the agent read it as is being built on the feat/writing-loop branch. It is
not on the release path and nothing on main depends on it.
Use the skills in Codex
The canonical skill tree is exposed at .agents/skills/ (the same files
Claude Code reads from .claude/skills/) — clone the repo, run
node scripts/setup.mjs, and point Codex at it, or let
node scaffold/awt.mjs init link the skills into a fresh thesis
workspace. On Windows, setup keeps
Git’s flattened link files intact and adds ignored awt-local-* directory
junctions to the same canonical skills.
To use the nine skills across your local Codex projects, install them in user scope from a source checkout (Python 3.9+, Node.js ^22.12 or >=24):
git clone https://github.com/yha9806/academic-writing-toolkit.git
cd academic-writing-toolkit
python scripts/install-codex-skills.py --install-deps
Use python3 if that is your Python command. This copies self-contained
skills to ~/.agents/skills with their helpers, creates a private Python
environment, and verifies the installed files. --install-deps downloads the
declared Python dependencies on first use; no model key is needed. Repeat the
last command after git pull --ff-only to update.
Existing skills from another installer, or locally edited skills, require
an explicit --replace-existing; their complete folders are backed up
before replacement. Use --dest to update an existing legacy Codex skills
directory. See global installation, verification and recovery.
The former packaged Codex plugin (plugins/)
was decommissioned with the v0.1 rebuild; it remains installable from the
immutable v0.5.0 tag.
Use the full repository from source
Use this route when you want the complete skill sources, examples, validators, and project templates, or when you plan to contribute to AWT. Most authors should start with the Codex install above or a plain clone.
Use git clone, not GitHub’s Download ZIP. AWT uses symlinks under .agents/skills/ so compatible local agents discover the same canonical skills.
The primary surface is an agent-native local agent skill workflow: the agent operates explicit files and validators inside the project you opened.
git clone https://github.com/yha9806/academic-writing-toolkit.git my-writing-project
cd my-writing-project
make setup
make doctor
Open the folder in your agent runtime and ask:
Show me the available academic-writing skills, explain which files each one reads or writes, and recommend the smallest safe workflow for my task.
Local discovery paths:
| Runtime | Discovery path | Setup guide |
|---|---|---|
| Claude Code | .claude/skills/ |
Claude Code |
| Codex | .agents/skills/ |
Codex CLI |
| Gemini CLI | .agents/skills/ |
Gemini CLI |
| Cursor | .cursor/rules/ baseline |
Cursor |
Run the 10-minute demo
The demo uses fictional public-safe sources and the same validators real projects use. It runs against local fixtures and needs no network.
python3 .claude/skills/verify-refs/scripts/verify-refs.py \
--bib examples/demo-project/references.bib --json
node .claude/skills/note/scripts/notes-lint.mjs examples/demo-project/literature/reading_notes/smith2024_NOTES.md
A valid run reports no blocking issues. The earlier governance-packet demos
and the lost-in-conversation comparison fixture were retired with their
skills; they remain inspectable under archive/skills/
and examples/ but are no longer presented as evaluations.
8 composable skills
The catalogue was triaged from 20 skills to 9 plus 3 reference documents on
2026-08-16 after an adversarial efficacy review (every skill had to beat the
unaided frontier model to stay). See
docs/specs/2026-08-16-awt-dsh-app-v0.1-design.md
for the per-skill verdicts; retired skills live under archive/skills/.
| Lane | Skills | What the lane produces |
|---|---|---|
| Read and ground | /read, /note, /map |
page-anchored notes with an evidence-status firewall, coverage matrix, progress dashboard |
| Write without losing the sources | /integrate |
notes woven into chapters with attribution; sources read only in part are refused as support |
| Review and ship | /review, /readers, /audit, /verify-refs, /export |
file:line review findings, what a reader panel carried away against the author’s intended points, the five audits, BibTeX checks, Word/ZIP exports |
The /review instructions distinguish external
review of another author’s submitted work from own-work review of the user’s
draft. Own-work clean-room review calls for a fresh-context subagent given only
the manuscript and explicitly listed evidence files. If no subagent is
available, the output must be labelled as not clean-room.
Reference documents (loaded on demand, no standing prompt cost):
references/argument-checklist.md,
references/evidence-vocabulary.md,
references/reframe-method.md.
Detailed, goal-oriented documentation lives in:
- Skill guides
- Use-case guides
- Write a literature review
- Audit thesis citations
- Verify references before submission
- Prepare a release-governance packet
What the checks guarantee — and what they do not
AWT’s deterministic helpers verify structural facts that software can check reliably:
- required fields, sections and allowed status values in a notes file
- source-note citation shape and in-text citation consistency
- malformed or duplicate BibTeX records
- that a quoted span is in the source the notes recorded, on the page they recorded
- that a claim, and a printed number, is bound to an artifact that still contains it
- the distribution of a manuscript’s prose devices against a baseline the author approved
- public-content boundaries and home-directory paths in anything this repository ships
They do not prove that a scientific claim is true, that evidence is sufficient for a venue, that a paper will be accepted, or that an AI-generated revision expresses the author’s intent. Those remain human scholarly judgments.
Safe fixers are deliberately narrow. They may normalise conservative citation punctuation or replace known US spellings with British forms; they do not invent references, rewrite arguments, or mark unresolved evidence as verified.
Deterministic quality gates
make setup # once per clone: configs, export backend, doctor
make doctor # read-only environment and project health
make test # regression suite, live surfaces (make test-all adds retired bundles)
node --test scripts/test-catalogue.mjs scripts/test-notes-lint.mjs scripts/test-citation-fidelity.mjs scripts/test-scaffold.mjs # also run by make test
node .claude/skills/note/scripts/notes-lint.mjs literature/reading_notes/*_NOTES.md
python3 scripts/audit-citations.py --base-dir . --style harvard --json
python3 scripts/audit-british-english.py --base-dir . --json
python3 scripts/audit-logic.py --base-dir . --json
python3 .claude/skills/audit/scripts/audit-prose-fingerprint.py --target chapters --baseline literature --exclude 'ourname*'
python3 .claude/skills/audit/scripts/audit-claim-positioning.py --base-dir . --json
node .claude/skills/audit/scripts/audit-citation-fidelity.mjs --base-dir . --json
python3 scripts/audit-public-content.py --base-dir .
python3 scripts/check-fails-closed.py # every check exits non-zero on an empty target
python3 scripts/session-scan.py --repo . --transcript # before staging, committing or planning a push
session-scan.py exists because two sessions of the same agent, working on
the same repositories one evening, each changed what the other should do next
without either noticing: one committed from a shared checkout and swept in the
other’s uncommitted edits; one merged and deleted a branch the other still meant
to push. The scan reads the remote as ls-remote reports it, not as the
tracking refs remember it, tells commits made in this worktree from commits
made elsewhere, flags a staged file older than the session, and lists the
other worktrees and the other sessions’ transcripts that name this repository.
It exits 2 when it cannot see a remote, because a scan that could not scan is
not a pass.
Reference verification is offline by default:
python3 .claude/skills/verify-refs/scripts/verify-refs.py --bib references.bib --json
python3 .claude/skills/verify-refs/scripts/verify-refs.py --bib references.bib --json --online
python3 .claude/skills/verify-refs/scripts/verify-refs.py --bib references.bib --json --online --metadata-dir path/to/metadata-fixtures
--exclude drops baseline files by glob. Point it at the authors’ own
papers: a baseline that contains them is partly the thing being measured, and
in practice it is often their own prior work that sets the extreme a target is
then judged against.
Two definitions in that audit are deliberate and worth knowing before the
numbers are read. A sentence may not begin with (, because in a PDF-derived
baseline that rule splits every inline author-year citation into a sentence,
and it does so in proportion to how much author-year citation each paper
happens to use. And sentence_length_lag1 only correlates spans that were
genuinely adjacent: filtering first and correlating afterwards joins the two
sentences on either side of anything dropped, which is enough to move a
manuscript from inside the published range to outside it.
The explicit --online mode can query Crossref, Semantic Scholar, and arXiv. CI uses local fixtures so the release gate stays deterministic.
Project structure
my-writing-project/
├── .claude/skills/ canonical nine-skill catalogue (single source)
├── .agents/skills/ 1:1 links — Codex and other Agent-Skills hosts read here
├── scaffold/ awt init: a clean thesis workspace linked to the catalogue
├── references/ on-demand reference documents
├── archive/skills/ retired skill bundles (history; validators run with make test-all)
├── chapters/ manuscript chapters (workspace demo)
├── literature/
│ └── reading_notes/ one structured notes file per source
├── final_output/ generated Word and ZIP outputs
├── scripts/ deterministic validators, maintenance tools, the Node tests
├── CLAUDE.md canonical project configuration
├── AGENTS.md generated agent configuration
└── GEMINI.md generated Gemini configuration
Edit CLAUDE.md for project-specific directories, page limits, British English policy, and citation style, then run make sync. Do not edit the generated AGENTS.md or GEMINI.md blocks by hand.
Release and distribution
mainafter 2026-09-20 — the dsh distribution is retired (decision record); the skills, the audits and the workspace scaffold continue here without it- v0.6.0-rc.2 — pre-release; the audit scripts fail closed and are registered as such in CI, a claim ledger and a number ledger, a venue-baseline builder, native Windows setup. Same evidence state as rc.1: Gate A §7 verified on macOS only (#56), no author chapter cycle yet (#35)
- v0.6.0-rc.1 — pre-release, the first tag of the dsh-distribution architecture; verified against Gate A §7 on macOS and not yet on Windows (#56)
- v0.5.0 stable release — the last release carrying the ChatGPT App, its privacy/terms documents, the Cloud Run/Render deployments, and the local workbench wheel; those surfaces are decommissioned on main (v0.1 design §13)
- README visual source in Figma
Every release should identify one exact Git ref, the packaged artifact and hash, its evidence state, the gate that approved it, and the owner of any remaining human decision.
This repository currently provides open-source, local software. It does not define a paid subscription, hosted processing service, support SLA, refund policy, or billing relationship. Those require a separate commercial offer and customer-facing terms before payment is accepted.
Development
make sync # regenerate AGENTS.md and GEMINI.md from CLAUDE.md
make repair # apply narrow, idempotent local repairs
make test
The canonical skill source is .claude/skills/. Read
CONTRIBUTING.md before opening a pull request: it carries
the evidence classes every claim here is stated in, the rule that installation
and first-run changes are verified on a machine that has never run this
toolkit, and the branch and review conventions. Each of those rules names the
incident that produced it.
License
MIT. See LICENSE.
추천 도구
다른 키워드를 입력하거나 필터를 제거해 보세요.
설치
npx skillfish add yha9806/academic-writing-toolkit