CD

ctxr-dev/llm-wiki-memory

Developer tools
126 stars 0 forks Качество 55 Тренд 55

Claude Code, Cursor, Codex, and every other MCP client forget everything when a session ends. LLM Wiki Memory fixes that: it captures your conversations, compiles them into durable project knowledge...

Обзор

Claude Code, Cursor, Codex, and every other MCP client forget everything when a session ends. LLM Wiki Memory fixes that: it captures your conversations, compiles them into durable project knowledge...

README

Install

Paste this one-liner into your AI coding agent (copy button on the right) — it covers both a fresh install and an update. The full procedure lives in AI-INSTALL-PROMPT.md; the agent fetches and follows it:

Set up llm-wiki-memory in this project: fetch https://raw.githubusercontent.com/ctxr-dev/llm-wiki-memory/main/AI-INSTALL-PROMPT.md and follow it EXACTLY (it covers fresh install and update; if already installed, the same file is local at @.llm-wiki-memory/src/AI-INSTALL-PROMPT.md).

Or run it yourself — macOS / Linux:

git clone https://github.com/ctxr-dev/llm-wiki-memory ./.llm-wiki-memory/src
./.llm-wiki-memory/src/bootstrap.sh                    # add --commit-memory to git-track the wiki (you commit it)
./.llm-wiki-memory/src/bootstrap.sh --schedule hourly  # optional: hourly cron / launchd

Windows (PowerShell — the native installer, same flags):

git clone https://github.com/ctxr-dev/llm-wiki-memory ./.llm-wiki-memory/src
./.llm-wiki-memory/src/bootstrap.ps1                    # add -CommitMemory to git-track the wiki (you commit it)
./.llm-wiki-memory/src/bootstrap.ps1 -Schedule hourly  # optional: hourly Task Scheduler job

Windows prerequisites:

  • Git for Windows. You already need it to git clone, and Claude Code runs the lifecycle hooks (and the git embedding-refresh hooks) through its bundled Git Bash — so capture/recall and index-warming work out of the box.
  • LLM provider: on Windows, set an API key (ANTHROPIC_API_KEY / OPENAI_API_KEY) or a base URL (e.g. a local Ollama) — these are fetch-based and fully supported. The subscription-auth CLI providers (claude / codex / cursor-agent) are not used on Windows (they are npm .cmd shims that can’t be spawned with the distillation prompt), so bootstrap won’t auto-select them there.

The bootstrap is idempotent — re-running preserves your edits to .env and your rule files.

What to expect

Here’s what you’ll actually experience session to session. The theme: your assistant carries your context forward on its own, you stay in control of what gets saved, and everything lives on your machine.

How each moment happens: Automatic = no action from you · Agent-led = your assistant does it in its normal flow · Asks first = it proposes and saves only on your explicit yes · Background = offline housekeeping while you’re away.

When you… What you get How
Open a session Your assistant opens already knowing where you left off: it’s handed a short briefing (visible in the transcript) with your recent notes, your in-progress plans and their checkbox progress (e.g. 4/12 done, in-progress), and the wiki leaves that match your current git branch. No re-explaining. Automatic
Start a real task (“implement X”, “fix the timeout”) Before working, it recalls the lessons it learned on similar past work and applies them silently — you’ll see a one-line applied lesson: when it does, so old mistakes don’t repeat. Agent-led
Say “remember this” (a fact, a decision, a convention) It’s saved as a plain-Markdown leaf in your project’s local wiki (e.g. knowledge/infra/decision/…md), versioned in git and shared with every AI tool on your machine (or with your team, if you install a shared repo wiki) — not a per-session scratchpad that vanishes. Agent-led
Correct it, or say “save that as a lesson” It proposes one lesson at a time and saves nothing until you say yes (on Claude Code you also get a one-click yes/no prompt). One approval covers one lesson. Asks first
Approve a plan The approved plan is captured as a tracked .plan.md with its checkboxes and status, so progress survives across sessions. Automatic
End or compact a session The conversation is distilled into dated notes under daily/; a later step folds those into the durable knowledge and lessons you’ll recall next time. Automatic
Enable the optional schedule Offline, it merges near-duplicate notes and archives stale ones — never a hard delete, always reversible, never interrupting you. Off by default. Background

What it actually looks like

At session start, your assistant is handed a briefing (Claude Code injects it automatically; other clients fetch the same context via the bundled rules). For example:

## Current-work context
**Branch**: `feature/DEV-129957-timeout` • **Repo**: `webhooks`
**Top wiki matches** (semantic):
- `plans/infra/observability/fix-timeout.plan.md` (0.81) — 4/12 done, in-progress
- `investigations/investigation-hermes-timeout.md` (0.69)

## 🧠  Recently — last 3 days
- **2026-07-14 15:30** — Fixed the Hermes Cassandra read timeout by bumping the socket pool… → daily-2026-07-14-153012123.md

When you correct it, it proposes the lesson and waits — saving nothing until your yes:

Want me to save this as a lesson? Title: await the async client before asserting, error_pattern: missing-await-on-async-call.

The Automatic rows are hooks in Claude Code; every other MCP client (Cursor, Codex, Claude Desktop) does the same steps by following the rules bundled at install, and gets the same MCP tools. The “asks first” consent holds on every client — the MCP server refuses an un-consented lesson save no matter the client.

Why a wiki instead of RAG

RAG memory stacks are powerful but heavy: a vector database, a container, an embedding service, ongoing ops. For small and medium projects that overhead is rarely worth it, yet you still want the agent to remember everything and improve itself across sessions.

llm-wiki-memory gives you that loop with a local hosted wiki as the substrate. Every category stays a nested tree (never a flat pile of files): non-daily categories nest by the metadata facets you search by; daily by date; an additional subject axis scatters leaves by what they’re about. Git history and validation come free, and the tree stays readable by humans. Recall runs on local embeddings — nothing leaves your machine.

Highlights

Each highlight links to its full section where one exists.

Everything lives in a local .llm-wiki-memory/ folder — no vector DB, no container, no cloud.

Every memory is a markdown leaf with full git history (one commit per operation); your project repo is never touched — unless you deliberately install a shared team wiki, whose knowledge files land in your working tree for you to commit (the engine still never runs git).

Self-improvement lessons save only with your explicit consent, one approval per lesson. → Memory write-gate

Long transcripts are chunked and distilled in pieces; a failed chunk is stashed and retried with no data loss. → Capture pipeline

An opt-in offline pass dedupes near-identical notes and refreshes stale ones — never a hard delete, always reversible. → Consolidate

Transformer embeddings rank queries on-device (default Xenova/bge-large-en-v1.5); nothing leaves your machine.

Every atom carries an apply-strength — P0 (guardrail) / P1 (strong default) / P2 (contextual). Relevance ranks first at recall; priority only breaks near-ties, so a guardrail is never crowded out.

Paste one prompt into your agent or run one script. Idempotent.

How it works

Three flows, all on-device: a write path (a session becomes durable notes), a read path (those notes come back as context), and offline upkeep (the store refines itself). Each is a separate focused diagram below; the deep dives live in docs/.

The pieces. An AI client talks to the MCP server (recall / search / save) and fires lifecycle hooks; an hourly cron does maintenance; both read and write one local, git-versioned Markdown tree, ranked by an on-device embedding index.

%%{init: {"theme":"base","flowchart":{"curve":"linear"},"themeVariables":{"lineColor":"#00B8C4","primaryColor":"#0D0D14","primaryTextColor":"#FCEE0A","primaryBorderColor":"#FCEE0A","secondaryColor":"#16161E","tertiaryColor":"#16161E","clusterBkg":"#16161E","clusterBorder":"#00B8C4","edgeLabelBackground":"#0D0D14","textColor":"#00B8C4"}}}%%
flowchart LR
    CL["AI client(Claude Code · Cursor · Codex)"]
    MCP["MCP serverrecall · search · save"]
    HK["hooks + hourly cron"]
    EM["embedding index(on-device)"]
    subgraph WIKI["wiki tree — local, git-versioned Markdown"]
      TREES["daily · knowledge · self_improvement · plans · investigations"]
    end
    CL  MCP
    CL --> HK
    MCP --> EM
    EM --> WIKI
    MCP  WIKI
    HK --> WIKI

Write path — capture then compile

A session is captured to dated daily/ notes by the flush hooks (and approved plans go straight to plans/); the hourly cron then promotes those atoms into the durable knowledge/ and self_improvement/ trees, superseding the daily source.

%%{init: {"theme":"base","flowchart":{"curve":"linear"},"themeVariables":{"lineColor":"#00B8C4","primaryColor":"#0D0D14","primaryTextColor":"#FCEE0A","primaryBorderColor":"#FCEE0A","secondaryColor":"#16161E","tertiaryColor":"#16161E","clusterBkg":"#16161E","clusterBorder":"#00B8C4","edgeLabelBackground":"#0D0D14","textColor":"#00B8C4"}}}%%
flowchart LR
    S[AI session]
    S -- "flush hooks(pre/post-compact, session-end)" --> DA[daily/]
    S -- "ExitPlanMode hook" --> PL[plans/]
    DA -- "compile(hourly cron + session-start)" --> KSI["knowledge/ + self_improvement/"]
    KSI -. supersedes daily source .-> DA

Read path — recall

Before a task (and at session start) the agent calls recall_lessons / search_memory; the embedding index ranks leaves across the trees and returns the top hits as a briefing or an applied lesson — nothing leaves the machine.

%%{init: {"theme":"base","flowchart":{"curve":"linear"},"themeVariables":{"lineColor":"#00B8C4","primaryColor":"#0D0D14","primaryTextColor":"#FCEE0A","primaryBorderColor":"#FCEE0A","secondaryColor":"#16161E","tertiaryColor":"#16161E","clusterBkg":"#16161E","clusterBorder":"#00B8C4","edgeLabelBackground":"#0D0D14","textColor":"#00B8C4"}}}%%
flowchart LR
    AG["session start /agent task"] --> RC["recall_lessons ·search_memory"]
    RC --> EM["embedding index(cosine rank)"]
    EM --> TR["knowledge · self_improvement · plans"]
    TR --> OUT["ranked hits →briefing / applied lesson"]

Offline upkeep. An opt-in hourly pass keeps the store from becoming a write-only graveyard (dedup, staleness refresh, housekeeping) and logs each attempt so failures surface next session — see Consolidate below and docs/consolidate.md. How embedding, ranking, and the vector cache work: docs/embeddings.md.

Private brain & shared team wikis

By default your memory is a private brain: one wiki in your home directory, gitignored, visible to every AI tool on your machine but to no one else. That’s the whole story for most installs.

You can also give a repo its own shared wiki — a knowledge tree you commit into the project so everyone who clones it inherits that knowledge (the engine never runs git on a shared wiki — you commit it). The two coexist: the engine discovers every .llm-wiki-memory wiki by walking up from where you’re working (your cwd and the repos in play), stacking them into a scope chain — your private brain (depth 0) plus any repo-owned wikis above it.

  • Reads fan out across the whole chain and merge, ranked so a comparably-relevant repo-local note outranks a general one from your brain — one embedding model serves the whole fan-out (no extra memory per level).
  • Writes pick one destination. A save targets your brain (private — the default choice) or a specific repo (shared). A shared write is opt-in: the agent asks first, then only stages the note in that repo’s working tree — it isn’t shared until you commit and push it. The engine never runs git on a shared wiki — not on save, recall, install, or uninstall — so your project’s git history is only ever changed by you.
  • Install shared by running the one home engine’s node ~/.llm-wiki-memory/src/scripts/mount-init.mjs "$PWD" inside the repo (no engine clone in the project — the same command sets up a fresh shared wiki or adopts one on clone). A shared wiki is auto-detected on any re-run and stays git-tracked — a bare re-run does not revert it to private, and the engine never runs git on it.

Full walkthrough — install, adoption, the scope chain, the ranking rules (confidence + locality + priority), upgrade/uninstall, and the team caveats → docs/shared-wikis.md.

Capture pipeline — chunked & recoverable

The flush worker (PreCompact / PostCompact / SessionEnd hooks) chunks oversized transcripts and runs each chunk through a provider/model chain, each under its own timeout budget. A clean “nothing durable” verdict writes no leaf at all (the breadcrumb log keeps visibility); a partial or total failure preserves the full body to a stash so cli.mjs redistill can re-attempt later with no data loss. Every leaf records audit frontmatter, so a distillation is reproducible from the file alone.

Deep dive → docs/capture.md: the chunk → map → reduce diagram, the audit frontmatter fields, the per-failure-mode behaviour, and how redistill re-attempts a stash.

Memory write-gate (read-freely, write-gated)

Self-improvement lessons are propose-then-confirm: the agent NEVER calls save_lesson (or save_to_dataset / write_memory into self_improvement) on its own. It proposes the save in chat, waits for an explicit user yes in the same turn, then calls the tool with gate.userRequested: true. The server refuses gated writes without the flag.

Three enforcement layers, defence-in-depth — any one can refuse a save:

Layer Where What it does
Instructions (probabilistic) MCP initialize + the rule files bundled at install Tells the model the rule, the exact wording to propose, and the consent contract. Reaches every MCP client — but not airtight alone, which is why the next two exist.
Claude Code hook (deterministic, Claude Code only) PreToolUse on the three gated writers (gate.claudeHookEnabled, default on) Inspects the latest user turn for a save phrase → allow; no match → ask (one-click yes/no). Per-lesson consent: one save phrase auto-allows only the FIRST gated write of a turn, so a batch can’t ride one yes.
MCP server gate (deterministic, every client) The save_lesson / save_to_dataset / write_memory handlers Refuses any call without userRequested: true, and refuses a path: that lands under self_improvement/… from a non-gated dataset:. The airtight bottom layer for hook-less clients (Cursor, Codex, generic).

Knowledge, plans, investigations, daily, and tracker-issue writes are not gated — their routing rules apply directly.

Consolidate (offline refinement)

Consolidate is the offline pass that stops memory from becoming a write-only graveyard — bug root-causes get fixed, feedback rules get reversed, and pattern-gotchas outlive the API they warned about. It is opt-in — off by default (consolidate.enabled: false). Once enabled it runs on the hourly maintenance cron (chained after compile) and at session end, over the categories the layout marks consolidate: refine. Each run does two kinds of work: cheap deterministic dedup + housekeeping, and — only when an LLM provider is reachable — a capped semantic refresh that re-reads aged leaves and returns keep / rewrite / archive, so recall keeps surfacing advice that matches today’s code. Nothing is ever hard-deleted: every removal is a reversible disableDocument status flip, and every LLM step falls back to a safe deterministic default when the provider is missing.

Eligibility is layout-declared: every category in /.layout/layout.yaml must say consolidate: refine or consolidate: none (no defaults — the orchestrator refuses to run if a category omits the field). consolidate: none categories (plans, investigations, daily) are owned by other lifecycles and never walked.

Deep dive → docs/consolidate.md: the per-leaf pipeline diagram, where local-embedding runs vs where the LLM runs, every pass and why, the staleness → refresh mechanism and its four verdicts (keep / rewrite / archive / fallback), cost controls, all tuning knobs, self-healing, and determinism.

Works with your agent

MCP client Hooks (Claude Code only) MCP tools Write-gate enforcement
Claude Code ✅ session-start / pre-compact / post-compact / session-end / exit-plan-mode / pre-tool-use ✅ instructions + hook + server (full three-layer)
Cursor ✗ ✅ instructions + server
Codex / OpenAI ✗ ✅ instructions + server
Claude Desktop ✗ ✅ instructions + server
Any MCP client ✗ ✅ instructions + server

Hook-driven auto-capture is Claude Code only; every other client gets the same MCP tools + the same discipline (and invokes cli.mjs cron-health at session start via the bundled rule to surface unresolved cron failures).

The LLM provider that extracts typed atoms during capture / compile / consolidate is set in .llm-wiki-memory/settings/.env, independent of the client:

openai-compatible covers ollama, vLLM, lm-studio, llama.cpp server, and litellm proxies — point MEMORY_LLM_BASE_URL at a local endpoint. The provider is auto-detected at install; an explicit --provider or a user-edited chain always wins. The cross-provider chain and per-provider models fallback lists live in settings.yaml (see Configuration) — provider model names live ONLY in YAML, never inlined in code.

MCP tools

Tool Purpose
recall_lessons Recall self-improvement lessons before a task (fall-back ladder drops error_pattern, then language, then task_type). Pass sections:["frontmatter"] for a compact glance view.
search_memory Cross-category embedding search with metadata pre-filtering. Each hit is annotated with its priority; relevance ranks first, priority breaks near-ties. Bodies are excerpted at the response boundary; fullContent: true for whole bodies, sections:["frontmatter"] for a glance.
save_lesson Write-gated. Persist a lesson after explicit user yes (requires gate.userRequested: true).
save_to_dataset Upsert a plan, investigation, knowledge artefact, or other category by name. Write-gated when dataset="self_improvement".
write_memory Create a memory leaf, optionally superseding an existing one. Write-gated when datasetId="self_improvement".
consolidate_memory Run the deterministic + LLM consolidation passes. System-maintenance; not write-gated.
disable_document / enable_document / delete_document Archive (reversible) or remove a leaf.
move_document Relocate a leaf within the curated (non-facet) zone, preserving content + embedding + both index.md files. Facet / topology categories relocate by metadata / compiler path instead, and are refused here.
audit_memory Surface duplicate keys, missing metadata, and cleanup candidates.
list_datasets, get_memory_config, reload_provider, reload_layout Inspect categories, config, LLM provider, and force-refresh caches.
validate_layout, validate_topology, test_path_compiler Layout + topology + placement-compiler sanity checks.

Every tool takes a required scopes (a string[]) — the directories you’re working in (your cwd plus any repos in play). It’s never optional: an empty or missing scopes is rejected before the tool runs, and the engine walks each scope up to your home wiki to resolve context. Claude Code seeds a default at session start; other clients compute it from the working directory + git (the bundled scope-seeding skill carries the procedure).

Every write names its destination — target is required. scopes says which wikis a call concerns; target says which one a write goes into: the literal "brain" for your private tree, or a resolved level’s wiki root / mount directory for a shared repo (discover them in get_memory_config’s levels). Omitting target is rejected, so the destination is always deterministic — which matters when two identical clones of one repo are in scope.

Caveat — cloud-synced workspaces. A sync daemon (Drive, Dropbox, iCloud, OneDrive) can relocate or half-replicate files mid-session. The wiki’s own git repo is the source of truth: recover with git reset --hard HEAD and run cli.mjs doctor after a suspected scramble. The bundled cloud-sync-safety rule carries the full checklist.

Configuration

Settings live in two files in ./.llm-wiki-memory/settings/:

  • .env — secrets, provider switches, deployment paths, workspace identity, test seams. Things that genuinely need shell precedence. See templates/env.example.
  • settings.yaml — every other knob, nested by concern (consolidate, flush, hook, embed, recall, compile, gc, gate, wiki, providers) plus the top-level crossCuttingAreas list. See templates/settings.yaml.

The .env file’s strict subset overrides the YAML where it overlaps (e.g. MEMORY_LLM_PROVIDER collapses the YAML chain). As of the 2026-06-03 v2 release, every MEMORY_* env var NOT on the strict allow-list is a silent no-op — application config moved into settings.yaml.

Strict-subset .env keys:

Key Default Meaning
ANTHROPIC_API_KEY / OPENAI_API_KEY (unset) Provider API keys (only for API providers).
MEMORY_LLM_PROVIDER auto claude / codex / cursor / anthropic / openai / openai-compatible / mock. When set, collapses the YAML chain to this one provider.
MEMORY_LLM_MODEL (unset) Provider-agnostic model override; prepends to the head provider’s models list.
ANTHROPIC_MODEL / OPENAI_MODEL (unset) Provider-specific model override; prepends to that provider’s models list.
MEMORY_LLM_BASE_URL (unset) OpenAI-compatible local endpoint (ollama, vLLM, lm-studio, llama.cpp, litellm).
MEMORY_LLM_TIMEOUT_MS 120000 Per-call CLI/API timeout.
MEMORY_DATA_DIR / LLM_WIKI_MEMORY_ROOT / MEMORY_EMBED_CACHE / MEMORY_EMBED_CACHE_DIR / MEMORY_SETTINGS_PATH derived Deployment + model-cache paths.
MEMORY_DEFAULT_PROJECT_MODULE / LLM_WIKI_MEMORY_PROJECT deterministic identity Workspace identity (scopes recall): the canonical git origin as org/repo, else file://, with basename(workspace) only as a last resort.
MEMORY_MCP_SERVER_NAME llm-wiki-memory MCP server name advertised at initialize.
MEMORY_LLM_MOCK_* (unset) Test seams for the mock provider.

Recall scoping is deterministic: project_module is derived from a declared project_id > the canonical git origin org/repo > file://mountDir (nested repos chain as org/repo//sub); cli.mjs migrate-identity restamps legacy leaves. A recall project_module filter matches the INNERMOST chain segment (a leaf stamped org/repo//sub matches a filter for sub or the full chain, not the outer org/repo), so clones of a sub-package still gather regardless of parent.

Highlights from settings.yaml (the knobs you’re most likely to flip — the full annotated set is in templates/settings.yaml):

Section.key Default Meaning
consolidate.enabled false Master switch for consolidation. Off by default (every path no-ops until you set true).
consolidate.intervalDays 1 Throttle for consolidate --if-due.
consolidate.llmPassesEnabled true Disable to run deterministic-only consolidation.
embed.model Xenova/bge-large-en-v1.5 Embedding model — see the model comparison below.
embed.backend transformers transformers (on-device bge) or lexical (no model download).
recall.recentActivityDays 3 SessionStart “🧠 Recently” window (days of recent notes surfaced). 0 disables.
recall.planContextMax 2 Max plans surfaced at SessionStart. 0 hides plans.
gate.selfImprovementEnabled true Operator escape hatch for the server-side write-gate.
gate.claudeHookEnabled true Enable/disable the Claude Code PreToolUse write-gate hook.
gate.perLessonConsent true One save phrase auto-allows only the first gated write of a turn (Claude Code).
wiki.autoCommit true Auto-commit every wiki change to the wiki’s own git repo.
flush.chunkTargetK 5 Target chunk count for map-reduce distillation.
flush.reduceModelPromote true Use a one-tier-stronger model for the reduce step.

Manual commands

cd .llm-wiki-memory/src

# Inspect what consolidate WOULD do (no mutations), then run it for real.
node scripts/cli.mjs consolidate --dry-run --force --json | jq
node scripts/cli.mjs consolidate --force --json | jq '.totals'

# Full cron-job (compile + consolidate + attempt log entry), and its health.
node scripts/cli.mjs cron-job
node scripts/cli.mjs cron-health | jq

# The classic ops trio.
node scripts/cli.mjs init       # materialise or repair the wiki shell
node scripts/cli.mjs validate   # skill-llm-wiki validate
node scripts/cli.mjs heal       # classify state and name the next command

# Recall / search from the terminal; resolved paths + provider.
node scripts/cli.mjs recall ""
node scripts/cli.mjs search ""
node scripts/cli.mjs where

# Recover a failed distillation (reads the stash, or the in-leaf raw fallback).
node scripts/cli.mjs redistill --leaf       # one daily leaf
node scripts/cli.mjs redistill --session      # newest stash for a session
node scripts/cli.mjs redistill --all              # every pending stash

# Schedule the hourly cron (or remove it).
./bootstrap.sh --schedule hourly   # cron on Linux, launchd on macOS, fires at :00 ('daily' = deprecated alias)
./bootstrap.sh --schedule off      # remove

On Linux the cron entry calls a generated wrapper (state/cron-daily.sh) — safe across workspaces whose paths contain single-quotes, percents, or spaces; on macOS the launchd job runs node … cli.mjs cron-job directly via discrete arguments (no wrapper).

Testing

npm test           # unit suite
npm run test:e2e   # full lifecycle against the real skill-llm-wiki CLI (LLM stubbed)

1563 tests in total (1454 unit + 109 e2e). The unit suite covers the chunker, the provider + model chain, the map-reduce distillation flow, the redistill CLI, the wiki auto-commit layer, and the entity-level self-healing pipeline; the e2e suite builds a wiki from scratch in a temp directory and asserts genesis, daily capture, lesson / knowledge / plan / investigation absorption, compile promotion + dedup, recall, tree-growth integrity, and idempotency against the real skill-llm-wiki CLI with mocked LLM responses.

Requirements

Node 20 or newer, and git. No Docker, no Python. The embedding model downloads on first recall (set embed.backend: lexical in settings.yaml to skip it entirely).

License

MIT

View this README on GitHub

Установка

This server does not publish a one-line install command.

Open the repository installation guide