AutoRename-PDF Automatically rename PDF files using AI and OCR. Extracts company names, dates, and document types from any PDF — renames to YYYYMMDD COMPANY DOCTYPE.pdf .
개요
AutoRename-PDF Automatically rename PDF files using AI and OCR. Extracts company names, dates, and document types from any PDF — renames to YYYYMMDD COMPANY DOCTYPE.pdf .
README
- Desktop GUI — Drag-and-drop files, preview renames before applying, undo if needed
- 5 AI Providers — OpenAI, Anthropic (Claude), Google Gemini, xAI (Grok), and Ollama for free offline use
- 3-Tier Extraction — Text extraction (free, instant) + OCR for scanned PDFs + vision for image-only PDFs
- Windows Integration — Right-click context menu in Windows Explorer for instant renaming
- Batch Processing — Rename hundreds of PDFs at once with automatic company name harmonization
- Privacy Option — Run fully offline with Ollama + PaddleOCR — no data leaves your machine
Quick Start
- Download the latest release ZIP
- Extract and run
setup.ps1(right-click → “Run with PowerShell”) - Configure — edit
config.yamlwith your AI provider and API key:ai: provider: "openai" # or anthropic, gemini, xai, ollama api_key: "your-key" - Launch
autorename-pdf-gui.exe— or right-click any PDF in Explorer
setup.ps1createsconfig.yamlfrom the template, adds context menu entries, and optionally installs PaddleOCR for offline OCR of scanned documents.
Common Questions
Does it work offline? Yes. Use Ollama (free, local AI) + PaddleOCR. No data leaves your machine, no API key needed. See Ollama Setup.
What types of PDFs does it handle? Any PDF: text-based (digital), scanned (OCR), or image-only (vision). All three extraction methods can run together.
How much does it cost? The tool is free and open source. Cloud AI providers charge ~$0.001 per PDF. Ollama is completely free.
Does it work on macOS or Linux? The CLI works cross-platform via Python. The GUI and context menu are Windows-only. See macOS / Linux.
Configuration
API Key Setup
There are two ways to configure your API key:
Method 1 — Directly in config.yaml (simple, portable):
ai:
api_key: "sk-your-actual-key-here"
Method 2 — Via environment variable (secure, easy to switch providers):
ai:
api_key: "${OPENAI_API_KEY}"
Then create a .env file next to your config.yaml:
OPENAI_API_KEY=sk-your-actual-key-here
ANTHROPIC_API_KEY=sk-ant-your-key-here
The ${VAR_NAME} syntax works in any string value in config.yaml. A .env file placed next to config.yaml is loaded automatically.
Tip: Method 2 lets you store all your API keys in
.envand switch providers inconfig.yamlby just changing the${VAR_NAME}reference. See.env.examplefor a template.
Recommended Setups
Text extraction (pdfplumber) always runs — it’s free and instant. OCR and vision are independent add-ons you enable based on your needs.
Cloud AI + PaddleOCR (Best Accuracy)
PaddleOCR runs locally for free, cloud AI handles smart extraction. Great for mixed document types.
ai:
provider: "openai" # or anthropic, gemini, xai
model: "gpt-5.4"
api_key: "your-api-key"
pdf:
ocr: true # PaddleOCR enhances scanned docs
vision: false # not needed — OCR covers it
Cost: ~$0.001/PDF. Requires: API key + PaddleOCR (~500 MB, installed via setup.ps1).
Cloud AI + Vision (No Local Setup)
No PaddleOCR needed — the LLM reads page images directly. Best for laptops or low-performance machines.
ai:
provider: "gemini" # or openai, anthropic, xai
model: "gemini-3.1-flash-lite"
api_key: "your-api-key"
pdf:
ocr: false
vision: true # send page images to LLM
Cost: ~$0.002/PDF. Requires: API key + vision-capable model.
Fully Offline (Max Privacy)
Everything runs on your machine. No data leaves your computer, no API keys, no cost.
ai:
provider: "ollama"
model: "qwen3:4b" # fast, fits in 3 GB VRAM
api_key: ""
pdf:
ocr: true # PaddleOCR for scanned docs
vision: false # text models are faster
Cost: Free. Requires: Ollama + PaddleOCR.
Provider Models
| Provider | Flagship Model | Budget Model |
|---|---|---|
| OpenAI | gpt-5.4 |
gpt-5-mini |
| Anthropic | claude-sonnet-4-6 |
claude-haiku-4-5-20251001 |
| Gemini | gemini-3.1-flash-lite |
gemini-3-flash-preview |
| xAI | grok-4.20-beta-0309-non-reasoning |
— |
| Ollama | qwen3:8b |
qwen3:4b / llama3.2:3b |
See config.yaml.example for full documentation of all settings.
Extraction Settings
| Setting | Values | Description |
|---|---|---|
pdf.ocr |
false / true / "auto" |
PaddleOCR for scanned PDFs |
pdf.vision |
false / true / "auto" |
Send page images to LLM |
pdf.text_quality_threshold |
0.0 – 1.0 |
Triggers OCR/vision in "auto" mode (default: 0.3) |
pdf.max_pages |
integer | Max pages to process per PDF (default: 3) |
false= disabled (default for both)true= always run alongside text extraction"auto"= run only when text quality falls below threshold
All enabled sources are combined before sending to the AI — maximizing extraction accuracy.
Usage
GUI
Launch autorename-pdf-gui.exe to open the desktop interface.
- Drag and drop PDF files or folders onto the window
- Dry-run preview shows proposed renames before applying
- Undo reverses the last rename operation
- Supports light and dark themes
Context Menu
After running setup.ps1, right-click in Windows Explorer:
- Single PDF: Right-click a PDF →
Auto Rename PDF - Folder of PDFs: Right-click a folder →
Auto Rename PDFs in Folder - Current Folder: Right-click folder background →
Auto Rename PDFs in This Folder
Windows 11 Note: Context menu entries appear under “Show more options” (Shift+F10).
Command Line
# Rename a single PDF
autorename-pdf-cli.exe "C:\path\to\file.pdf"
# Preview what would be renamed (no changes made)
autorename-pdf-cli.exe --dry-run "C:\path\to\folder"
# Process folders recursively
autorename-pdf-cli.exe --recursive "C:\path\to\folder"
# Undo the last rename operation
autorename-pdf-cli.exe undo
# Override AI provider/model for one run
autorename-pdf-cli.exe --provider anthropic --model claude-sonnet-4-6 "file.pdf"
# Enable vision and/or OCR
autorename-pdf-cli.exe --vision --ocr "scanned_document.pdf"
# JSON output (for scripting / GUI integration)
autorename-pdf-cli.exe rename --output json "C:\path\to\folder"
Company Name Harmonization
Standardize company name variations using harmonized-company-names.yaml:
ACME:
- "ACME Corp"
- "ACME Inc."
- "ACME Corporation"
XYZ:
- "XYZ Ltd"
- "XYZ LLC"
- "XYZ Enterprises"
The tool uses fuzzy matching (Jaro-Winkler similarity) to automatically map extracted names to their standardized form, even with OCR typos.
Quick Tip: Copy your PDF filenames, paste them into ChatGPT/Claude/Gemini with “Create a harmonized-company-names.yaml mapping these company name variations to standardized names” — then save the result.
Support the Project
If AutoRename-PDF saves you time, consider supporting its development:
- ⭐ Star this repo on GitHub
- 💖 Sponsor on GitHub
- ☕ Buy me a coffee on Ko-fi
- 💛 Donate via PayPal
Also check out PhraseVault — a text expander and snippet manager by the same developer.
Thank You to Our Supporters
- @claus82 — Thank you for your generous donation!
MIT License — Made by Gerhard Petermeir, SPQRK Web Solutions
추천 도구
다른 키워드를 입력하거나 필터를 제거해 보세요.
설치
npx skillfish add ptmrio/autorename-pdf