PA

ptmrio/autorename-pdf

Developer tools
116 stars 품질 40 트렌드 40

AutoRename-PDF Automatically rename PDF files using AI and OCR. Extracts company names, dates, and document types from any PDF — renames to YYYYMMDD COMPANY DOCTYPE.pdf .

개요

AutoRename-PDF Automatically rename PDF files using AI and OCR. Extracts company names, dates, and document types from any PDF — renames to YYYYMMDD COMPANY DOCTYPE.pdf .

README

  • Desktop GUI — Drag-and-drop files, preview renames before applying, undo if needed
  • 5 AI Providers — OpenAI, Anthropic (Claude), Google Gemini, xAI (Grok), and Ollama for free offline use
  • 3-Tier Extraction — Text extraction (free, instant) + OCR for scanned PDFs + vision for image-only PDFs
  • Windows Integration — Right-click context menu in Windows Explorer for instant renaming
  • Batch Processing — Rename hundreds of PDFs at once with automatic company name harmonization
  • Privacy Option — Run fully offline with Ollama + PaddleOCR — no data leaves your machine

Quick Start

  1. Download the latest release ZIP
  2. Extract and run setup.ps1 (right-click → “Run with PowerShell”)
  3. Configure — edit config.yaml with your AI provider and API key:
    ai:
      provider: "openai"       # or anthropic, gemini, xai, ollama
      api_key: "your-key"
    
  4. Launch autorename-pdf-gui.exe — or right-click any PDF in Explorer

setup.ps1 creates config.yaml from the template, adds context menu entries, and optionally installs PaddleOCR for offline OCR of scanned documents.

Common Questions

Does it work offline? Yes. Use Ollama (free, local AI) + PaddleOCR. No data leaves your machine, no API key needed. See Ollama Setup.

What types of PDFs does it handle? Any PDF: text-based (digital), scanned (OCR), or image-only (vision). All three extraction methods can run together.

How much does it cost? The tool is free and open source. Cloud AI providers charge ~$0.001 per PDF. Ollama is completely free.

Does it work on macOS or Linux? The CLI works cross-platform via Python. The GUI and context menu are Windows-only. See macOS / Linux.

Configuration

API Key Setup

There are two ways to configure your API key:

Method 1 — Directly in config.yaml (simple, portable):

ai:
  api_key: "sk-your-actual-key-here"

Method 2 — Via environment variable (secure, easy to switch providers):

ai:
  api_key: "${OPENAI_API_KEY}"

Then create a .env file next to your config.yaml:

OPENAI_API_KEY=sk-your-actual-key-here
ANTHROPIC_API_KEY=sk-ant-your-key-here

The ${VAR_NAME} syntax works in any string value in config.yaml. A .env file placed next to config.yaml is loaded automatically.

Tip: Method 2 lets you store all your API keys in .env and switch providers in config.yaml by just changing the ${VAR_NAME} reference. See .env.example for a template.

Text extraction (pdfplumber) always runs — it’s free and instant. OCR and vision are independent add-ons you enable based on your needs.

Cloud AI + PaddleOCR (Best Accuracy)

PaddleOCR runs locally for free, cloud AI handles smart extraction. Great for mixed document types.

ai:
  provider: "openai"           # or anthropic, gemini, xai
  model: "gpt-5.4"
  api_key: "your-api-key"
pdf:
  ocr: true                    # PaddleOCR enhances scanned docs
  vision: false                # not needed — OCR covers it

Cost: ~$0.001/PDF. Requires: API key + PaddleOCR (~500 MB, installed via setup.ps1).

Cloud AI + Vision (No Local Setup)

No PaddleOCR needed — the LLM reads page images directly. Best for laptops or low-performance machines.

ai:
  provider: "gemini"           # or openai, anthropic, xai
  model: "gemini-3.1-flash-lite"
  api_key: "your-api-key"
pdf:
  ocr: false
  vision: true                 # send page images to LLM

Cost: ~$0.002/PDF. Requires: API key + vision-capable model.

Fully Offline (Max Privacy)

Everything runs on your machine. No data leaves your computer, no API keys, no cost.

ai:
  provider: "ollama"
  model: "qwen3:4b"            # fast, fits in 3 GB VRAM
  api_key: ""
pdf:
  ocr: true                    # PaddleOCR for scanned docs
  vision: false                # text models are faster

Cost: Free. Requires: Ollama + PaddleOCR.

Provider Models

Provider Flagship Model Budget Model
OpenAI gpt-5.4 gpt-5-mini
Anthropic claude-sonnet-4-6 claude-haiku-4-5-20251001
Gemini gemini-3.1-flash-lite gemini-3-flash-preview
xAI grok-4.20-beta-0309-non-reasoning —
Ollama qwen3:8b qwen3:4b / llama3.2:3b

See config.yaml.example for full documentation of all settings.

Extraction Settings

Setting Values Description
pdf.ocr false / true / "auto" PaddleOCR for scanned PDFs
pdf.vision false / true / "auto" Send page images to LLM
pdf.text_quality_threshold 0.0 – 1.0 Triggers OCR/vision in "auto" mode (default: 0.3)
pdf.max_pages integer Max pages to process per PDF (default: 3)
  • false = disabled (default for both)
  • true = always run alongside text extraction
  • "auto" = run only when text quality falls below threshold

All enabled sources are combined before sending to the AI — maximizing extraction accuracy.

Usage

GUI

Launch autorename-pdf-gui.exe to open the desktop interface.

  • Drag and drop PDF files or folders onto the window
  • Dry-run preview shows proposed renames before applying
  • Undo reverses the last rename operation
  • Supports light and dark themes

Context Menu

After running setup.ps1, right-click in Windows Explorer:

  • Single PDF: Right-click a PDF → Auto Rename PDF
  • Folder of PDFs: Right-click a folder → Auto Rename PDFs in Folder
  • Current Folder: Right-click folder background → Auto Rename PDFs in This Folder

Windows 11 Note: Context menu entries appear under “Show more options” (Shift+F10).

Command Line

# Rename a single PDF
autorename-pdf-cli.exe "C:\path\to\file.pdf"

# Preview what would be renamed (no changes made)
autorename-pdf-cli.exe --dry-run "C:\path\to\folder"

# Process folders recursively
autorename-pdf-cli.exe --recursive "C:\path\to\folder"

# Undo the last rename operation
autorename-pdf-cli.exe undo

# Override AI provider/model for one run
autorename-pdf-cli.exe --provider anthropic --model claude-sonnet-4-6 "file.pdf"

# Enable vision and/or OCR
autorename-pdf-cli.exe --vision --ocr "scanned_document.pdf"

# JSON output (for scripting / GUI integration)
autorename-pdf-cli.exe rename --output json "C:\path\to\folder"

Company Name Harmonization

Standardize company name variations using harmonized-company-names.yaml:

ACME:
    - "ACME Corp"
    - "ACME Inc."
    - "ACME Corporation"

XYZ:
    - "XYZ Ltd"
    - "XYZ LLC"
    - "XYZ Enterprises"

The tool uses fuzzy matching (Jaro-Winkler similarity) to automatically map extracted names to their standardized form, even with OCR typos.

Quick Tip: Copy your PDF filenames, paste them into ChatGPT/Claude/Gemini with “Create a harmonized-company-names.yaml mapping these company name variations to standardized names” — then save the result.

Support the Project

If AutoRename-PDF saves you time, consider supporting its development:

Also check out PhraseVault — a text expander and snippet manager by the same developer.

Thank You to Our Supporters

  • @claus82 — Thank you for your generous donation!

MIT License — Made by Gerhard Petermeir, SPQRK Web Solutions

View this README on GitHub

추천 도구

다른 키워드를 입력하거나 필터를 제거해 보세요.

설치

npx skillfish add ptmrio/autorename-pdf