XT

xuanhuaz1022/traeofflinerag

开发工具
47 stars 质量 40 趋势 40

A powerful document analysis and query system based on Retrieval-Augmented Generation (RAG) technology, integrated as a Trae SKILL.

概览

A powerful document analysis and query system based on Retrieval-Augmented Generation (RAG) technology, integrated as a Trae SKILL. - 📚 : TXT, Markdown, and PDF documents - 🔍 : Advanced semantic search with hybrid retrieval strategy - 🧠 : Intelligent context building for better LLM responses - 🚀 : Works seamlessly with Trae Agent - ⚡ : Built-in caching mechanism - 🔧 : Configurable chunking strategies and retrieval parameters - OPENAI_API_KEY: For OpenAI embedding model (optional) - Python 3.9+ - numpy - chromadb - sentence-transformers - PyPDF2 - tiktoken - markdown 1. Fork the repository 2. Create a feature branch (git checkout -b feature/amazing-feature) 3. Commit your changes (git commit -m 'Add some amazing feature') 4. Push to the branch (git push origin feature/amazing-feature) 5. Open a Pull Request This project is licensed under the MIT License - see the LICENSE file for details. For support, please open an issue on GitHub or contact the author.

README

RAG Document Assistant

A powerful document analysis and query system based on Retrieval-Augmented Generation (RAG) technology, integrated as a Trae SKILL.

Features

  • 📚 Multi-format support: TXT, Markdown, and PDF documents
  • 🔍 Smart retrieval: Advanced semantic search with hybrid retrieval strategy
  • 🧠 Context-aware: Intelligent context building for better LLM responses
  • 🚀 Easy integration: Works seamlessly with Trae Agent
  • ⚡ Fast performance: Built-in caching mechanism
  • 🔧 Customizable: Configurable chunking strategies and retrieval parameters

Installation

Option 1: Install from GitHub

git clone https://github.com/yourusername/rag-document-assistant.git
cd rag-document-assistant
pip install -e .

Option 2: Install from PyPI (coming soon)

pip install rag-document-assistant

Quick Start

In Trae Agent

  1. Load the SKILL:

    from skill_rag import load_document, query_document, get_status, clear_database
    
  2. Load a document:

    result = load_document("/path/to/your/document.pdf")
    print(f"✅ Loaded: {result['success']}")
    
  3. Query the document:

    result = query_document("What is this document about?", top_k=5)
    if result['success']:
        print(f"📝 Answer: {result['data']['context']}")
    
  4. Check status:

    status = get_status()
    print(f"📊 Vector count: {status['data']['vector_count']}")
    

Using Slash Commands

from slash_commands import process_slash_command

# Load document
print(process_slash_command('/load --file /path/to/document.pdf'))

# Query document
print(process_slash_command('/query --question "Your question?"'))

# Get status
print(process_slash_command('/status'))

# Clear database
print(process_slash_command('/clear'))

Configuration

Configuration File (config.json)

{
    "embedding_model": "sentence-transformers",
    "vector_db_type": "chromadb",
    "chunk_size": 512,
    "chunk_overlap": 50,
    "top_k": 5
}

Environment Variables

  • OPENAI_API_KEY: For OpenAI embedding model (optional)

Project Structure

RAG/
├── rag/
│   ├── core/              # Core RAG components
│   │   ├── loader.py      # Document loaders
│   │   ├── chunker.py     # Document chunkers
│   │   ├── embedder.py    # Embedding models
│   │   ├── retriever.py   # Retrieval strategies
│   │   └── context_builder.py  # Context building
│   ├── db/                # Database implementations
│   └── scripts/           # CLI scripts
├── trae/                 # Trae integration
│   ├── handlers.py       # Command handlers
│   └── plugin.py         # RAG plugin
├── .trae/
│   └── skills/
│       └── rag_document_assistant/
│           └── SKILL.md   # SKILL definition
├── skill_rag.py          # SKILL implementation
├── slash_commands.py     # Slash command handler
├── requirements.txt      # Dependencies
└── setup.py              # Package configuration

Dependencies

  • Python 3.9+
  • numpy
  • chromadb
  • sentence-transformers
  • PyPDF2
  • tiktoken
  • markdown

Usage Examples

Basic Document Analysis

from skill_rag import load_document, query_document

# Load document
load_result = load_document("report.pdf")
if load_result["success"]:
    print(f"✅ Loaded {load_result['data']['chunks_processed']} chunks")
    
    # Query document
    query_result = query_document("What are the key findings?")
    if query_result["success"]:
        print(f"📝 Context: {query_result['data']['context']}")

Advanced Usage

from skill_rag import execute

# Using execute function
result = execute("query_document", {
    "question": "Summarize the main points",
    "top_k": 3
})

Contributing

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add some amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

License

This project is licensed under the MIT License - see the LICENSE file for details.

Support

For support, please open an issue on GitHub or contact the author.

View this README on GitHub

推荐工具

换一个关键词,或者移除筛选条件。

安装

npx skillfish add xuanhuaz1022/traeofflinerag