XT

xuanhuaz1022/traeofflinerag

Developer tools
47 stars 품질 40 트렌드 40

A powerful document analysis and query system based on Retrieval-Augmented Generation (RAG) technology, integrated as a Trae SKILL.

개요

A powerful document analysis and query system based on Retrieval-Augmented Generation (RAG) technology, integrated as a Trae SKILL. - 📚 : TXT, Markdown, and PDF documents - 🔍 : Advanced semantic search with hybrid retrieval strategy - 🧠 : Intelligent context building for better LLM responses - 🚀 : Works seamlessly with Trae Agent - ⚡ : Built-in caching mechanism - 🔧 : Configurable chunking strategies and retrieval parameters - OPENAI_API_KEY: For OpenAI embedding model (optional) - Python 3.9+ - numpy - chromadb - sentence-transformers - PyPDF2 - tiktoken - markdown 1. Fork the repository 2. Create a feature branch (git checkout -b feature/amazing-feature) 3. Commit your changes (git commit -m 'Add some amazing feature') 4. Push to the branch (git push origin feature/amazing-feature) 5. Open a Pull Request This project is licensed under the MIT License - see the LICENSE file for details. For support, please open an issue on GitHub or contact the author.

README

RAG Document Assistant

A powerful document analysis and query system based on Retrieval-Augmented Generation (RAG) technology, integrated as a Trae SKILL.

Features

  • 📚 Multi-format support: TXT, Markdown, and PDF documents
  • 🔍 Smart retrieval: Advanced semantic search with hybrid retrieval strategy
  • 🧠 Context-aware: Intelligent context building for better LLM responses
  • 🚀 Easy integration: Works seamlessly with Trae Agent
  • ⚡ Fast performance: Built-in caching mechanism
  • 🔧 Customizable: Configurable chunking strategies and retrieval parameters

Installation

Option 1: Install from GitHub

git clone https://github.com/yourusername/rag-document-assistant.git
cd rag-document-assistant
pip install -e .

Option 2: Install from PyPI (coming soon)

pip install rag-document-assistant

Quick Start

In Trae Agent

  1. Load the SKILL:

    from skill_rag import load_document, query_document, get_status, clear_database
    
  2. Load a document:

    result = load_document("/path/to/your/document.pdf")
    print(f"✅ Loaded: {result['success']}")
    
  3. Query the document:

    result = query_document("What is this document about?", top_k=5)
    if result['success']:
        print(f"📝 Answer: {result['data']['context']}")
    
  4. Check status:

    status = get_status()
    print(f"📊 Vector count: {status['data']['vector_count']}")
    

Using Slash Commands

from slash_commands import process_slash_command

# Load document
print(process_slash_command('/load --file /path/to/document.pdf'))

# Query document
print(process_slash_command('/query --question "Your question?"'))

# Get status
print(process_slash_command('/status'))

# Clear database
print(process_slash_command('/clear'))

Configuration

Configuration File (config.json)

{
    "embedding_model": "sentence-transformers",
    "vector_db_type": "chromadb",
    "chunk_size": 512,
    "chunk_overlap": 50,
    "top_k": 5
}

Environment Variables

  • OPENAI_API_KEY: For OpenAI embedding model (optional)

Project Structure

RAG/
├── rag/
│   ├── core/              # Core RAG components
│   │   ├── loader.py      # Document loaders
│   │   ├── chunker.py     # Document chunkers
│   │   ├── embedder.py    # Embedding models
│   │   ├── retriever.py   # Retrieval strategies
│   │   └── context_builder.py  # Context building
│   ├── db/                # Database implementations
│   └── scripts/           # CLI scripts
├── trae/                 # Trae integration
│   ├── handlers.py       # Command handlers
│   └── plugin.py         # RAG plugin
├── .trae/
│   └── skills/
│       └── rag_document_assistant/
│           └── SKILL.md   # SKILL definition
├── skill_rag.py          # SKILL implementation
├── slash_commands.py     # Slash command handler
├── requirements.txt      # Dependencies
└── setup.py              # Package configuration

Dependencies

  • Python 3.9+
  • numpy
  • chromadb
  • sentence-transformers
  • PyPDF2
  • tiktoken
  • markdown

Usage Examples

Basic Document Analysis

from skill_rag import load_document, query_document

# Load document
load_result = load_document("report.pdf")
if load_result["success"]:
    print(f"✅ Loaded {load_result['data']['chunks_processed']} chunks")
    
    # Query document
    query_result = query_document("What are the key findings?")
    if query_result["success"]:
        print(f"📝 Context: {query_result['data']['context']}")

Advanced Usage

from skill_rag import execute

# Using execute function
result = execute("query_document", {
    "question": "Summarize the main points",
    "top_k": 3
})

Contributing

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add some amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

License

This project is licensed under the MIT License - see the LICENSE file for details.

Support

For support, please open an issue on GitHub or contact the author.

View this README on GitHub

추천 도구

다른 키워드를 입력하거나 필터를 제거해 보세요.

설치

npx skillfish add xuanhuaz1022/traeofflinerag