XT

xuanhuaz1022/traeofflinerag

Developer tools
47 stars Quality 40 Trend 40

A powerful document analysis and query system based on Retrieval-Augmented Generation (RAG) technology, integrated as a Trae SKILL.

Overview

A powerful document analysis and query system based on Retrieval-Augmented Generation (RAG) technology, integrated as a Trae SKILL. - 📚 : TXT, Markdown, and PDF documents - 🔍 : Advanced semantic search with hybrid retrieval strategy - 🧠 : Intelligent context building for better LLM responses - 🚀 : Works seamlessly with Trae Agent - ⚡ : Built-in caching mechanism - 🔧 : Configurable chunking strategies and retrieval parameters - OPENAI_API_KEY: For OpenAI embedding model (optional) - Python 3.9+ - numpy - chromadb - sentence-transformers - PyPDF2 - tiktoken - markdown 1. Fork the repository 2. Create a feature branch (git checkout -b feature/amazing-feature) 3. Commit your changes (git commit -m 'Add some amazing feature') 4. Push to the branch (git push origin feature/amazing-feature) 5. Open a Pull Request This project is licensed under the MIT License - see the LICENSE file for details. For support, please open an issue on GitHub or contact the author.

README

RAG Document Assistant

A powerful document analysis and query system based on Retrieval-Augmented Generation (RAG) technology, integrated as a Trae SKILL.

Features

  • 📚 Multi-format support: TXT, Markdown, and PDF documents
  • 🔍 Smart retrieval: Advanced semantic search with hybrid retrieval strategy
  • 🧠 Context-aware: Intelligent context building for better LLM responses
  • 🚀 Easy integration: Works seamlessly with Trae Agent
  • ⚡ Fast performance: Built-in caching mechanism
  • 🔧 Customizable: Configurable chunking strategies and retrieval parameters

Installation

Option 1: Install from GitHub

git clone https://github.com/yourusername/rag-document-assistant.git
cd rag-document-assistant
pip install -e .

Option 2: Install from PyPI (coming soon)

pip install rag-document-assistant

Quick Start

In Trae Agent

  1. Load the SKILL:

    from skill_rag import load_document, query_document, get_status, clear_database
    
  2. Load a document:

    result = load_document("/path/to/your/document.pdf")
    print(f"✅ Loaded: {result['success']}")
    
  3. Query the document:

    result = query_document("What is this document about?", top_k=5)
    if result['success']:
        print(f"📝 Answer: {result['data']['context']}")
    
  4. Check status:

    status = get_status()
    print(f"📊 Vector count: {status['data']['vector_count']}")
    

Using Slash Commands

from slash_commands import process_slash_command

# Load document
print(process_slash_command('/load --file /path/to/document.pdf'))

# Query document
print(process_slash_command('/query --question "Your question?"'))

# Get status
print(process_slash_command('/status'))

# Clear database
print(process_slash_command('/clear'))

Configuration

Configuration File (config.json)

{
    "embedding_model": "sentence-transformers",
    "vector_db_type": "chromadb",
    "chunk_size": 512,
    "chunk_overlap": 50,
    "top_k": 5
}

Environment Variables

  • OPENAI_API_KEY: For OpenAI embedding model (optional)

Project Structure

RAG/
├── rag/
│   ├── core/              # Core RAG components
│   │   ├── loader.py      # Document loaders
│   │   ├── chunker.py     # Document chunkers
│   │   ├── embedder.py    # Embedding models
│   │   ├── retriever.py   # Retrieval strategies
│   │   └── context_builder.py  # Context building
│   ├── db/                # Database implementations
│   └── scripts/           # CLI scripts
├── trae/                 # Trae integration
│   ├── handlers.py       # Command handlers
│   └── plugin.py         # RAG plugin
├── .trae/
│   └── skills/
│       └── rag_document_assistant/
│           └── SKILL.md   # SKILL definition
├── skill_rag.py          # SKILL implementation
├── slash_commands.py     # Slash command handler
├── requirements.txt      # Dependencies
└── setup.py              # Package configuration

Dependencies

  • Python 3.9+
  • numpy
  • chromadb
  • sentence-transformers
  • PyPDF2
  • tiktoken
  • markdown

Usage Examples

Basic Document Analysis

from skill_rag import load_document, query_document

# Load document
load_result = load_document("report.pdf")
if load_result["success"]:
    print(f"✅ Loaded {load_result['data']['chunks_processed']} chunks")
    
    # Query document
    query_result = query_document("What are the key findings?")
    if query_result["success"]:
        print(f"📝 Context: {query_result['data']['context']}")

Advanced Usage

from skill_rag import execute

# Using execute function
result = execute("query_document", {
    "question": "Summarize the main points",
    "top_k": 3
})

Contributing

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add some amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

License

This project is licensed under the MIT License - see the LICENSE file for details.

Support

For support, please open an issue on GitHub or contact the author.

View this README on GitHub

Recommended Tools

Try a different keyword or remove a filter.

Install

npx skillfish add xuanhuaz1022/traeofflinerag