A powerful document analysis and query system based on Retrieval-Augmented Generation (RAG) technology, integrated as a Trae SKILL.
概览
A powerful document analysis and query system based on Retrieval-Augmented Generation (RAG) technology, integrated as a Trae SKILL. - 📚 : TXT, Markdown, and PDF documents - 🔍 : Advanced semantic search with hybrid retrieval strategy - 🧠 : Intelligent context building for better LLM responses - 🚀 : Works seamlessly with Trae Agent - ⚡ : Built-in caching mechanism - 🔧 : Configurable chunking strategies and retrieval parameters - OPENAI_API_KEY: For OpenAI embedding model (optional) - Python 3.9+ - numpy - chromadb - sentence-transformers - PyPDF2 - tiktoken - markdown 1. Fork the repository 2. Create a feature branch (git checkout -b feature/amazing-feature) 3. Commit your changes (git commit -m 'Add some amazing feature') 4. Push to the branch (git push origin feature/amazing-feature) 5. Open a Pull Request This project is licensed under the MIT License - see the LICENSE file for details. For support, please open an issue on GitHub or contact the author.
README
RAG Document Assistant
A powerful document analysis and query system based on Retrieval-Augmented Generation (RAG) technology, integrated as a Trae SKILL.
Features
- 📚 Multi-format support: TXT, Markdown, and PDF documents
- 🔍 Smart retrieval: Advanced semantic search with hybrid retrieval strategy
- 🧠 Context-aware: Intelligent context building for better LLM responses
- 🚀 Easy integration: Works seamlessly with Trae Agent
- ⚡ Fast performance: Built-in caching mechanism
- 🔧 Customizable: Configurable chunking strategies and retrieval parameters
Installation
Option 1: Install from GitHub
git clone https://github.com/yourusername/rag-document-assistant.git
cd rag-document-assistant
pip install -e .
Option 2: Install from PyPI (coming soon)
pip install rag-document-assistant
Quick Start
In Trae Agent
-
Load the SKILL:
from skill_rag import load_document, query_document, get_status, clear_database -
Load a document:
result = load_document("/path/to/your/document.pdf") print(f"✅ Loaded: {result['success']}") -
Query the document:
result = query_document("What is this document about?", top_k=5) if result['success']: print(f"📝 Answer: {result['data']['context']}") -
Check status:
status = get_status() print(f"📊 Vector count: {status['data']['vector_count']}")
Using Slash Commands
from slash_commands import process_slash_command
# Load document
print(process_slash_command('/load --file /path/to/document.pdf'))
# Query document
print(process_slash_command('/query --question "Your question?"'))
# Get status
print(process_slash_command('/status'))
# Clear database
print(process_slash_command('/clear'))
Configuration
Configuration File (config.json)
{
"embedding_model": "sentence-transformers",
"vector_db_type": "chromadb",
"chunk_size": 512,
"chunk_overlap": 50,
"top_k": 5
}
Environment Variables
OPENAI_API_KEY: For OpenAI embedding model (optional)
Project Structure
RAG/
├── rag/
│ ├── core/ # Core RAG components
│ │ ├── loader.py # Document loaders
│ │ ├── chunker.py # Document chunkers
│ │ ├── embedder.py # Embedding models
│ │ ├── retriever.py # Retrieval strategies
│ │ └── context_builder.py # Context building
│ ├── db/ # Database implementations
│ └── scripts/ # CLI scripts
├── trae/ # Trae integration
│ ├── handlers.py # Command handlers
│ └── plugin.py # RAG plugin
├── .trae/
│ └── skills/
│ └── rag_document_assistant/
│ └── SKILL.md # SKILL definition
├── skill_rag.py # SKILL implementation
├── slash_commands.py # Slash command handler
├── requirements.txt # Dependencies
└── setup.py # Package configuration
Dependencies
- Python 3.9+
- numpy
- chromadb
- sentence-transformers
- PyPDF2
- tiktoken
- markdown
Usage Examples
Basic Document Analysis
from skill_rag import load_document, query_document
# Load document
load_result = load_document("report.pdf")
if load_result["success"]:
print(f"✅ Loaded {load_result['data']['chunks_processed']} chunks")
# Query document
query_result = query_document("What are the key findings?")
if query_result["success"]:
print(f"📝 Context: {query_result['data']['context']}")
Advanced Usage
from skill_rag import execute
# Using execute function
result = execute("query_document", {
"question": "Summarize the main points",
"top_k": 3
})
Contributing
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add some amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
License
This project is licensed under the MIT License - see the LICENSE file for details.
Support
For support, please open an issue on GitHub or contact the author.
推荐工具
换一个关键词,或者移除筛选条件。
安装
npx skillfish add xuanhuaz1022/traeofflinerag