A powerful document analysis and query system based on Retrieval-Augmented Generation (RAG) technology, integrated as a Trae SKILL.
Обзор
A powerful document analysis and query system based on Retrieval-Augmented Generation (RAG) technology, integrated as a Trae SKILL. - 📚 : TXT, Markdown, and PDF documents - 🔍 : Advanced semantic search with hybrid retrieval strategy - 🧠 : Intelligent context building for better LLM responses - 🚀 : Works seamlessly with Trae Agent - ⚡ : Built-in caching mechanism - 🔧 : Configurable chunking strategies and retrieval parameters - OPENAI_API_KEY: For OpenAI embedding model (optional) - Python 3.9+ - numpy - chromadb - sentence-transformers - PyPDF2 - tiktoken - markdown 1. Fork the repository 2. Create a feature branch (git checkout -b feature/amazing-feature) 3. Commit your changes (git commit -m 'Add some amazing feature') 4. Push to the branch (git push origin feature/amazing-feature) 5. Open a Pull Request This project is licensed under the MIT License - see the LICENSE file for details. For support, please open an issue on GitHub or contact the author.
README
RAG Document Assistant
A powerful document analysis and query system based on Retrieval-Augmented Generation (RAG) technology, integrated as a Trae SKILL.
Features
- 📚 Multi-format support: TXT, Markdown, and PDF documents
- 🔍 Smart retrieval: Advanced semantic search with hybrid retrieval strategy
- 🧠 Context-aware: Intelligent context building for better LLM responses
- 🚀 Easy integration: Works seamlessly with Trae Agent
- ⚡ Fast performance: Built-in caching mechanism
- 🔧 Customizable: Configurable chunking strategies and retrieval parameters
Installation
Option 1: Install from GitHub
git clone https://github.com/yourusername/rag-document-assistant.git
cd rag-document-assistant
pip install -e .
Option 2: Install from PyPI (coming soon)
pip install rag-document-assistant
Quick Start
In Trae Agent
-
Load the SKILL:
from skill_rag import load_document, query_document, get_status, clear_database -
Load a document:
result = load_document("/path/to/your/document.pdf") print(f"✅ Loaded: {result['success']}") -
Query the document:
result = query_document("What is this document about?", top_k=5) if result['success']: print(f"📝 Answer: {result['data']['context']}") -
Check status:
status = get_status() print(f"📊 Vector count: {status['data']['vector_count']}")
Using Slash Commands
from slash_commands import process_slash_command
# Load document
print(process_slash_command('/load --file /path/to/document.pdf'))
# Query document
print(process_slash_command('/query --question "Your question?"'))
# Get status
print(process_slash_command('/status'))
# Clear database
print(process_slash_command('/clear'))
Configuration
Configuration File (config.json)
{
"embedding_model": "sentence-transformers",
"vector_db_type": "chromadb",
"chunk_size": 512,
"chunk_overlap": 50,
"top_k": 5
}
Environment Variables
OPENAI_API_KEY: For OpenAI embedding model (optional)
Project Structure
RAG/
├── rag/
│ ├── core/ # Core RAG components
│ │ ├── loader.py # Document loaders
│ │ ├── chunker.py # Document chunkers
│ │ ├── embedder.py # Embedding models
│ │ ├── retriever.py # Retrieval strategies
│ │ └── context_builder.py # Context building
│ ├── db/ # Database implementations
│ └── scripts/ # CLI scripts
├── trae/ # Trae integration
│ ├── handlers.py # Command handlers
│ └── plugin.py # RAG plugin
├── .trae/
│ └── skills/
│ └── rag_document_assistant/
│ └── SKILL.md # SKILL definition
├── skill_rag.py # SKILL implementation
├── slash_commands.py # Slash command handler
├── requirements.txt # Dependencies
└── setup.py # Package configuration
Dependencies
- Python 3.9+
- numpy
- chromadb
- sentence-transformers
- PyPDF2
- tiktoken
- markdown
Usage Examples
Basic Document Analysis
from skill_rag import load_document, query_document
# Load document
load_result = load_document("report.pdf")
if load_result["success"]:
print(f"✅ Loaded {load_result['data']['chunks_processed']} chunks")
# Query document
query_result = query_document("What are the key findings?")
if query_result["success"]:
print(f"📝 Context: {query_result['data']['context']}")
Advanced Usage
from skill_rag import execute
# Using execute function
result = execute("query_document", {
"question": "Summarize the main points",
"top_k": 3
})
Contributing
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add some amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
License
This project is licensed under the MIT License - see the LICENSE file for details.
Support
For support, please open an issue on GitHub or contact the author.
Рекомендуемые инструменты
Попробуйте другой запрос или уберите фильтр.
Установка
npx skillfish add xuanhuaz1022/traeofflinerag