# rag-backend **Repository Path**: HealthCare/rag-backend ## Basic Information - **Project Name**: rag-backend - **Description**: 备份: https://github.com/jianl-RH/rag-backend https://github.com/lj2tj/rag-backend - **Primary Language**: Unknown - **License**: Not specified - **Default Branch**: master - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2026-04-30 - **Last Updated**: 2026-04-30 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # RAG Backend Project This is a RAG (Retrieval-Augmented Generation) backend project that provides document parsing, chunking, and vector storage capabilities for intelligent question-answering. ## Features - Multi-format document parsing: Supports .docx, .pdf, .md, .txt, .xlsx files - Archive file support: Automatically extracts .zip, .gz, .tar, .tar.gz, .tgz, .tar.bz2, .tar.xz - Intelligent document chunking: Configurable chunk size and overlap for different document types - Vector store integration: Supports both FAISS and Chroma for efficient similarity search - RESTful API: Provides endpoints for document upload, vectorization, and question answering - Configurable settings: All settings are stored in config.yaml - Modular architecture: Clean, organized code structure with separated components - Ollama integration: Supports local LLM and embedding models via Ollama ## Project Structure ``` rag-backend/ main.py # FastAPI application entry config.yaml # Configuration file enums.py # Enumeration definitions logger.py # Logging setup pyproject.toml # Project dependencies requirements.txt # Alternative dependencies list src/ api/ # API endpoints __init__.py ask_question.py # Question answering endpoint embed_document.py # Document vectorization endpoint get_files.py # File listing endpoint upload_document.py # Document upload endpoint rag_backend/ # Core modules chunkers/ # Document chunkers __init__.py base.py factory.py markdown_chunker.py pdf_chunker.py table_chunker.py text_chunker.py word_chunker.py config/ # Configuration management __init__.py settings.py embedding/ # Embedding models factory.py parsers/ # Document parsers __init__.py base.py excel_parser.py factory.py markdown_parser.py pdf_parser.py text_parser.py word_parser.py util/ # Utility functions __init__.py api_helper.py file_helper.py file_status_helper.py vector_helper.py vector_store/ # Vector storage implementations base.py chroma.py factory.py faiss.py __init__.py ``` ## Installation ### Prerequisites - Python 3.11+ - uv package manager (recommended) or pip - Ollama (for local LLM and embedding models) ### Steps 1. Clone the repository 2. Install Ollama: https://ollama.com/download 3. Pull required models: ```bash ollama pull qwen3:8b ollama pull nomic-embed-text-v2-moe:latest ``` 4. Install dependencies: ```bash # Using uv (recommended) uv sync # Using pip pip install -r requirements.txt ``` 5. Run the application: ```bash # Development mode (auto-reload) uv run uvicorn main:app --port 3001 --reload # Production mode uv run uvicorn main:app --workers 4 --port 3001 ``` The API will be available at `http://localhost:3001`. ## Configuration Update `config.yaml` to customize settings: ```yaml document_parser: supported_types: - ".docx" - ".pdf" - ".md" - ".txt" - ".xlsx" document_chunker: chunk_settings: ".docx": chunk_size: 500 overlap: 50 ".pdf": chunk_size: 500 overlap: 50 ".md": chunk_size: 500 overlap: 50 ".txt": chunk_size: 500 overlap: 50 ".xlsx": chunk_size: 500 overlap: 50 docs_paths: - "./assets/docs" - "./assets/uploaded_docs" vector_store_path: "./vector_store_db" max_batch_size: 100 chat_model: "qwen3:8b" embedding_model: "nomic-embed-text-v2-moe:latest" embeddings_provider: "ollama" vector_store_type: "chroma" logging: log_file: "./logs/app.log" log_level: "INFO" ``` ## API Endpoints ### POST /upload_document Upload a document to the server. Supports archive files (.zip, .tar, etc.) which will be automatically extracted. **Request:** - `file`: File to upload (multipart/form-data) **Response:** ```json { "message": "File 'example.docx' uploaded successfully.", "file_path": "uploaded_docs/example.docx" } ``` ### GET /files Get all files from file-status.json with their embedding status. **Response:** ```json [ { "name": "example.docx", "path": "uploaded_docs/example.docx", "embeded": "not_started" }, { "name": "test.pdf", "path": "uploaded_docs/test.pdf", "embeded": "completed" } ] ``` ### POST /embed_document Vectorize a document that has been uploaded but not yet vectorized. **Request:** - `file_path`: Relative path to the file (multipart/form-data) **Response:** ```json { "message": "Document 'uploaded_docs/example.docx' vectorized successfully.", "file_path": "uploaded_docs/example.docx" } ``` ### POST /ask Ask a question based on the knowledge base. **Request:** ```json { "query": "What is the project about?" } ``` **Response:** ```json { "answer": "This is a RAG backend project that provides document parsing, chunking, and vector storage capabilities for intelligent question-answering." } ``` ## Usage Examples ### Uploading a Document ```bash curl -X POST "http://localhost:3001/upload_document" \ -F "file=@example.docx" ``` ### Uploading an Archive (ZIP/TAR) ```bash curl -X POST "http://localhost:3001/upload_document" \ -F "file=@documents.zip" ``` ### Getting All Files ```bash curl -X GET "http://localhost:3001/files" ``` ### Vectorizing a Document ```bash curl -X POST "http://localhost:3001/embed_document" \ -F "file_path=uploaded_docs/example.docx" ``` ### Asking a Question ```bash curl -X POST "http://localhost:3001/ask" \ -H "Content-Type: application/json" \ -d '{"query": "What are the key features?"}' ``` ## Supported File Types | File Type | Description | |-----------|-------------| | .docx | Microsoft Word documents | | .pdf | PDF documents | | .md | Markdown files | | .txt | Plain text files | | .xlsx | Excel spreadsheets | | .zip/.tar/.gz/.tar.gz/.tgz/.tar.bz2/.tar.xz | Archive files (automatically extracted) | ## Chunking Strategies | File Type | Default Chunk Size | Default Overlap | |-----------|-------------------|----------------| | .docx | 500 | 50 | | .pdf | 500 | 50 | | .md | 500 | 50 | | .txt | 500 | 50 | | .xlsx | 500 | 50 | ## Vector Store Options | Option | Description | |--------|-------------| | FAISS | Fast, lightweight vector store suitable for smaller datasets | | Chroma | Feature-rich vector store with more advanced capabilities (default) | ## Technology Stack - Framework: FastAPI - LLM: Ollama (Qwen3) - Embeddings: Ollama (Nomic Embed Text) - Vector Store: Chroma / FAISS - Document Parsing: python-docx, pypdf, pdfplumber, openpyxl, markdown - Configuration: PyYAML - Package Manager: uv ## License MIT