# rag-chatbot **Repository Path**: likongze/rag-chatbot ## Basic Information - **Project Name**: rag-chatbot - **Description**: No description available - **Primary Language**: Python - **License**: Not specified - **Default Branch**: main - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2026-05-03 - **Last Updated**: 2026-05-03 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # 🧠 RAG-Driven Question Answering API (TinyLlama + FAISS + LoRA) ## 🎯 Objective: Why Retrieval-Augmented Generation (RAG)? Modern language models have strong reasoning skills, but limited access to **real-world or domain-specific knowledge**. To solve this, **Retrieval-Augmented Generation (RAG)** combines: > πŸ”Ή *Information retrieval* (from documents) > πŸ”Ή *Generative reasoning* (via an LLM) so the chatbot can answer based on **grounded, factual context**. This project demonstrates how to build a **domain-adapted RAG pipeline** that can: * Ingest PDFs and text files into a vector database. * Retrieve semantically similar passages to user questions. * Use a local language model to generate coherent, evidence-based answers. * Optionally improve model understanding with **LoRA fine-tuning**. The example use case focuses on **mental health and social policy in Indonesia**, using a curated dataset of PDF articles and text reports. --- ## 🧩 Architecture Overview The architecture consists of **three core modules**: 1. 🧱 **Document Indexing (Memory Builder)** * Loads PDFs and TXT files from the `/data/sample_docs` folder. * Splits content into smaller chunks using LangChain’s `RecursiveCharacterTextSplitter`. * Encodes text using HuggingFace **Sentence Transformers** (`all-MiniLM-L6-v2`) into embeddings. * Stores all embeddings in a **FAISS** vector database for fast semantic search. 2. 🧠 **Model Fine-Tuning (LoRA Adapter)** * A lightweight fine-tuning layer on top of the **TinyLlama** model using **PEFT + LoRA**. * Trained with a small curated dataset (`mental_health_qa.jsonl`) to specialize in Indonesian mental health topics. * The resulting adapter weights are automatically detected and loaded during runtime. 3. ⚑ **RAG Chatbot Runtime (FastAPI)** * A user sends a question through the `/ask` endpoint. * The question is embedded and compared in FAISS to retrieve the most relevant text chunks. * A prompt is dynamically composed with retrieved context and fed to **TinyLlama-LoRA** for grounded generation. * The API responds with the final answer and the sources used for context. --- ### πŸ–ΌοΈ System Architecture Diagram Below is a visual overview of the system pipeline: ![RAG Chatbot Architecture](assets/architecture_diagram.png) *(Diagram by Muhammad Zakaria Saputra β€” showing the relationship between Document Indexing, LoRA Fine-Tuning, and FastAPI Runtime.)* --- ### πŸ”§ Component Flow Summary | Phase | Process | Tools / Libraries | | ------------------ | ----------------------------------------------- | ------------------------------------------- | | **Data Ingestion** | Load PDFs & TXT files | `LangChain Community Loaders` | | **Chunking** | Split text into overlapping segments | `RecursiveCharacterTextSplitter` | | **Embedding** | Convert chunks to dense vectors | `Sentence Transformers (Hugging Face)` | | **Vector Storage** | Store and retrieve top-k relevant chunks | `FAISS` | | **Fine-Tuning** | Train LoRA adapter for domain Q&A | `PEFT`, `TinyLlama`, `Transformers` | | **Inference** | Answer user queries with retrieval + generation | `FastAPI`, `TinyLlama-LoRA`, `Transformers` | --- ## βš™οΈ Tech Stack | Component | Purpose | Framework/Library | | ------------------- | ---------------------------- | ------------------------------------------------------------ | | **Backend API** | Serve endpoints | [FastAPI](https://fastapi.tiangolo.com/) | | **Vector Store** | Context retrieval | [FAISS](https://github.com/facebookresearch/faiss) | | **Embeddings** | Semantic similarity encoding | [HuggingFace Sentence Transformers](https://www.sbert.net/) | | **Model Runtime** | Local generation | [Transformers](https://huggingface.co/docs/transformers) | | **Fine-tuning** | Lightweight adaptation | [PEFT + LoRA](https://huggingface.co/docs/peft) | | **Docs Processing** | PDF & TXT parsing | [LangChain Community Loaders](https://python.langchain.com/) | --- ## 🧱 Project Structure ``` rag-chatbot/ β”œβ”€β”€ app/ β”‚ β”œβ”€β”€ main.py # FastAPI server and routes β”‚ β”œβ”€β”€ rag_pipeline.py # Core RAG logic (retriever + generator) β”‚ β”œβ”€β”€ indexing.py # Builds FAISS vector index β”‚ β”œβ”€β”€ config.py # Model and path configuration β”‚ β”œβ”€β”€ finetune/ β”‚ β”œβ”€β”€ train_lora.py # LoRA fine-tuning script β”‚ └── adapter/ # Fine-tuned adapter weights (if available) β”‚ β”œβ”€β”€ data/ β”‚ β”œβ”€β”€ sample_docs/ # PDF/TXT source documents β”‚ └── mental_health_qa.jsonl # Fine-tuning dataset β”‚ β”œβ”€β”€ requirements.txt └── README.md ``` --- ## πŸš€ How to Run the Project ### 1️⃣ Clone the Repository ```bash git clone https://github.com/zakariasaputra/rag-chatbot.git cd rag-chatbot ``` ### 2️⃣ Create a Virtual Environment ```bash python -m venv venv # On Windows venv\Scripts\activate # On macOS/Linux source venv/bin/activate ``` ### 3️⃣ Install Dependencies ```bash pip install -r requirements.txt ``` ### 4️⃣ Prepare the Documents Add your `.pdf` and `.txt` files to: ``` data/sample_docs/ ``` For example: ``` data/sample_docs/ β”œβ”€β”€ Free from pasung.pdf β”œβ”€β”€ Barriers and facilitators.pdf β”œβ”€β”€ Cultural diversity in beliefs.pdf β”œβ”€β”€ Problems Among Indonesian Adolescents.pdf β”œβ”€β”€ First 1000 Days.pdf └── Pasung.txt ``` ### 5️⃣ Build the FAISS Index ```bash python -m app.indexing ``` Expected output: ``` πŸ“š Indexed 6 files β†’ 477 chunks total. πŸ’Ύ FAISS index saved to vector_index/ ``` ### 6️⃣ Run the API ```bash uvicorn app.main:app --reload ``` Visit: πŸ‘‰ [http://127.0.0.1:8000/docs](http://127.0.0.1:8000/docs) to open the Swagger UI and test the `/ask` endpoint. ### 7️⃣ Ask a Question ```bash curl -X POST http://127.0.0.1:8000/ask \ -H "Content-Type: application/json" \ -d '{"question": "What is pasung?"}' ``` Output: ```json { "question": "What is pasung?", "answer": "Pasung is a long-standing custom in Indonesia that involves shackling individuals with mental illness.", "context_sources": [ "Free from pasung.pdf (page 1)", "Pasung.txt (page ?)" ], "metadata": { "model": "tinyllama-LoRA", "retrieval_engine": "FAISS", "timestamp": "2025-10-19T09:00:00Z" } } ``` --- ## 🧠 Fine-Tuning with LoRA (Optional) You can enhance the model’s factual alignment using your own Q&A dataset. ### Dataset Format `data/mental_health_qa.jsonl` ```json {"instruction": "What is pasung?", "output": "Pasung is a traditional practice in Indonesia where people with mental illness are restrained using shackles or wooden blocks."} {"instruction": "What are the main barriers to mental health care?", "output": "Limited access, stigma, and insufficient health facilities are the main challenges."} ``` ### Run Fine-Tuning ```bash python finetune/train_lora.py ``` This creates LoRA adapter weights in: ``` finetune/adapter/ ``` Once generated, the main pipeline automatically detects and loads them: ``` 🧩 Found LoRA adapter. Loading adapted weights... βœ… RAG pipeline ready. Using model: tinyllama-LoRA ``` --- ## πŸ’¬ Example Questions to Try Once the API is running, open [http://127.0.0.1:8000/docs](http://127.0.0.1:8000/docs) and test the `/ask` endpoint. Here are some questions you can try, all answerable based on the indexed documents: | Category | Example Question | Description | | --------------------------- | ---------------------------------------------------------------------------- | ------------------------------------------------------------------ | | 🧠 **Mental Health Policy** | `What is pasung in Indonesia?` | Tests contextual understanding of cultural practices. | | πŸ₯ **Healthcare Barriers** | `What are the main challenges to accessing mental health care in Indonesia?` | Retrieves key findings from the β€œBarriers and Facilitators” paper. | | πŸ‘Ά **Early Development** | `Why is the first 1000 days of life important?` | Checks retrieval from the β€œFirst 1000 Days” document. | | 🌏 **Cultural Perspective** | `How do cultural beliefs affect mental health treatment?` | Combines reasoning with data from β€œCultural Diversity in Beliefs”. | | 🧩 **Adolescent Wellbeing** | `What mental health problems are common among Indonesian adolescents?` | Pulls content from β€œProblems Among Indonesian Adolescents.pdf”. | | βš–οΈ **Policy Evaluation** | `What steps has Indonesia taken to eliminate pasung?` | Uses retrieved context to summarize national initiatives. | --- ## πŸ“Š Example Use Cases * Summarizing insights from internal PDFs * Answering domain-specific questions (e.g., policy, research) * Training lightweight AI assistants with local documents * Demonstrating fine-tuning workflows for small models --- ## πŸ“ˆ Design Highlights | Feature | Description | | --------------------------- | ------------------------------------------------- | | **RAG with FAISS** | Efficient semantic retrieval from document chunks | | **HuggingFace Integration** | Local text generation using Transformers | | **Auto LoRA Loading** | Detects and merges fine-tuned adapter weights | | **Explainable Responses** | Includes document source references | | **Modular Components** | Easy to extend with new models or datasets | --- ## βš™οΈ Environment Requirements | Requirement | Version | | ------------ | ------- | | Python | β‰₯ 3.10 | | torch | β‰₯ 2.0 | | transformers | β‰₯ 4.40 | | fastapi | latest | | uvicorn | latest | | faiss-cpu | latest | --- ## πŸ‘€ Author **Muhammad Zakaria Saputra** TensorFlow Developer | AI Product Builder | Applied Research Enthusiast πŸ’¬ *β€œRAG bridges what models know and what they should know, making AI systems both factual and grounded.”*