# agentic-rag-jdemo-5-router **Repository Path**: createmaker/agentic-rag-jdemo-5-router ## Basic Information - **Project Name**: agentic-rag-jdemo-5-router - **Description**: agentic-rag-jdemo-5-router - **Primary Language**: Unknown - **License**: Not specified - **Default Branch**: master - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2026-05-03 - **Last Updated**: 2026-05-09 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # agentic-rag-jdemo-5-router Demo 5 (Java) — **Multi-corpus router**. The agent picks between `docs`, `code`, and `tickets` per query. Mirror of the Python sibling. ## Why three corpora? A real assistant rarely has one homogeneous knowledge base. You usually have at least: | corpus | shape | typical question | citation form | |---------|------------------------------------------------------|--------------------------------------|-----------------------| | `docs/` | TaskSaaS architecture / API / deployment markdown | "How does auth work?" | `[01-architecture#2]` | | `code/` | Annotated source snippets | "How is JWT issued?" | `auth_service.py` | | `tickets/` | Issue-tracker tickets `TS-NNN` | "Why do refresh tokens drop on restart?" | `[TS-204]` | Mixing them into one BM25 index pollutes rankings — ticket text full of stack traces beats clean doc prose for keyword queries it shouldn't win. Splitting also lets each source have a different refresh cadence. ## Architecture ``` question ──► Agent ──► search_router ──► (BM25 each corpus, pick highest top-1) ──► top-k from winner ├─► search_docs (force docs only) ├─► search_code (force code only) └─► search_tickets (force tickets only) ``` The router itself is intentionally dumb — BM25 each corpus, take whichever has the highest top-1 score, ties broken by registration order. Production systems use a small classifier or LLM-as-router; this version's pedagogical point is **transparency**: every router result includes per-corpus top-1 evidence so you can see _why_ a corpus won. ## Setup ```powershell cd D:\javacode\agentic-rag-jdemo-5-router mvn package copy .env.example .env ``` ## Run ```powershell java -jar target\arag.jar --ping java -jar target\arag.jar --list # 3 lists, one per corpus java -jar target\arag.jar --route "JWT refresh fails on restart" # see the router pick + evidence java -jar target\arag.jar "Why are issue assignment emails sent twice?" ``` ## Files ``` corpora/ ├── docs/ # TaskSaaS markdown (5 files, byte-identical to the Python sibling) ├── code/ # 3 annotated source snippets └── tickets/ # 3 tickets (TS-101, TS-204, TS-318) src/main/java/com/agentic/rag/ ├── Router.java # NEW — Corpus, RouteDecision, Router, loadDefault() ├── Tools.java # 4 search tools + read_doc cross-corpus ├── Agent.java # system prompt explains the 3 corpora + per-corpus citation forms ├── App.java # --list per corpus; --route Q to probe the router └── BM25Retriever.java / LLMClient.java / Document.java / ... # unchanged src/test/java/com/agentic/rag/ ├── RouterTest.java # NEW — 11 tests (load, route, evidence, get-across-corpora) └── SafeEvalTest.java ``` ## Tests ```powershell mvn test # 11 Router + 10 SafeEval = 21 tests ``` ## Cross-language sibling Mirror at `D:\PythonProjects\agentic-rag-demo-5-router` — corpora layout is byte-identical so you can compare router decisions across runtimes on the exact same inputs.