# neural-code-atlas **Repository Path**: felixagents/neural-code-atlas ## Basic Information - **Project Name**: neural-code-atlas - **Description**: fork:https://github.com/Bytefield/neural-code-atlas.git - **Primary Language**: Unknown - **License**: MIT - **Default Branch**: main - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2026-09-23 - **Last Updated**: 2026-10-08 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # Neural Code Atlas (NCA) > Measures how AI coding agents actually behave in your codebase and turns that into > decisions you can defend — including the decision to build nothing. Underneath it: a > local structural index of your code (tree-sitter + SQLite), exposed via CLI and MCP. Local-first. No cloud, no embeddings, no LLM calls. Everything runs on your machine and stays there. ## What it does today Two layers, one CLI. ### 1. Behavioural instrument — `nca corpus` *(new in 1.6)* Reads your Claude Code session logs and keeps only their *shape*: file paths, tool names, timestamps, session ids, counts. Never prompts, never code, never tool output. From that it answers, with frozen heuristics and no LLM: - Where does the agent repeatedly orient? What does it reread across sessions? - What kind of work consumes orientation — implementation, diagnosis, search, shell? - Which tools are actually in use, and are they helping? (NCA's own `nca_ask` included.) - What should change — and what should **not** be built? Each answer is a recommendation object (`dont_build` / `do` / `fix` / `no_intervention` / `insufficient_evidence`) with its evidence, its rule and threshold, a categorical confidence (`EVIDENCE_STRONG` … `INSUFFICIENT`, never a percentage), and the metric to remeasure. Markdown is just the renderer. ```bash nca corpus extract --project myapp --project-root /path/to/myapp --phase baseline --dry-run nca corpus extract --project myapp --project-root /path/to/myapp --phase baseline nca corpus orientation --project myapp --phase baseline nca corpus insights --project myapp # → ~/.nca/metrics/myapp/insights/ (never inside your repo) ``` On the one real corpus it has been validated against (93 sessions), it reproduced three decisions that had previously taken hours of manual analysis — two "don't build this" and one "fix this first" — purely from the events. That is the whole claim so far. Whether it generalizes to codebases and people it has never seen is the next test, not a done deal. ### 2. Structural index — `nca scan` and friends A persistent SQLite graph of functions, arrows, methods and classes plus your Markdown docs, with graph analytics computed on every scan (Louvain communities, PageRank, betweenness, god nodes). ```bash nca scan . # build/update the index; writes .nca/SKILL.md nca ask handleRequest # exact-name lookup; module, PageRank rank, god-node flag nca flow handleRequest # execution path from an entry point nca impact # callers, docs, security and silent-fallback risk per changed symbol nca evolve # architectural warnings: complexity, cycles, deep chains, god nodes ``` `nca_ask` is **exact-name retrieval** over indexed nodes. It is not semantic search and does not answer natural-language questions; on no match it says so explicitly and returns no graph content. `nca impact` is the more useful entry point for "what does this change touch". ## Where it is going NCA's roadmap is gated on evidence, not dates. Each step only starts once the previous one has produced measurable results: 1. **Decision ledger** — every recommendation → decision → outcome, so confidence can be calibrated against what actually happened. 2. **Treatment loop** — before/after measurement of an intervention (e.g. a context change) on the same codebase, per lane, with effect sizes and *n*. Quasi-experiment, not A/B, and labelled as such. 3. **Generalization** — other repos, then a corpus from someone who isn't the author. This is the test of whether NCA discovers knowledge or reconstructs one person's habits. 4. **Machine-readable judgment** — recommendation objects via `--format json`; proposed context changes as diffs a human accepts. NCA never edits your files on its own. 5. **Advisory preflight** — before a task, what to read first, what it touches, what prior evidence says. Read-only. Must be able to answer *no policy applies*. Anything beyond that — agents specialised per task class, orchestration — is explicitly not scheduled. If the evidence never justifies it, it never gets built. ## Install ```bash npm i -g @synio-es/neural-code-atlas nca --help ``` Native modules (`better-sqlite3`, `tree-sitter`) need build tools; see `INSTALL.md`. ## Commands **Corpus** — `nca corpus extract` (`--project`, `--project-root`, `--phase baseline|treatment`, `--since`, `--until`, `--dry-run`, `--include-worktrees`, `--min-events`) · `nca corpus orientation` (`--project`, `--phase`, `--json`, `--csv`, `--include-no-write`) · `nca corpus insights` (`--project`). **Code** — `nca ask ` · `nca flow ` · `nca impact [diff-spec]` (`--json`, `--air`, `--out`) · `nca evolve`. **Docs & vault** — `nca vault scan ` · `nca vault search ` · `nca vault get ` · `nca related ` · `nca docs audit`. **Context** — `nca task [description]` (`--show`, `--clear`) · `nca brief [--light]`. **Index** — `nca scan [path]` · `nca status` · `nca watch [path]` · `nca insights` · `nca projects` · `nca migrate`. **Server** — `nca mcp`. `nca --help` for full options. ## MCP server (Claude Code) ```json { "mcpServers": { "nca": { "command": "nca", "args": ["mcp"] } } } ``` Tools: `nca_ask`, `nca_flow`, `nca_status`, `nca_evolve`, `nca_insights`, `nca_projects`. The project is autodetected from the working directory; pass `project` to target another. The MCP server is a long-lived process. After rebuilding or upgrading NCA, restart the connection — a stale server against a migrated database reports `NCA schema version mismatch: db_version=… build_version=… db_path=…` with the fix. ## Graph analytics Computed on every scan, stored in the index, surfaced in `nca ask` and `SKILL.md`: **Louvain communities** (which files form natural modules), **PageRank** (load-bearing nodes), **betweenness** (bottlenecks everything routes through), **god nodes** (coupling above the p95 of the graph). `SKILL.md` — written to `.nca/` on every scan — is a token-efficient map: modules with node counts, top nodes by PageRank, god nodes, cycle and deep-chain counts, indexed docs. Safe to commit; it contains only structural metadata. ### Why a query is fast The graph is computed once, at scan, and written into the index: `edges` holds the resolved dependencies between nodes, `node_ranks` holds each node's PageRank position and god-node score. `nca ask` reads those tables instead of rebuilding the whole graph to print fifty lines. Measured on a real 50,670-node / 6,724-file index: | path | before | after | |---|---|---| | `nca_ask` over MCP (warm server) | 3,700–4,000 ms | 10–37 ms | | `nca ask` from a shell (new process) | ~6,500 ms | 170–260 ms, of which ~130 ms is Node starting up | | `nca scan` on unchanged files | four full graph builds | one shared snapshot: walk 4.4 s, link+derive 4.5 s, flows 1.3 s, SKILL.md 0.7 s | Two consequences worth knowing. Expansion follows the graph's resolution rules (same file → unique component alias → globally unique name → no edge) rather than matching every node that shares a name, and is capped at 400 nodes / 200 per hop — a second hop past that is noise inside a 600-token answer. And `nca ask`'s flow and warning blocks draw from at most 500 flows and 200 warnings, so a large index cannot make one answer a data dump; a query that matches no symbol returns empty `flows` and `warnings` in `--json` too, since nothing there is related to the question. An index written before these tables existed repairs itself: the first query rebuilds the derived graph once — slow, once — and prints `NCA|derived_rebuild|first_query_this_index` on stderr. `nca status` shows `derived: N edges / M ranks`, and flags ranks of zero as the state that needs a scan. ## Configuration `.nca/config.json` in the project root: ```json { "exclude": ["generated", "vendor"], "include_extensions": [".java", ".vue"], "exclude_extensions": [".cjs"], "max_file_size_kb": 256, "respect_gitignore": true, "path_aliases": { "@": "src", "@/store": "packages/store" }, "evolve": { "complexityThreshold": 10, "maxParamsThreshold": 7, "maxDepsThreshold": 15, "maxChainDepth": 6 } } ``` `include_extensions` **adds** to the built-in list (`.ts .tsx .js .jsx .mjs .cjs .py .java .vue`), it does not replace it — writing `[".vue"]` still indexes your TypeScript. Use `exclude_extensions` to take a built-in back out. ### What gets indexed `.gitignore` is followed by default, with git's own rules: a trailing `/` means directories, a leading or interior `/` anchors the pattern to the directory holding the file, a bare name matches at any depth, `!` re-includes, and a `.gitignore` in a subdirectory governs only that subtree (never the directory that holds it). Ignored directories are pruned rather than descended, so a large `node_modules` costs nothing, and files that fall out of scope are removed from the index on the next scan — a node for code git does not track is a phantom, not a gap. `nca scan -v` reports `gitignored_files:N|gitignored_dirs:M` and the summary block gains a `gitignored` line whenever scope was actually trimmed. Set `"respect_gitignore": false` to index the whole tree anyway — the case it exists for is generated code that is ignored by git but is the thing you want mapped. Ignored files are never listed as `unindexed`: that counter is for blind spots the project did not ask for, and an ignore rule is the project asking. Beyond `.gitignore`, directory names are excluded by a built-in table (`node_modules`, `dist`, `build`, `target`, `vendor`, `coverage`, `__pycache__`, `site-packages`, `.venv`, `venv`, `env`, anything starting with a dot) plus `exclude`, which appends to it. `exclude` matches directory *names* at any depth — write `"store"`, not `"packages/store"`. That table covers build output under its usual name. A repository that ships a minified library somewhere unusual — `public/static/jquery.min.js`, `web-app/js/lib/` — indexes it as its own code, and the several thousand one-letter functions a bundle declares dilute every ranking and fill the flow table with paths nobody wrote. Those files are usually in `.gitignore` already; when they are not, put the directory name in `exclude`. `path_aliases` tells the linker what the build tool already knows: `@/store/user` is a path, not a package, so it resolves to `/store/user.` and becomes a real dependency edge. Prefixes match on a `/` boundary — `@` never swallows `@element-plus/icons-vue` — and the longest one wins. An aliased specifier that names no file prints `NCA|unresolved||from:` once and stays a dependency; a specifier no alias covers is a package and stays quiet. Applying it needs `nca scan` (resolution is a write-path step), not a reparse. Languages: TypeScript (`.ts`, `.tsx`), JavaScript (`.js`, `.jsx`, `.mjs`, `.cjs`), Python, Java (`.java` — classes, interfaces, enums, records and annotation types, with `extends` / `implements` as dependencies; abstract and interface-declared methods are contracts, so they are not indexed as nodes), Vue (`.vue` — each `