Skip to content

Muundo — From Code Graph to Change Evidence

Direction

Muundo computes facts about a source tree, declares what it did and did not read, and lets people or systems verify those facts later.

What it refuses to copy is a mechanism: prompt hooks, injection into a model's context, relevance ranking, a viewer. Not the questions. "What calls this?" and "what does this file declare?" are asked by every tool in this area because they are the right questions, and Milestone 2 answers them — as cited structural queries, computed and replayable.

Its product question is:

After a code change, what did it touch, what evidence supports that answer, and what must be reconsidered before the change is trusted?

This serves ordinary teams reviewing agent-produced changes and, through optional adapters, high-assurance workflows. Those audiences must not force the same product on one another.

Product boundary

Layer Responsibility
Muundo Core Compute and reproduce entities, imports, calls, dependencies, metrics, reachability, completeness and provenance.
Change Evidence Compare two verifiable snapshots and report structural change plus revalidation scope.
Specialist adapters Import CodeQL, Joern, Semgrep, compiler and language-server evidence without duplicating their semantic analysis.
Domain adapters Link facts to requirements, hazards, standards and audits for Jagora and Litatoli. Optional.

Muundo is not a compiler, language server, security scanner, code-search product, LLM memory system, prompt hook or viewer. Agent-context tools may use Muundo evidence, but Muundo does not need to reproduce their user experience.

Principles

  1. Computed before described. Every result identifies source files, configuration and analyzer version before explanations are attached.
  2. Incomplete is an answer, and it needs a value of its own. Skipped files, unresolved imports and partial stages remain visible in derived results. A step that did not run says so in its own field, with a stable reason code beside the sentence — never through the emptiness of a list that means something else. The types stay per operation; the rule is shared, the vocabulary is not.
  3. No invented semantics. Precision comes from a proper extractor, compiler, language server or specialist adapter; heuristics are labelled.
  4. Local first. Core works without a build, network, database or model.
  5. Small evidence, not large reports. CI and agents need bounded, task-shaped output instead of a full graph in a prompt.
  6. One fact, many consumers. CLI, MCP, HTTP, Python and SDKs expose one report contract.
  7. Reproducibility is a feature. A result that cannot be replayed against the same tree is advice, not evidence.

The rule in the three places it applies

Found by holding the product to Principle 2 on 2026-09-20. One of the three was a live defect: verify answered reproduced: true about a re-analysis it had skipped, because the divergence list was empty and empty was read as agreement. Fixed — the verdict now carries a reproduction state and a reason code.

Where The states The gap
A call edge in_tree, outside, unresolved, undetermined Done. undetermined is 61-71 % of the edges measured, and saying so is the point: only an import the file writes makes a callee external.
A reproduction matched, diverged, not attempted Done. Published in REFERENCE.md, guarded by a test that fails the build on an unpublished code.
A compact view fields selected, fields omitted, truncation Done. A family that was not selected is absent rather than empty, and the view names it; a cut says what it kept and what there was; skipped_files and partial_analysis ride in every view whatever was asked for.

A projection must never weaken a verification. A diff explains the changes worth reading; a verification detects every difference in the fields it promises to check. They may share a comparator; a bounded diff can never stand in for the full comparison.

Success measures

Outcome Evidence Non-goal
Reviewers see patch impact A bounded report names changed entities, reverse callers, dependencies and gaps. Predicting every runtime behaviour.
CI knows whether evidence is usable It distinguishes verified, changed, incomplete and unable-to-answer. Treating analysis failure as a pass.
Agents receive structural context Queries return a cited subgraph with declared limits. Automatic prompt injection.
Specialist findings remain attributable Imported findings carry tool, version, configuration, revision and locations. Reimplementing CodeQL or Joern.
Domain traceability stays optional Jagora/Litatoli map facts to their own requirements and standards. Making requirements management a Core dependency.

First measured 2026-09-20, and the harnesses are in scripts/. Change evidence against git on pallets/click, dtolnay/anyhow and sindresorhus/ky and twelve more, in thirteen languages, over 25 commits each — 334 commits and 935 declarations: recall and precision 1.000 on every one, with nothing taken out of the question. Declarations a name cannot identify used to be excluded and were the one defect left; they are paired now, and measured like anything else. Call edges against the PyCG micro-benchmark, a prepared corpus of 112 reviewed cases under the same licence: precision 0.989 on edges called in_tree, honesty 0.916 on what it could not place, recall 0.326 — and 87 % of what it misses needs to know what a name holds, which is the gap a specialist adapter would close rather than a defect. Rust call edges against rustc, compared by FILE AND LINE rather than by name — -C debuginfo=2 says where every compiled function was written and where every call was made, which ends the mangling problem outright: on clap_builder, 3 569 compiled call sites, precision 0.953, honesty 0.984. C++ call edges against clang by the same route, since clang emits the same debug metadata: on fmt, 1 819 compiled call sites, precision 0.820, honesty 0.860. C call edges against clang as the oracle (-S -emit-llvm -O0 names the target of every direct call, so a prepared corpus is not needed) on sds, linenoise and cJSON: 990 calls in 10 files, precision 1.000, not one call placed wrongly, and every difference with the compiler accounted for — and four defects found on the way there, three of which made code vanish from the graph rather than answer wrongly. And a third way of measuring that needs no second tool at all: twist the source and check the answer moved. A call site is renamed to a name that exists once in the universe, that name is then declared in a place chosen in advance — the caller's own file, another file, a type's body, or nowhere — and the law each choice implies is absolute rather than comparative. 796 probes across C, C++, Rust, Python and TypeScript; every law held, after the method found that this->m() resolved nothing in C++ and five defects in its own harness. This is what reaches the languages no oracle covers. Every number, what it does not say, and the command that reproduces it: What has been measured.

Every milestone must be measured on at least three external repositories:

  • exactness of changed-entity and affected-caller sets against reviewed patches;
  • false-positive and false-negative rates for each declared heuristic;
  • median latency and report size for representative changes;
  • incomplete-analysis rate and stated causes;
  • reproducibility from recorded provenance.

Do not publish token-saving or quality claims without a versioned corpus, reproducible harness, raw results and an explicit baseline.

Milestones

0. Make the existing contract easy to adopt

Goal: a new team understands what Muundo proves today and can use it in CI.

  • Publish compact examples for analysis, verification, hotspots and failure on incomplete analysis.
  • Add a stable machine-readable completeness summary: complete, incomplete, or unable-to-answer, with reason codes.
  • Publish a per-language compatibility matrix: entities, imports, calls, metrics, reachability and static limits.
  • Keep resource ceilings, containment and cancellation contracts tested.
  • ~~Expose verification on every surface.~~ Done 2026-09-20: one engine in muundo_core::verify, called by the command line, by the MCP verify tool and by POST /verify, each with the guards that surface already applies.

Exit: a CI job fails closed on incomplete or stale evidence without parsing prose or guessing from an exit code, and an agent can verify what it was told.

1. Snapshot and change evidence

Goal: make Muundo useful after a pull request or agent edit.

  • ~~Define a versioned AnalysisSnapshot.~~ Done. Most of it is already in every report — per-file BLAKE3 hashes, analyzer and grammar versions, the full configuration, the configuration files that were read and the places one had to stay absent. What is missing is a single identity for the set, so a snapshot can be named in one string. A Git revision is context beside it, never the identity: an uncommitted tree has no revision, and the same revision analysed with different options is a different snapshot.
  • ~~Define compact views before freezing the schema, not after.~~ Done. A full analysis of muundo's own core/ is 10.1 MB — entities 58 %, call edges 37 %, and inside the entities body_features alone is 3.6 MB. Freezing today's report as the contract freezes a 10 MB default. A view selects field families and declares its source snapshot, its selection and its omissions.
  • ~~Add structural comparison between before and after reports.~~ Done. It is not an addition: verify compares by serialising both reports and testing whole field subtrees for equality, so it can say entities differs and never which entity. The diff replaces that comparison — without weakening it.
  • ~~Report supported deltas only: entities, imports, call edges, dependencies, metric deltas, new incompleteness and reverse callers to a chosen depth.~~ Done, plus scores and use/def modules.
  • Result-count limits are in (--max-items on views, diffs and queries, every cut declared). A BYTE-SIZE limit is not: a caller can bound how many items come back and not how many bytes, and one enormous entity is still enormous.
  • ~~Keep ambiguity visible: unresolved calls and partial analyses cannot become definitive impact claims.~~ Done: every answer carries complete, the count of unplaced calls that could belong in it, and whether the analysis behind it read the whole tree.

Exit: ~~a multi-file change yields a reproducible, bounded evidence document with locations and an explicit confidence boundary.~~ Met 2026-09-20 and checked on a real three-file change to core/: 12.9 KB against an 11.1 MB report, the files named, each declaration with its location and which of its fields moved, 35 entities to re-check, complete: false with 10 unplaced calls that could reach them, both snapshot ids, nothing unexplained.

2. Task-shaped structural queries

Goal: expose Core facts without sending an entire report to an agent.

Half of it is built, and the measurement that justified it now confirms it. The prerequisite — the snapshot kept between questions — exists, and so does an index beside each snapshot: the declarations, the edges and the totals without the bodies. muundo query entity|callers|callees|file|summary|freshness answers a 259 000-line tree in 85 ms, where analysing it to answer the same question takes 3.90 s. Every answer names its snapshot and says where it stops.

~~The same operations over MCP and HTTP~~ — done the same day: the snapshot and query tools, and POST /snapshot and POST /query, all reading the same index through the same core. What is left of this milestone is impact as its own answer rather than a traversal plus dependencies read separately.

Expose the same typed operations through CLI, MCP and HTTP:

  • ~~entity: declaration, metrics and direct relationships;~~ done (CLI)
  • ~~callers and callees: bounded traversal, resolved and unresolved edges apart;~~ done (CLI) — the traversal walks placed edges only and counts the unplaced ones that could have belonged in the answer
  • impact: reverse callers, imports, dependencies and limits;
  • ~~file API: declarations and signatures without bodies;~~ done (CLI)
  • ~~repository summary: totals, hubs, hotspots and completeness;~~ done (CLI)
  • ~~freshness: whether a snapshot still matches the tree and why it does not.~~ done (CLI) — the cheap half of a verification, and it names what it cannot see

Core does not perform natural-language ranking. A caller may translate a question to identifiers, but answers remain cited structural queries.

Exit: an MCP client receives bounded, reproducible context from the same Core report as every other surface.

3. Incremental analysis without weaker proof

Goal: avoid recomputing unchanged code while preserving replayability.

Measured 2026-09-20, and it is now the priority rather than an optimisation. Cold analysis is 0.95 s for muundo's core/ (86 700 lines) and 3.90 s for a 259 000-line tree. That is tolerable once and intolerable per question — and every question is a fresh analysis today, so this is what stands between the product and Milestone 2. The two share a prerequisite: a snapshot that survives between questions. Build the store first, then the queries over it, then the incremental recompute under it.

  • Cache parser output per content hash, grammar version and analyzer options.
  • Record dependency and configuration inputs that invalidate cached results.
  • Recompute affected graph regions; rebuild globally when invalidation cannot be bounded honestly.
  • Make verification compare cache-backed snapshots with clean replay.
  • Publish cold, warm and invalidated-file benchmarks separately.

Exit: a one-file edit performs less work than a cold run and a clean replay produces the same supported claims.

4. Specialist evidence adapters

Goal: complement stronger analyzers instead of competing with them.

Start with immutable import adapters:

  • CodeQL SARIF/path findings: rule, locations, source/sink path and database provenance;
  • Semgrep SARIF findings: rule, locations, configuration and revision;
  • Joern findings or exported traversals: query, locations, graph version and evidence boundary;
  • compiler and language-server diagnostics: version, configuration, locations and severity.

Muundo normalizes identity, revision and completeness metadata. It must never claim that its tree-sitter graph independently proved a third-party finding.

Exit: one change-evidence report can show Core facts beside a specialist finding while retaining the source and confidence of each claim.

5. Optional domain adapters

Goal: serve high-assurance workflows without making them mandatory.

  • Jagora adapter: map generated structure and requirement identifiers to entities and snapshots.
  • Litatoli adapter: attach a finding to a versioned standard clause and exact structural claims.
  • Port-comparison adapter: compare native ports while retaining language limits.
  • Data and state adapters: expose contracts, tables and state-flow points only where extractors support them.

Exit: each adapter stays outside muundo-core, can be removed without a base-schema change, and has fixtures proving both a finding and a limitation.

Product validation

Do not assume Kionowo needs are the market.

Run both segments at the same time, not one then the other. They fail for opposite reasons, and running them in sequence hides one of the two answers for months.

  1. Recruit five teams that use coding agents but do not use Jagora, Litatoli or regulated standards.
  2. Give them a Milestone 1 report on a real pull request. Measure whether it changed a review or test decision, not enthusiasm.
  3. In parallel, recruit three teams where an omission is expensive: medical, industrial, security-sensitive or public-sector software. Test the same Core evidence plus one adapter. C and Ada belong here — no comparable tool covers them — but the language is the way in, not the argument. The argument is that refusing to answer has a value where a wrong answer has a cost.
  4. Keep workflows that reduce measured review, regression or audit cost. Retire the rest.

Explicit non-goals

  • Copying another product's hooks, prompt injection, terminology, viewer or query-ranking behaviour.
  • Rebuilding CodeQL or Joern semantic depth inside tree-sitter extractors.
  • Requiring an LLM, embeddings, cloud service or database for Core.
  • Presenting a hotspot, heuristic reachability result or unresolved edge as a vulnerability or proof of runtime behaviour.
  • Making standards, requirements, tickets or human decisions mandatory inputs to a basic source-tree analysis.

First implementation slice

Contracts before code, and each step is small enough to land on its own.

  1. ~~Say what the verifier did not do.~~ Done 2026-09-20: verify carries a reproduction state — matched, diverged, not_attempted — with a published reason code. It answered reproduced: true about a replay it had skipped.
  2. ~~Measure the verifier by stage before attributing its cost.~~ Done 2026-09-20, and it moved: reading the report was 9.2 s of 10.8 s, because serde_json::from_reader pulls one byte at a time and nothing buffered the file. Buffered, a verify of core/ is 1.6 s — 0.14 s to read, 0.76 s to re-analyse, 0.64 s to compare. A unit test now counts reads per byte, since no test parsing an in-memory document could ever have seen it.
  3. ~~Give a call edge its resolution state~~ Done 2026-09-20: in_tree, outside, unresolved, undetermined, with undetermined 61-71 % of the edges measured and only an import the file writes making a callee external.
  4. ~~Fix the snapshot and view contract~~ Done 2026-09-20. snapshot_id names one analysis in one string — a digest of the sources, configurations, absences, options and engine, and of nothing that was found — and verify refuses a report whose id is not the id of its own inputs, before opening anything. analyze --fields answers with a view that names the snapshot it was cut from, the families it carries, the ones it does not and every collection it cut. core/ full is 11 MB; --fields entities is 947 KB.
  5. ~~Share the verifier across CLI, MCP and HTTP~~ Done 2026-09-20. The engine moved from the command line into muundo_core::verify, so all three surfaces run the same code and answer the same verdict: verify on MCP, POST /verify on HTTP, both behind the guards /analyze already had, and the replay is now cancellable because a surface that can drop a request has to be able to stop one.
  6. ~~Build the structural diff~~ Done 2026-09-20. muundo diff --before A --after B answers in 5 KB what two 11 MB reports differ by: files, declarations with the fields that moved, edges, metrics, scores, and what to re-check. It does not weaken the exhaustive comparison — it runs it too, and names every field the structure did not explain, so it can never answer "nothing changed" about two documents that are not the same. The revalidation set is walked over resolved edges only and counts the unplaced calls that could reach it, which makes it a stated lower bound rather than a claim.
  7. ~~Measure which pays most~~ Done 2026-09-20, and it changed the plan. A view saves bytes and no time: --fields entities --max-items 50 on a 259 000-line tree is 19 KB instead of 30.8 MB and takes 3.87 s where the full analysis takes 3.90 s. The output was never the cost — the tree is, and every question re-reads it. So task-shaped queries and incremental analysis are not two milestones but one lever: the snapshot kept between questions. Reading a report back is 0.14 s for 11 MB, so a stored snapshot turns a four-second question into a fraction of a second before any incremental work exists. That is the next thing to build. (Numbers from one machine on one day, reproducible with scripts/bench_workflows.sh.)

This created an original product surface: evidence-backed change analysis. It exists now — analyze, verify, diff and bounded views, on the command line, over MCP and over HTTP, each saying what it did not establish.

~~What the measurement says to do next: a snapshot store.~~ Built 2026-09-20. muundo snapshot save|list|show|ref|prune, a directory of reports addressed by snapshot id, outside the analysed tree — a store under the root would be read by the next analysis and change the snapshot it holds. A name is how a person says which one is the base, and a pinned base is never collected. muundo diff --before @base --after @head then answers after the code has moved on, which before this was simply impossible: the earlier tree was gone.

And the next lever, which it unblocks. A question still costs a full read of the document: of the 3.9 s a diff of two 30 MB reports takes, 1.6 s is the whole-document comparison run for honesty. Milestone 2 is now worth building — answer from something stored beside the snapshot instead of from the snapshot itself.