Skip to content

What has been measured

Every number here was produced by a script in scripts/, against a repository anybody can clone. Nothing is an estimate, and the commands are given so you can disagree with them.

Three instruments, three questions

The question How
Change evidence Did it name the declarations a commit touched? git says which lines changed; muundo diff is checked against it
Call edges Is an edge the edge the compiler compiled? clang or rustc, compared by file and line so no name has to be matched
Probes Does it see a declaration placed where the language makes it visible? the source is twisted in a way chosen in advance, and only muundo's own two reports are read

The third needs no second tool, which is why it reaches every language.

Per language

Language Change evidence, 25 commits Call edges Probes
C 3 repositories, 1.000 / 1.000 clang, 990 sites, precision 1.000 2 188
C++ fmt, 1.000 / 1.000 clang, 1 819 sites, precision 0.968, 582 placed 665
Rust 3 crates, 1.000 / 1.000 rustc, 3 569 sites, precision 0.971, honesty 0.957, 1 979 placed 1 692
Python click, 1.000 / 1.000 PyCG corpus, precision 0.989, honesty 0.915 720
TypeScript ky, 1.000 / 1.000 tsc, 10 536 sites, precision 0.995, 2 042 placed 664
JavaScript express, 1.000 / 1.000 tsc, 14 490 sites, precision 0.999, 964 placed 1 367
TSX excalidraw, 1.000 / 1.000 tsc, 56 070 sites, precision 0.993, honesty 0.983, 18 085 placed 1 621
ArkTS photos, 1.000 / 1.000 tsc, 5 141 sites, precision 1.000, honesty 1.000, 2 383 placed 598
Go uuid, 1.000 / 1.000 callgraph on cobra, 817 sites, precision 1.000, honesty 1.000, 780 placed 2 271
Java gson, 1.000 / 1.000 javac, 23 291 sites, precision 1.000, honesty 0.997, 10 357 placed 1 482
C# Newtonsoft.Json, 1.000 / 1.000 Roslyn, 42 454 sites, precision 0.985, honesty 0.986, 13 113 placed 1 275
Kotlin okio, 1.000 / 1.000 klaxon's own bytecode, 848 sites, precision 1.000, honesty 1.000, 388 placed 1 573
Ada ada-util, 1.000 / 1.000 GNAT .ali, 2 130 sites, precision 1.000, honesty 1.000, 2 088 placed 702

Of the calls whose target really IS in the tree, how many are named

One script produces every row — scripts/corpus_score.py <language> <corpus> <muundo> [<server>] — because comparing per-language buckets by hand is how a report comes out wrong. Each harness counts calls whose target is a LIBRARY alongside calls that land in the tree, and reading the total as "what muundo found" overstates the gap by half.

coverage is named / (named + missed + wrong): of the edges that end inside the tree, the share muundo names at the right declaration. left is counted apart — a call whose target is outside the tree and that muundo said so about, which answers a different question.

What is left, and what is a wall

A coverage figure says how much is missed. It does not say whether the rest is work. scripts/corpus_ceiling.py <language> <corpus> asks one question of every miss: does muundo declare anything at the position the oracle names? Where it declares nothing, there was nothing muundo could have answered.

corpus named reachable misses outside the model ceiling
excalidraw (TSX) 91.9 % 762 449 96.5 %
ky (TypeScript) 74.1 % 421 416 87.0 %
express (JavaScript) 86.3 % 91 60 94.5 %
fmt (C++) 28.8 % observed 1 858 130 94.9 %

Half of what ky still misses is in the second column. The shape is always the same:

new Promise((resolve) => { … resolve(value) … })
const Footer = ({ onChange }: { onChange: () => void }) => { … onChange() … }

resolve and onChange are PARAMETERS. The TypeScript compiler resolves the call to the parameter's own declaration; muundo's model carries entities for declarations — functions, classes, members — and a parameter holding a callback is not one. A report cannot name what it does not carry, so closing that gap means deciding that a parameter is a declaration, which is a different product and a different set of numbers everywhere else.

Reading the source alone, with no compiler and no build:

Language corpus the right overload the right method named overload-only precision
C cJSON 100 % 100 % 306 0 1.000
ArkTS photos 92.7 % 92.7 % 2 224 0 1.000
TSX excalidraw 91.9 % 92.5 % 15 347 91 0.993
JavaScript express 86.3 % 86.3 % 967 0 0.999
Go cobra 82.8 % 82.8 % 453 0 1.000
Rust clap 78.9 % 81.0 % 1 799 47 0.975
TypeScript ky 74.1 % 74.2 % 2 419 3 0.994
Java gson 73.0 % 96.0 % 8 179 2 577 0.9998
Kotlin klaxon 68.1 % 72.2 % 515 31 0.992
C# Newtonsoft.Json 66.3 % 94.3 % 12 851 5 414 0.986
Python PyCG corpus 63.5 % 63.5 % 160 0 0.994
Kotlin KotlinPoet 45.9 % 75.9 % 383 251 0.955
Ada ada-util 98.0 % observed 98.0 % 2 088 0 1.000
C++ fmt 28.8 % observed 42.4 % 826 390 0.987
C++ leveldb 51.3 % observed 51.3 % 533 0 0.945

Two columns because there are two questions. The first asks whether the call reaches the right DECLARATION — which overload of SerializeObject, of eight. The second asks whether it reaches the right METHOD, which is what a reader of a call graph follows. muundo's qualified names do not distinguish overloads, so a call it places perfectly still cannot say which one: Src/…/JsonConvert.cs::JsonConvert::SerializeObject is eight declarations.

The gap between the columns IS that, and it is the largest single lever left: 5 077 of C#'s 6 371 misses and 2 577 of Java's 3 015 are calls placed on the right method whose overload is unnamed. Kotlin and C++ have it too (251 and 379). Languages without overloading — C, Go, Python, ArkTS — have two identical columns.

The last two rows publish no headline: 30 of ada-util's 222 units and 29 of fmt's 47 do not compile, so the oracle saw part of the tree.

These are protocol v2 and are NOT comparable to anything published before 2026-09-27. What changed is the measurement, not the engine — see below.

The last two publish no headline: 30 of ada-util's 222 units and 29 of fmt's 47 do not compile, so the oracle saw part of the tree and a rate over part of a tree is not that tree's rate. observed_coverage carries the number with oracle_status: partial beside it.

These are protocol v2 and are NOT comparable to anything published before 2026-09-27. What changed is the measurement, not the engine — see below.

A compiler calls what the author did not write

Status::Corruption("msg") takes a const Slice&, so passing a string builds a Slice — and clang records a call to Slice::Slice on a line that writes Slice nowhere. A Rust #[derive] is the same shape, and so is every C++ copy constructor and destructor: real calls in the binary, and no call a reader of that line could point at.

muundo reads source. An answer it cannot see in the source is not a miss, and counting it as one measures the distance between two definitions of the word "call". Those expectations are counted apart now, under not_written, and the number excluded is published beside the score:

corpus judged excluded as not written
leveldb 1 936 545
fmt 1 016 230

Only a CONSTRUCTOR or a TYPE is excluded, and only where the line does not name it. That is the whole phenomenon: a temporary built to convert an argument, a copy, a destructor. s.ok() stays judged because ok is a method; new leveldb_t stays judged because leveldb_t is right there.

A first version of this rule was wrong and the numbers it produced were published for an hour. It excluded any expectation whose declaration the line did not name — which is exactly what a renaming import does, so it dropped 655 of express's calls and read 66.8 % where the truth is 86.2 %. The same blind spot as keying a site on the resolved name, in a new guise. The figures in this document are the narrow rule's.

Two bugs in the C++ oracle, and both hid the same thing

C++ read 51.4 % on fmt, and a second corpus was added to see whether that was the language or the project. leveldb — ordinary object-oriented C++, the kind a regulated codebase is written in — read 28.6 % with precision 0.586. A precision that bad is either a serious defect or a broken instrument, and it was the instrument, twice.

One translation unit at a time. db/c.cc calls NewBloomFilterPolicy, which util/bloom.cc defines — and in db/c.cc's own module that is a declare and nothing more. The oracle read each unit alone, so every call crossing a .cc boundary was scored as leaving the tree, and muundo was right about every one of them. The tree is all of its units: the harness now gathers what they all define before judging any of them. 123 of 227 disagreements.

An invoke writes its !dbg on the next line. LLVM spells a call that can throw as invoke … to label … unwind label …, !dbg !N, and the pattern reading the IR could not cross the newline. Every call that can throw was invisible to this oracle — which in ordinary C++ is most of them. The oracle saw 650 more calls on leveldb once it could read them, and 158 more on fmt.

Neither was visible on fmt alone: it is header-only, so nothing crosses a unit, and its hot paths are noexcept. fmt's published figure was 51.4 % measured against an oracle missing 158 calls; it is 50.1 % against one that sees them.

leveldb now reads 38.7 % with precision 0.865. What is left is named in the design record: a construction of a plain struct, which muundo reports as a dependency by design and the compiler runs no code for, and a method reached through a member whose type a header declares.

A <builtin> was counted as an in-tree expectation

PyCG's answer files list nine edges whose target is <builtin>.len or <builtin>.map. Those name no node of the tree — len is the interpreter's — so muundo cannot reach them with an in-tree edge, and outside is the right answer rather than a miss. expected_edges drops them now, which is worth 2.4 points of the Python row and is a correction to the instrument, not to the engine.

Protocol v2 — what the old numbers were counting

Every figure above moved on 2026-09-27, most of them down, and no line of the engine changed that day. Two flaws in the instrument:

A site muundo said nothing about left the score. The rate divided by named + honest + wrong, and a position where muundo produced no edge at all landed in missing, which was in none of those. So removing an edge could take the site out of the denominator and LIFT the percentage. That is not a hypothetical: when casts stopped being calls, fmt's population fell from 1 079 to 1 060 while named stayed at 614, and the rate rose from 56.90 % to 57.92 % with not one additional call named. The denominator is now every expectation the oracle can judge, which does not move when the answers do — scripts/test_corpus_protocol.py asserts it: the same comparison with four answers instead of ten keeps the same denominator and scores 0.4.

An external call paid for an internal one. The class-based adapters — Go, C, Ada, ArkTS — counted agreement about a call that LEAVES the tree in the same total as one that lands in it. Reading that as "what muundo found" overstated it: cobra's 855 agreements were 453 internal and 402 external, and internal recall is 453 of 547.

Two smaller ones came with it. Precision now counts only answers that DREW an in-tree edge, because saying outside about an internal call draws nothing — it costs recall and cannot cost precision. And the Kotlin oracle reads javap -s and keys on (owner, name, descriptor), so two overloads are two identities: they used to collapse under one name, which let an answer naming the wrong overload count as right.

The numbers are lower and they answer the question the table claims to answer. A row also carries the digest of the expectations it was compared against and of the sources they were read from, so it can be re-derived rather than believed.

And they are a LOWER bound, for one reason worth knowing. v1 excluded an expectation whose target the call's own line does not name — a temporary built to convert an argument, a copy constructor, a #[derive]: real calls in the binary that no reader of that line could point at. v2 stopped, because the exclusion was computed from muundo's OWN report, so the set of excluded expectations moved as the engine changed. That is the one thing a denominator must not do, so dropping it was right — but the phenomenon is real, and those calls are now counted as misses. An engine-independent version would ask the ORACLE for the target's name rather than asking muundo which names it recorded at that position. what_the_source_wrote is kept, out of use, with that written on it.

C++'s figure is bounded by PARSING, not by resolution

fmt is 750 syntax errors across 39 files. leveldb is 116 across 28 — six times fewer for a tree of similar size. That is the whole distance between them, and it is why a language whose C sibling scores 100 % sits at 57.9 %.

What that costs, concretely: class buffer in core.h is recorded as spanning lines 1759 to 1773 because the parser stopped at the first constructor it could not read. push_back is at line 1829, so it is not an entity at all — and no call to it can resolve, however good the resolver is. muundo says so in every report: partial_analysis names each file, its error count, and that "code under those regions was NOT read".

Three macro shapes were tried on a rewritten copy, each length-preserving so no position moved: the namespace-opening macro (FMT_BEGIN_NAMESPACE), the SFINAE macro inside a template parameter list (FMT_ENABLE_IF), and the pragma macro at namespace scope (FMT_PRAGMA_GCC). Masking all three took fmt from 750 errors to 658. The remainder is spread across the grammar's handling of template metaprogramming, not concentrated in a construct that could be masked.

So the honest statement about C++ is two numbers, not one: what the resolver does with the code it can read, and how much of the code it can read. fmt is the extreme of the second.

leveldb's oracle shrank, so its row is gone

scripts/corpus_cpp.py compiles one translation unit at a time and skips the ones that do not build. On 2026-09-26 it builds 29 of leveldb's 76 and sees 1 229 call sites; the published 65.8 % was measured when it saw 1 936. Three different muundo binaries — before this work, after it, and the commit in between — all score 69.4 % against the smaller oracle, so nothing in muundo moved. A number measured over 38 % of a project is not that project's score, and fmt's row is the C++ figure to read until leveldb builds again.

Kotlin's row was one small corpus

klaxon judges 238 calls. That is a hundred-file JSON library, and 45.0 % of 238 is not a language's score — it is one project's shape, measured once.

KotlinPoet doubles the population and answers 76.0 % over 491 in-tree calls: builders, sealed types, extension functions and a DSL, which is what modern Kotlin looks like. Both rows stay, because the pair says more than either.

Two rules moved that row from 58.0 %, and they are the two shapes Kotlin code is written in. A val b = Thing.builder() says what b holds, so the calls on b reach the nested Builder — that is the same value_bindings fact Rust and TypeScript already filled, now filled by the Kotlin parser too. And a call on a receiver whose name is declared exactly once at FILE scope is an extension function, xs.toImmutableList(), which is a member of nothing.

klaxon gained less (41.5 → 45.0 %) and lost precision (1.000 → 0.947): six of its calls now name the wrong target, all of them a basename declared once at file scope in a tree that also declares it as a member elsewhere. The rule is right about the shape and wrong about those six, and the trade is published rather than hidden.

Building it takes one command, and the corpus is the compiled sources:

kotlinc -Xjvm-default=all -cp kotlin-reflect.jar \
  -d /var/tmp/muundo/kotlinpoet-classes \
  $(find …/kotlinpoet/kotlinpoet/src/commonMain -name '*.kt')
MUUNDO_KOTLIN_CLASSES=/var/tmp/muundo/kotlinpoet-classes \
  scripts/corpus_score.py kotlin …/kotlinpoet/kotlinpoet/src/commonMain <muundo>

Two bigger corpora were tried and refused. okio is multiplatform, and kotlinc rejects an expect and its actual in one module — a multiplatform build is the only way to compile it, and this oracle reads what one kotlinc run wrote. OkHttp 4.12 compiles down to thirteen errors with five compile-only jars on the classpath (okio, animal-sniffer, Conscrypt, BouncyCastle, OpenJSSE, an Android stub), and the last of them need Android APIs no published stub jar carries. Neither is a muundo limit; both are what the corpus costs.

One flaw of this instrument, and what it cost

The TypeScript comparison keys each site on the token the source wrote, because JavaScript writes several calls on one line and the line alone cannot say which is which. Both sides read the same token out of the same line — until muundo resolves a call through an import that RENAMES it. express() reaches a declaration called createApplication, the report publishes the declaration, and the site then matches nothing on either side: it leaves the scored population altogether.

That reads as a score going up. On express the first CommonJS measurement said 66.7 % where the true figure was 86.2 %, because 654 calls had quietly stopped being counted — including the ones that had just been fixed.

The report now carries wrote, the name the source used, on the edges where resolution replaced it — absent everywhere else, which is almost every edge. The harness keys on that. A reader of a report gets the same thing back: the declaration a call reaches AND the token on the line it was written on.

The share a reader sees

Every table above asks one question: of the calls whose target the compiler says is in this tree, how many does muundo name? That is the right question for the resolver and the wrong one for a person looking at a graph of their codebase. They see every edge the report has, and an edge muundo said nothing about looks exactly like one it got wrong.

scripts/corpus_answered.py <tree> counts the whole report: answered = (in_tree + outside) / all. It is not a quality score — an outside answer can be wrong — but it is what fills or empties a picture.

tree calls in tree outside no answer answered
jagora-worker (Rust) 513 70 415 28 94.5 %
gson (Java) 23 539 10 897 11 056 1 586 93.3 %
ada-util (Ada) 13 538 8 195 1 496 3 847 71.6 %
newtonsoft (C#) 48 269 21 435 15 288 11 546 76.1 %
cobra (Go) 4 430 1 281 1 700 1 449 67.3 %
clap (Rust) 30 309 6 559 13 006 10 744 64.6 %
ky (TS) 10 563 2 071 4 648 3 844 63.6 %
lakisa (TS + Python) 8 320 1 451 3 560 3 309 60.2 %
excalidraw (TSX) 56 209 17 254 15 810 23 145 58.8 %
photos (ArkTS) 17 186 4 881 3 656 8 649 49.7 %
leveldb (C++) 9 951 3 635 1 100 5 216 47.6 %
express (JS) 11 330 977 3 807 6 546 42.2 %
click (Python) 6 869 2 716 2 836 1 317 80.8 %
PyCG corpus (Python) 2 197 605 612 980 55.4 %

This is the number that was hidden. A graph of jagora-worker draws 70 edges from a report of 513, and a reader concludes muundo found almost nothing — while 390 of those calls DO have an answer, outside, and are left out of the picture on purpose because drawing them would invent a node for String::new.

Two lessons in the same table. A projection has to say what it left out and why. And Python, which the corpus tables put at 35 %, is at the bottom here too — the two measures agree about where the work is.

And what the misses ARE

A percentage points at no work. scripts/corpus_misses.py <language> <corpus> reads the same oracle and the same report, keeps only the calls muundo did not name, and files each one under what it is. Three questions, in this order, because each rules out a different repair:

bucket what it means what to fix
no_edge no call edge at that line at all the grammar walk does not recognise the shape
undeclared · … no entity in the tree carries that name the declaration was never extracted
unplaced · bare one segment written a rule about scope and visibility
unplaced · through_one / _two / _many one, two, three or more dots a rule about types, components and chains

The buckets are the same words in every language; only the harness that produces the misses differs. On ada-util the first run read:

bucket count share
unplaced · through_one 104 33 %
unplaced · through_two 75 24 %
unplaced · bare 63 20 %
no_edge 39 12 %
unplaced · through_many 26 8 %
undeclared · … 11 3 %

Grouping those 318 misses by the declaration GNAT says each one reaches named the work: two packages held a fifth of them, and both were generic packages instantiated in a spec and called from its body. Seventeen rules later the count is 42 and Ada reads 98.5 %, every answer right — see the design record for what each rule reads. The buckets moved as the causes were removed:

bucket first run now
unplaced · through_one 104 4
unplaced · through_two 75 13
unplaced · bare 63 6
no_edge 39 11
unplaced · through_many 26 8
undeclared · … 11 0

What is left is what needs a compiler: a generic's formal type standing for a different type at each instantiation, an expression in an aspect on a declaration, and a handful of lines that write two calls where the oracle and muundo cannot be paired by position alone.

Run on clap the same instrument read 48 % unplaced · through_one and 38 % unplaced · bare — a different shape of gap and a different list of rules. Nine rules later Rust reads 84.2 %, and the biggest bucket left is a target the compiler GENERATED: a #[derive(Default)] writes an impl the source never did, so no declaration exists to name. 89 of the 283.

Which rule placed it, and what that rule is worth

resolution says what became of a callee. It does not say how that was decided — and the second question is what tells a reader whether to trust the first. Every edge carries placed_by, the family of rule that placed it, in EVERY language: the field is set in the shared resolver, not in any one language's code.

Each family's precision is measured on the same oracles as everything else. scripts/corpus_rules.py <language> <corpus> does a whole language in one pass — the oracle is built once and muundo runs once, only the comparison is repeated. A harness that keeps its own tally takes --only <family> instead (corpus_ada.py, corpus_go.py, corpus_arkts.py).

Each cell is answers · precision, counting ALL of a family's answers, in_tree and outside together, because some families only ever give one of the two.

Rule Java C# TSX TS ArkTS Rust C++ Go Kotlin Ada
callers_file 2 099 · 1.000 3 205 · 0.997 4 800 · 0.9996 273 · 0.960 69 · 1.000 60 · 0.983 375 · 0.979 76 · 1.000 — 710 · 1.000
import 1 369 · 1.000 3 691 · 0.997 7 546 · 0.988 638 · 1.000 1 018 · 1.000 224 · 0.987 — 234 · 1.000 185 · 1.000 —
written_type 5 068 · 0.9998 3 554 · 0.957 1 159 · 0.999 27 · 1.000 357 · 1.000 588 · 0.968 13 · 0.769 311 · 1.000 26 · 1.000 12 · 1.000
nothing_declares_it 2 · 0.000 1 112 · 0.989 3 744 · 0.998 3 643 · 1.000 345 · 1.000 169 · 1.000 32 · 0.906 110 · 1.000 109 · 1.000 —
visible_unique 1 019 · 1.000 835 · 0.999 — — — 8 · 1.000 96 · 0.906 74 · 1.000 47 · 1.000 44 · 1.000
self_receiver 46 · 1.000 58 · 1.000 845 · 1.000 89 · 1.000 457 · 1.000 500 · 0.992 4 · 0.750 — — —
inheritance 73 · 0.986 773 · 0.987 10 · 1.000 — 60 · 1.000 8 · 1.000 — — — 83 · 1.000
return_type 930 · 1.000 19 · 0.947 6 · 1.000 — 74 · 1.000 494 · 0.966 — 19 · 1.000 20 · 1.000 —
package — 6 · 1.000 — — — 8 · 1.000 18 · 0.944 — — 1 322 · 1.000
local_scope 3 · 1.000 34 · 1.000 72 · 1.000 — 3 · 1.000 45 · 0.956 — — — 3 · 1.000
bare_name — — — — — — — — — —

Over all ten corpora: 55 158 answers, 385 wrong, 0.9930. By family:

Rule answers wrong precision
package 1 354 1 0.9993
self_receiver 1 999 5 0.9975
nothing_declares_it 9 266 24 0.9974
callers_file 11 667 32 0.9973
visible_unique 2 123 10 0.9953
import 14 905 105 0.9930
inheritance 1 007 11 0.9891
return_type 1 562 18 0.9885
local_scope 160 2 0.9875
written_type 11 115 177 0.9841

return_type answers three times as often as it did — 494 of those are Rust's, where it had 44 — and its precision fell from 0.9954 to 0.9885 in the process. That is the trade the Rust rules made: a rule used eleven times as much, wrong four times more often, for 427 further calls placed.

bare_name answers nothing here any more. It was Ada's parameterless calls, and every one of them is now placed by a rule that reads the source — the package a name belongs to, or the type a receiver is written as.

What this is for. A consumer who needs certainty keeps the families measured at 1.000 on its language and treats the rest as leads. Before the field, in_tree read off an import the file writes and in_tree chosen because a name was unique were the same word.

Two cells read 0.000 and neither is a rule that fails. A family's precision has to be read over the population it answers on. nothing_declares_it only ever says outside, so against javac and GNAT — which record only in-tree calls — every one of its answers is judged against a target the oracle says is inside. The same family is 1.000 on ky, TSX, ArkTS, Go, Kotlin and Rust. Ada's import row is the same shape: seven answers, all of them judged against an in-tree target.

This table was wrong twice before it was right, and both mistakes are worth knowing. Go read 0.000 everywhere because its oracle compares by CLASS and was being run through the by-position comparison. And families that answer nothing read 8 · 1.000 on Rust and 5 · 1.000 on C++, because a site where the compiler itself says undetermined matches muundo's undetermined whatever rule did or did not apply — a trivial agreement no family can claim. Both are excluded now.

One row was a corpus and not a language, and then it was a defect. ky read 31.6 % against excalidraw's 86.0 %, measured by the same checker and the same script, and the difference was written off as ky being a two-thousand-line library built on callbacks. Most of it was one shape muundo could not read: const ky = createInstance(); is what every consumer of that module calls into, and a module-level constant bound to a call was not a declaration at all. ky now reads 64.1 %. The gap that remains — a value holding a function, an option object whose members are callables — is the part that really is the corpus.

The same call in every language: which rules are missing

A corpus says how much is missed. It does not say what. scripts/corpus_shapes.py asks the other question: a call graph has a small number of SHAPES — a receiver whose type the source writes, a chain through a declared return type, a member declared on a base type, an overload told apart by what receives its result — and each exists in every language, spelled differently in each. Where muundo places a shape in one language and not in another, the difference is never the language: it is a rule read in one grammar and not in the other.

The matrix names the rule that placed each cell, or MISS:

shape Ada C++ C# Go Java Python Rust TS
written_type ok ok ok ok ok ok ok ok
return_chain ok ok ok ok ok ok ok ok
base_member ok ok ok — ok ok — ok
result_type ok — — — ok — — —
interface_member ok — ok ok ok — ok ok
field_type ok — ok ok ok — ok ok
named_arguments ok — — — ok ok — —
aliased_namespace ok — ok — ok ok ok —

Every cell it can ask about is filled. Ten gaps on its first run, all ten closed:

  • Go's struct fields were read as identifier where the grammar writes field_identifier, so type Holder struct{ part Inner } recorded nothing.
  • Go's interface methods were not entities at all, so s.area() on a Shape named something muundo had never heard of — the same gap Rust's trait methods and Ada's spec subprograms had.
  • import p.deep.Deep; then Deep.work() built p.deep.Deep::work, a name no entity carries, and called it outside.
  • using p.deep; then Deep.work() looked for the member beside the namespace instead of inside the type. 1 962 of Newtonsoft.Json's calls.
  • An Ada call written without parentheses lost its prefix. C.Describe is how Ada calls a parameterless function, and only the leaf survived — so the type of C was thrown away before anything could use it, and H.Part.Deep, with three segments, produced no edge at all.
  • An inherited Ada primitive needs no with. A file that writes with Child; calls C.Describe without ever naming Parent, and the visibility check refused it.
  • C++ emitted no inheritance edge. struct Child : Parent produced no Extends dependency, so the rule every other language has had nothing to walk.

Three of those were in the matrix before they were in any corpus. The Ada pair is worth 16 calls and three fewer wrong answers on ada-util — base_member and field_type looked like one gap each and were one cause.

And with --lsp, the same corpus and the same score

The same command with a server named. Measured 2026-09-24 in this devcontainer; the time is the whole run, muundo included.

Language server coverage precision was time
Go gopls 0.16.2 100 % 1.000 80.8 % 2 min
C clangd 100 % 1.000 100 % seconds
Java jdtls (JDK 25) 98.9 % 0.9995 95.8 % 15 min
Rust rust-analyzer 92.3 % 0.953 84.2 % 1 min
ArkTS typescript-language-server 91.6 % 1.000 71.2 % 3 min
Kotlin kotlin-language-server 90.7 % 1.000 41.5 % 1 min
TSX typescript-language-server 90.1 % 0.993 86.0 % 5 min
C++ clangd 75.2 % 0.911 50.1 % 9 min
TypeScript typescript-language-server 67.6 % 0.995 64.1 % 1 min
Ada ada_language_server 99.2 % 1.000 98.5 % 3 min
JavaScript typescript-language-server 34.6 % 1.000 86.2 % 1 min
Python pyright 67.2 % 0.787 27.4 % 1 min

C++'s row moved with the oracle, and only C++'s. clangd read 79.9 % at precision 0.897 against the oracle that missed every call that can throw; it reads 75.2 % at 0.911 against the one that sees them. Rust's and C's rows are unchanged, because a Rust module compiles whole and cJSON's calls do not throw.

Two are not in the table, and neither is a muundo result:

  • C# — OmniSharp 2.0.0 starts and answers initialize, then places 0 before the ask budget runs out: the project has not opened. The provenance says so in the server's own words.
  • Python — corpus_calls.py reads a reviewed corpus of hard cases and has no --lsp flag. The click row is scored by corpus_python.py instead, which compares against the RUN. Its --lsp pyright row is stale: re-run after the two rules below, pyright added exactly nothing, digit for digit, which is what a server that failed to start looks like. Nobody has confirmed it ran.

The Python oracle RUNS the program. py_runtime_calls.py is a measuring instrument for this repository's own development, pointed at corpora chosen and read beforehand, inside a container. muundo analyze executes nothing, ever, in any language — that is the property the engine is built on, and nothing here is wired into it.

Python is measured twice, and the two say different things. The PyCG corpus is a catalogue of hard cases by construction — a fair test of the limit and an unfair one of ordinary code. scripts/py_runtime_calls.py records what the interpreter actually called while a project's own test suite ran, and scripts/corpus_python.py compares muundo against that. It is the strongest evidence there is for a call that HAPPENED, and it is silent about branches the run never took, so only the sites it recorded are judged.

It also sees what no reader of the source can: a method found through __getattr__, an object a test monkeypatched. Those count against muundo, and they are the honest reason pyright's own answers disagree with the run 501 times — a type checker names what is declared, a run names what happened.

jdtls needs JDK 25. Started on 21 it fails initialize with an OSGi Require-Capability error, and muundo then reports unavailable and changes not one edge — which is the right behaviour and looks exactly like a server that added nothing.

Change evidence is recall / precision over the declarations git says changed: 384 commits and 1 265 declarations in all, with nothing excluded from the count. Precision on call edges means an edge called in_tree really is one; honesty means that of the calls it did not place, it said it could not place them rather than placing them wrongly. 16 832 probes across 22 projects; every law held.

All thirteen have a compiler to answer for them. Two answer through what they WRITE rather than through an API: kotlinc keeps a line-number table and a source name in every class, which javap -c -l -p prints, and GNAT writes its cross-references into an .ali file, where a reference marked s is a subprogram call. ArkTS is measured on the files the TypeScript checker can READ: its declarative ArkUI body — build() { Column() { Text(this.name) } } — is not TypeScript at all, so the tree is handed over with .ets renamed to .ts and only the files that parse with no syntax error are read back. In a HarmonyOS application that is the models, the services and the utilities; the screens are covered by the probes. Each oracle found defects on its first run — which is the whole reason for building one — and the largest was shared by seven languages: new Thing() produced no call edge at all, in every language that spells construction with a keyword.

YAML is supported and is not in this table: it has no declarations and no calls, so none of the three instruments has a question to ask of it.

What a site is, and where that is not enough

A call site is a FILE AND A LINE, which is all debug info gives. A line can write several calls, and where both sides have answers left over that disagree among themselves, no rule pairs them except by order — and order is not evidence. Those are counted apart as ambiguous and judged neither way: 96 of fmt's sites, 57 of clap's, and 0 to 6 everywhere else. The harnesses that read the source through a compiler's own API — TypeScript, Java, C#, Kotlin, ArkTS — put the NAME the source writes in the key as well, and so almost never reach it. Both sides read that name off the same line; it is not a name matched between two vocabularies.

The limit, and its size

muundo does not track what a name holds. Two consequences, both measured:

  • Overload resolution. C++ and Rust pick between same-named declarations by argument type. What is left of it after the ambiguous sites are set aside is 27 of fmt's 1 819 compiled call sites and 8 of clap's 3 569, counted against muundo rather than excused.
  • Value flow. Against the PyCG corpus, muundo resolves 33 % of the edges a value-flow analysis finds. 87 % of what it misses needs to know what a name holds — a parameter called as a function, a method on a receiver whose type is not tracked.

Asking a compiler instead

--lsp <command> asks a language server about the calls reading the source cannot place. On anyhow, against rustc's own debug info, the call sites muundo places correctly go from 59 to 119 of 120 — in 29 seconds, where the analysis alone takes under one.

Every language muundo speaks has a server it can drive. One corpus each, counting the calls a server placed that reading the source had not:

Language Server Placed
C clangd 68 on sds
C++ clangd 1 835 on fmt
Rust rust-analyzer 815 on anyhow — undetermined 720 → 34
Python pyright 3 221 on click
TypeScript typescript-language-server 1 253 on ky
TSX typescript-language-server 2 020 on excalidraw
ArkTS typescript-language-server 783 on OpenHarmony photos
JavaScript typescript-language-server 1 120 on axios
Go gopls 573 on uuid — undetermined 505 → 7
Java jdtls 590 on gson
Ada ada_language_server 14 on ada-util
Kotlin kotlin-language-server answers; okio's Gradle model had not opened after 20 minutes
C# OmniSharp answers; Newtonsoft.Json had not opened after 20 minutes

The counts are not comparable across rows: each ran under its own --lsp-budget, and what a server can add depends on how much the language lets muundo place by itself.

The command is the whole configuration, arguments included, because that is how a server is told where the project is — --lsp "ada_language_server --config als.json". Ada places 14 and not 0 because of that file: without it the server opens no project. And it costs what muundo otherwise does not: a project the server can open. Where it cannot be started, the provenance says so in the server's own words and not one edge changes.

It also changes the snapshot id, because the server's identity joins the hashed inputs — two runs under different toolchains are two analyses and must not share an id. A report made this way cannot be replayed by verify on a machine without the same server, and the verdict says resolver_absent rather than comparing against a weaker analysis. Every edge a server placed carries resolved_by.

Neither limit is a defect to fix within this design. undetermined exists to say so per edge, and it is the honest majority: across six trees and six languages, in_tree is 8–19 % of call edges, outside 0–25 %, and undetermined the rest.

Reproduce

scripts/corpus_change.py <repo> --commits 25      # any git repository
scripts/corpus_c.py      <tree>                   # C, against clang
scripts/corpus_cpp.py    <tree> --include <dir>   # C++, against clang
scripts/corpus_rust.py   <crate>                  # Rust, against rustc
scripts/corpus_go.py     <module>                 # Go, against x/tools callgraph
scripts/corpus_ts.py     <tree>                   # TS/JS, against the tsc checker
scripts/corpus_java.py   <tree>                   # Java, against javac
scripts/corpus_csharp.py <tree>                   # C#, against Roslyn
scripts/corpus_arkts.py  <tree>                   # ArkTS, against tsc
scripts/corpus_kotlin.py <tree> --classes <dir>   # Kotlin, against its bytecode
scripts/corpus_ada.py    <tree>                   # Ada, against GNAT's cross-references
scripts/corpus_calls.py  PyCG/micro-benchmark/snippets
scripts/corpus_probe.py  <tree> --lang <language> [--kind member]

What these numbers do not say

A 1.000 over 25 commits of one repository means muundo named the right declarations there. It is enough to find defects — it found eleven — and it is not a rate to quote about your code.

A held probe means the machinery answers when a declaration is placed where the language makes it visible. Only a compiler oracle says a project's own calls are all resolved. All thirteen now have one.