What has been measured¶
Every number here was produced by a script in scripts/, against a repository
anybody can clone. Nothing is an estimate, and the commands are given so you can
disagree with them.
Three instruments, three questions¶
| The question | How | |
|---|---|---|
| Change evidence | Did it name the declarations a commit touched? | git says which lines changed; muundo diff is checked against it |
| Call edges | Is an edge the edge the compiler compiled? | clang or rustc, compared by file and line so no name has to be matched |
| Probes | Does it see a declaration placed where the language makes it visible? | the source is twisted in a way chosen in advance, and only muundo's own two reports are read |
The third needs no second tool, which is why it reaches every language.
Per language¶
| Language | Change evidence, 25 commits | Call edges | Probes |
|---|---|---|---|
| C | 3 repositories, 1.000 / 1.000 | clang, 990 sites, precision 1.000 | 2 188 |
| C++ | fmt, 1.000 / 1.000 | clang, 1 819 sites, precision 0.968, 582 placed | 665 |
| Rust | 3 crates, 1.000 / 1.000 | rustc, 3 569 sites, precision 0.971, honesty 0.957, 1 979 placed | 1 692 |
| Python | click, 1.000 / 1.000 | PyCG corpus, precision 0.989, honesty 0.915 | 720 |
| TypeScript | ky, 1.000 / 1.000 | tsc, 10 536 sites, precision 0.995, 2 042 placed | 664 |
| JavaScript | express, 1.000 / 1.000 | tsc, 14 490 sites, precision 0.999, 964 placed | 1 367 |
| TSX | excalidraw, 1.000 / 1.000 | tsc, 56 070 sites, precision 0.993, honesty 0.983, 18 085 placed | 1 621 |
| ArkTS | photos, 1.000 / 1.000 | tsc, 5 141 sites, precision 1.000, honesty 1.000, 2 383 placed | 598 |
| Go | uuid, 1.000 / 1.000 | callgraph on cobra, 817 sites, precision 1.000, honesty 1.000, 780 placed | 2 271 |
| Java | gson, 1.000 / 1.000 | javac, 23 291 sites, precision 1.000, honesty 0.997, 10 357 placed | 1 482 |
| C# | Newtonsoft.Json, 1.000 / 1.000 | Roslyn, 42 454 sites, precision 0.985, honesty 0.986, 13 113 placed | 1 275 |
| Kotlin | okio, 1.000 / 1.000 | klaxon's own bytecode, 848 sites, precision 1.000, honesty 1.000, 388 placed | 1 573 |
| Ada | ada-util, 1.000 / 1.000 | GNAT .ali, 2 130 sites, precision 1.000, honesty 1.000, 2 088 placed |
702 |
Of the calls whose target really IS in the tree, how many are named¶
One script produces every row — scripts/corpus_score.py <language> <corpus>
<muundo> [<server>] — because comparing per-language buckets by hand is how a
report comes out wrong. Each harness counts calls whose target is a LIBRARY
alongside calls that land in the tree, and reading the total as "what muundo
found" overstates the gap by half.
coverage is named / (named + missed + wrong): of the edges that end
inside the tree, the share muundo names at the right declaration. left is
counted apart — a call whose target is outside the tree and that muundo said
so about, which answers a different question.
What is left, and what is a wall¶
A coverage figure says how much is missed. It does not say whether the rest is
work. scripts/corpus_ceiling.py <language> <corpus> asks one question of every
miss: does muundo declare anything at the position the oracle names? Where it
declares nothing, there was nothing muundo could have answered.
| corpus | named | reachable misses | outside the model | ceiling |
|---|---|---|---|---|
| excalidraw (TSX) | 91.9 % | 762 | 449 | 96.5 % |
| ky (TypeScript) | 74.1 % | 421 | 416 | 87.0 % |
| express (JavaScript) | 86.3 % | 91 | 60 | 94.5 % |
| fmt (C++) | 28.8 % observed | 1 858 | 130 | 94.9 % |
Half of what ky still misses is in the second column. The shape is always the same:
new Promise((resolve) => { … resolve(value) … })
const Footer = ({ onChange }: { onChange: () => void }) => { … onChange() … }
resolve and onChange are PARAMETERS. The TypeScript compiler resolves the call
to the parameter's own declaration; muundo's model carries entities for
declarations — functions, classes, members — and a parameter holding a callback is
not one. A report cannot name what it does not carry, so closing that gap means
deciding that a parameter is a declaration, which is a different product and a
different set of numbers everywhere else.
Reading the source alone, with no compiler and no build:
| Language | corpus | the right overload | the right method | named | overload-only | precision |
|---|---|---|---|---|---|---|
| C | cJSON | 100 % | 100 % | 306 | 0 | 1.000 |
| ArkTS | photos | 92.7 % | 92.7 % | 2 224 | 0 | 1.000 |
| TSX | excalidraw | 91.9 % | 92.5 % | 15 347 | 91 | 0.993 |
| JavaScript | express | 86.3 % | 86.3 % | 967 | 0 | 0.999 |
| Go | cobra | 82.8 % | 82.8 % | 453 | 0 | 1.000 |
| Rust | clap | 78.9 % | 81.0 % | 1 799 | 47 | 0.975 |
| TypeScript | ky | 74.1 % | 74.2 % | 2 419 | 3 | 0.994 |
| Java | gson | 73.0 % | 96.0 % | 8 179 | 2 577 | 0.9998 |
| Kotlin | klaxon | 68.1 % | 72.2 % | 515 | 31 | 0.992 |
| C# | Newtonsoft.Json | 66.3 % | 94.3 % | 12 851 | 5 414 | 0.986 |
| Python | PyCG corpus | 63.5 % | 63.5 % | 160 | 0 | 0.994 |
| Kotlin | KotlinPoet | 45.9 % | 75.9 % | 383 | 251 | 0.955 |
| Ada | ada-util | 98.0 % observed | 98.0 % | 2 088 | 0 | 1.000 |
| C++ | fmt | 28.8 % observed | 42.4 % | 826 | 390 | 0.987 |
| C++ | leveldb | 51.3 % observed | 51.3 % | 533 | 0 | 0.945 |
Two columns because there are two questions. The first asks whether the call
reaches the right DECLARATION — which overload of SerializeObject, of eight. The
second asks whether it reaches the right METHOD, which is what a reader of a call
graph follows. muundo's qualified names do not distinguish overloads, so a call it
places perfectly still cannot say which one: Src/…/JsonConvert.cs::JsonConvert::SerializeObject
is eight declarations.
The gap between the columns IS that, and it is the largest single lever left: 5 077 of C#'s 6 371 misses and 2 577 of Java's 3 015 are calls placed on the right method whose overload is unnamed. Kotlin and C++ have it too (251 and 379). Languages without overloading — C, Go, Python, ArkTS — have two identical columns.
The last two rows publish no headline: 30 of ada-util's 222 units and 29 of fmt's 47 do not compile, so the oracle saw part of the tree.
These are protocol v2 and are NOT comparable to anything published before 2026-09-27. What changed is the measurement, not the engine — see below.
The last two publish no headline: 30 of ada-util's 222 units and 29 of fmt's 47
do not compile, so the oracle saw part of the tree and a rate over part of a tree
is not that tree's rate. observed_coverage carries the number with
oracle_status: partial beside it.
These are protocol v2 and are NOT comparable to anything published before 2026-09-27. What changed is the measurement, not the engine — see below.
A compiler calls what the author did not write¶
Status::Corruption("msg") takes a const Slice&, so passing a string builds a
Slice — and clang records a call to Slice::Slice on a line that writes
Slice nowhere. A Rust #[derive] is the same shape, and so is every C++ copy
constructor and destructor: real calls in the binary, and no call a reader of
that line could point at.
muundo reads source. An answer it cannot see in the source is not a miss, and
counting it as one measures the distance between two definitions of the word
"call". Those expectations are counted apart now, under not_written, and
the number excluded is published beside the score:
| corpus | judged | excluded as not written |
|---|---|---|
| leveldb | 1 936 | 545 |
| fmt | 1 016 | 230 |
Only a CONSTRUCTOR or a TYPE is excluded, and only where the line does not
name it. That is the whole phenomenon: a temporary built to convert an
argument, a copy, a destructor. s.ok() stays judged because ok is a method;
new leveldb_t stays judged because leveldb_t is right there.
A first version of this rule was wrong and the numbers it produced were published for an hour. It excluded any expectation whose declaration the line did not name — which is exactly what a renaming import does, so it dropped 655 of express's calls and read 66.8 % where the truth is 86.2 %. The same blind spot as keying a site on the resolved name, in a new guise. The figures in this document are the narrow rule's.
Two bugs in the C++ oracle, and both hid the same thing¶
C++ read 51.4 % on fmt, and a second corpus was added to see whether that was the language or the project. leveldb — ordinary object-oriented C++, the kind a regulated codebase is written in — read 28.6 % with precision 0.586. A precision that bad is either a serious defect or a broken instrument, and it was the instrument, twice.
One translation unit at a time. db/c.cc calls NewBloomFilterPolicy,
which util/bloom.cc defines — and in db/c.cc's own module that is a
declare and nothing more. The oracle read each unit alone, so every call
crossing a .cc boundary was scored as leaving the tree, and muundo was right
about every one of them. The tree is all of its units: the harness now gathers
what they all define before judging any of them. 123 of 227 disagreements.
An invoke writes its !dbg on the next line. LLVM spells a call that can
throw as invoke … to label … unwind label …, !dbg !N, and the pattern reading
the IR could not cross the newline. Every call that can throw was invisible to
this oracle — which in ordinary C++ is most of them. The oracle saw 650 more
calls on leveldb once it could read them, and 158 more on fmt.
Neither was visible on fmt alone: it is header-only, so nothing crosses a unit,
and its hot paths are noexcept. fmt's published figure was 51.4 % measured
against an oracle missing 158 calls; it is 50.1 % against one that sees them.
leveldb now reads 38.7 % with precision 0.865. What is left is named in the design record: a construction of a plain struct, which muundo reports as a dependency by design and the compiler runs no code for, and a method reached through a member whose type a header declares.
A <builtin> was counted as an in-tree expectation¶
PyCG's answer files list nine edges whose target is <builtin>.len or
<builtin>.map. Those name no node of the tree — len is the interpreter's —
so muundo cannot reach them with an in-tree edge, and outside is the right
answer rather than a miss. expected_edges drops them now, which is worth 2.4
points of the Python row and is a correction to the instrument, not to the
engine.
Protocol v2 — what the old numbers were counting¶
Every figure above moved on 2026-09-27, most of them down, and no line of the engine changed that day. Two flaws in the instrument:
A site muundo said nothing about left the score. The rate divided by
named + honest + wrong, and a position where muundo produced no edge at all
landed in missing, which was in none of those. So removing an edge could take
the site out of the denominator and LIFT the percentage. That is not a
hypothetical: when casts stopped being calls, fmt's population fell from 1 079 to
1 060 while named stayed at 614, and the rate rose from 56.90 % to 57.92 % with
not one additional call named. The denominator is now every expectation the oracle
can judge, which does not move when the answers do —
scripts/test_corpus_protocol.py asserts it: the same comparison with four
answers instead of ten keeps the same denominator and scores 0.4.
An external call paid for an internal one. The class-based adapters — Go, C, Ada, ArkTS — counted agreement about a call that LEAVES the tree in the same total as one that lands in it. Reading that as "what muundo found" overstated it: cobra's 855 agreements were 453 internal and 402 external, and internal recall is 453 of 547.
Two smaller ones came with it. Precision now counts only answers that DREW an
in-tree edge, because saying outside about an internal call draws nothing — it
costs recall and cannot cost precision. And the Kotlin oracle reads javap -s
and keys on (owner, name, descriptor), so two overloads are two identities: they
used to collapse under one name, which let an answer naming the wrong overload
count as right.
The numbers are lower and they answer the question the table claims to answer. A row also carries the digest of the expectations it was compared against and of the sources they were read from, so it can be re-derived rather than believed.
And they are a LOWER bound, for one reason worth knowing. v1 excluded an
expectation whose target the call's own line does not name — a temporary built to
convert an argument, a copy constructor, a #[derive]: real calls in the binary
that no reader of that line could point at. v2 stopped, because the exclusion was
computed from muundo's OWN report, so the set of excluded expectations moved as the
engine changed. That is the one thing a denominator must not do, so dropping it was
right — but the phenomenon is real, and those calls are now counted as misses. An
engine-independent version would ask the ORACLE for the target's name rather than
asking muundo which names it recorded at that position. what_the_source_wrote is
kept, out of use, with that written on it.
C++'s figure is bounded by PARSING, not by resolution¶
fmt is 750 syntax errors across 39 files. leveldb is 116 across 28 — six times fewer for a tree of similar size. That is the whole distance between them, and it is why a language whose C sibling scores 100 % sits at 57.9 %.
What that costs, concretely: class buffer in core.h is recorded as spanning
lines 1759 to 1773 because the parser stopped at the first constructor it
could not read. push_back is at line 1829, so it is not an entity at all — and
no call to it can resolve, however good the resolver is. muundo says so in every
report: partial_analysis names each file, its error count, and that "code under
those regions was NOT read".
Three macro shapes were tried on a rewritten copy, each length-preserving so no
position moved: the namespace-opening macro (FMT_BEGIN_NAMESPACE), the SFINAE
macro inside a template parameter list (FMT_ENABLE_IF), and the pragma macro at
namespace scope (FMT_PRAGMA_GCC). Masking all three took fmt from 750 errors to
658. The remainder is spread across the grammar's handling of template
metaprogramming, not concentrated in a construct that could be masked.
So the honest statement about C++ is two numbers, not one: what the resolver does with the code it can read, and how much of the code it can read. fmt is the extreme of the second.
leveldb's oracle shrank, so its row is gone¶
scripts/corpus_cpp.py compiles one translation unit at a time and skips the
ones that do not build. On 2026-09-26 it builds 29 of leveldb's 76 and sees
1 229 call sites; the published 65.8 % was measured when it saw 1 936. Three
different muundo binaries — before this work, after it, and the commit in
between — all score 69.4 % against the smaller oracle, so nothing in muundo
moved. A number measured over 38 % of a project is not that project's score,
and fmt's row is the C++ figure to read until leveldb builds again.
Kotlin's row was one small corpus¶
klaxon judges 238 calls. That is a hundred-file JSON library, and 45.0 % of 238 is not a language's score — it is one project's shape, measured once.
KotlinPoet doubles the population and answers 76.0 % over 491 in-tree calls: builders, sealed types, extension functions and a DSL, which is what modern Kotlin looks like. Both rows stay, because the pair says more than either.
Two rules moved that row from 58.0 %, and they are the two shapes Kotlin code is
written in. A val b = Thing.builder() says what b holds, so the calls on b
reach the nested Builder — that is the same value_bindings fact Rust and
TypeScript already filled, now filled by the Kotlin parser too. And a call on a
receiver whose name is declared exactly once at FILE scope is an extension
function, xs.toImmutableList(), which is a member of nothing.
klaxon gained less (41.5 → 45.0 %) and lost precision (1.000 → 0.947): six of its calls now name the wrong target, all of them a basename declared once at file scope in a tree that also declares it as a member elsewhere. The rule is right about the shape and wrong about those six, and the trade is published rather than hidden.
Building it takes one command, and the corpus is the compiled sources:
kotlinc -Xjvm-default=all -cp kotlin-reflect.jar \
-d /var/tmp/muundo/kotlinpoet-classes \
$(find …/kotlinpoet/kotlinpoet/src/commonMain -name '*.kt')
MUUNDO_KOTLIN_CLASSES=/var/tmp/muundo/kotlinpoet-classes \
scripts/corpus_score.py kotlin …/kotlinpoet/kotlinpoet/src/commonMain <muundo>
Two bigger corpora were tried and refused. okio is multiplatform, and
kotlinc rejects an expect and its actual in one module — a multiplatform
build is the only way to compile it, and this oracle reads what one kotlinc
run wrote. OkHttp 4.12 compiles down to thirteen errors with five compile-only
jars on the classpath (okio, animal-sniffer, Conscrypt, BouncyCastle, OpenJSSE,
an Android stub), and the last of them need Android APIs no published stub jar
carries. Neither is a muundo limit; both are what the corpus costs.
One flaw of this instrument, and what it cost¶
The TypeScript comparison keys each site on the token the source wrote,
because JavaScript writes several calls on one line and the line alone cannot
say which is which. Both sides read the same token out of the same line — until
muundo resolves a call through an import that RENAMES it. express() reaches a
declaration called createApplication, the report publishes the declaration,
and the site then matches nothing on either side: it leaves the scored
population altogether.
That reads as a score going up. On express the first CommonJS measurement said 66.7 % where the true figure was 86.2 %, because 654 calls had quietly stopped being counted — including the ones that had just been fixed.
The report now carries wrote, the name the source used, on the edges where
resolution replaced it — absent everywhere else, which is almost every edge.
The harness keys on that. A reader of a report gets the same thing back: the
declaration a call reaches AND the token on the line it was written on.
The share a reader sees¶
Every table above asks one question: of the calls whose target the compiler says is in this tree, how many does muundo name? That is the right question for the resolver and the wrong one for a person looking at a graph of their codebase. They see every edge the report has, and an edge muundo said nothing about looks exactly like one it got wrong.
scripts/corpus_answered.py <tree> counts the whole report:
answered = (in_tree + outside) / all. It is not a quality score — an outside
answer can be wrong — but it is what fills or empties a picture.
| tree | calls | in tree | outside | no answer | answered |
|---|---|---|---|---|---|
| jagora-worker (Rust) | 513 | 70 | 415 | 28 | 94.5 % |
| gson (Java) | 23 539 | 10 897 | 11 056 | 1 586 | 93.3 % |
| ada-util (Ada) | 13 538 | 8 195 | 1 496 | 3 847 | 71.6 % |
| newtonsoft (C#) | 48 269 | 21 435 | 15 288 | 11 546 | 76.1 % |
| cobra (Go) | 4 430 | 1 281 | 1 700 | 1 449 | 67.3 % |
| clap (Rust) | 30 309 | 6 559 | 13 006 | 10 744 | 64.6 % |
| ky (TS) | 10 563 | 2 071 | 4 648 | 3 844 | 63.6 % |
| lakisa (TS + Python) | 8 320 | 1 451 | 3 560 | 3 309 | 60.2 % |
| excalidraw (TSX) | 56 209 | 17 254 | 15 810 | 23 145 | 58.8 % |
| photos (ArkTS) | 17 186 | 4 881 | 3 656 | 8 649 | 49.7 % |
| leveldb (C++) | 9 951 | 3 635 | 1 100 | 5 216 | 47.6 % |
| express (JS) | 11 330 | 977 | 3 807 | 6 546 | 42.2 % |
| click (Python) | 6 869 | 2 716 | 2 836 | 1 317 | 80.8 % |
| PyCG corpus (Python) | 2 197 | 605 | 612 | 980 | 55.4 % |
This is the number that was hidden. A graph of jagora-worker draws 70 edges
from a report of 513, and a reader concludes muundo found almost nothing — while
390 of those calls DO have an answer, outside, and are left out of the picture
on purpose because drawing them would invent a node for String::new.
Two lessons in the same table. A projection has to say what it left out and why. And Python, which the corpus tables put at 35 %, is at the bottom here too — the two measures agree about where the work is.
And what the misses ARE¶
A percentage points at no work. scripts/corpus_misses.py <language> <corpus>
reads the same oracle and the same report, keeps only the calls muundo did not
name, and files each one under what it is. Three questions, in this order,
because each rules out a different repair:
| bucket | what it means | what to fix |
|---|---|---|
no_edge |
no call edge at that line at all | the grammar walk does not recognise the shape |
undeclared · … |
no entity in the tree carries that name | the declaration was never extracted |
unplaced · bare |
one segment written | a rule about scope and visibility |
unplaced · through_one / _two / _many |
one, two, three or more dots | a rule about types, components and chains |
The buckets are the same words in every language; only the harness that produces the misses differs. On ada-util the first run read:
| bucket | count | share |
|---|---|---|
unplaced · through_one |
104 | 33 % |
unplaced · through_two |
75 | 24 % |
unplaced · bare |
63 | 20 % |
no_edge |
39 | 12 % |
unplaced · through_many |
26 | 8 % |
undeclared · … |
11 | 3 % |
Grouping those 318 misses by the declaration GNAT says each one reaches named the work: two packages held a fifth of them, and both were generic packages instantiated in a spec and called from its body. Seventeen rules later the count is 42 and Ada reads 98.5 %, every answer right — see the design record for what each rule reads. The buckets moved as the causes were removed:
| bucket | first run | now |
|---|---|---|
unplaced · through_one |
104 | 4 |
unplaced · through_two |
75 | 13 |
unplaced · bare |
63 | 6 |
no_edge |
39 | 11 |
unplaced · through_many |
26 | 8 |
undeclared · … |
11 | 0 |
What is left is what needs a compiler: a generic's formal type standing for a different type at each instantiation, an expression in an aspect on a declaration, and a handful of lines that write two calls where the oracle and muundo cannot be paired by position alone.
Run on clap the same instrument read 48 % unplaced · through_one and 38 %
unplaced · bare — a different shape of gap and a different list of rules.
Nine rules later Rust reads 84.2 %, and the biggest bucket left is a target
the compiler GENERATED: a #[derive(Default)] writes an impl the source
never did, so no declaration exists to name. 89 of the 283.
Which rule placed it, and what that rule is worth¶
resolution says what became of a callee. It does not say how that was
decided — and the second question is what tells a reader whether to trust the
first. Every edge carries placed_by, the family of rule that placed it, in
EVERY language: the field is set in the shared resolver, not in any one
language's code.
Each family's precision is measured on the same oracles as everything else.
scripts/corpus_rules.py <language> <corpus> does a whole language in one
pass — the oracle is built once and muundo runs once, only the comparison is
repeated. A harness that keeps its own tally takes --only <family> instead
(corpus_ada.py, corpus_go.py, corpus_arkts.py).
Each cell is answers · precision, counting ALL of a family's answers,
in_tree and outside together, because some families only ever give one of
the two.
| Rule | Java | C# | TSX | TS | ArkTS | Rust | C++ | Go | Kotlin | Ada |
|---|---|---|---|---|---|---|---|---|---|---|
callers_file |
2 099 · 1.000 | 3 205 · 0.997 | 4 800 · 0.9996 | 273 · 0.960 | 69 · 1.000 | 60 · 0.983 | 375 · 0.979 | 76 · 1.000 | — | 710 · 1.000 |
import |
1 369 · 1.000 | 3 691 · 0.997 | 7 546 · 0.988 | 638 · 1.000 | 1 018 · 1.000 | 224 · 0.987 | — | 234 · 1.000 | 185 · 1.000 | — |
written_type |
5 068 · 0.9998 | 3 554 · 0.957 | 1 159 · 0.999 | 27 · 1.000 | 357 · 1.000 | 588 · 0.968 | 13 · 0.769 | 311 · 1.000 | 26 · 1.000 | 12 · 1.000 |
nothing_declares_it |
2 · 0.000 | 1 112 · 0.989 | 3 744 · 0.998 | 3 643 · 1.000 | 345 · 1.000 | 169 · 1.000 | 32 · 0.906 | 110 · 1.000 | 109 · 1.000 | — |
visible_unique |
1 019 · 1.000 | 835 · 0.999 | — | — | — | 8 · 1.000 | 96 · 0.906 | 74 · 1.000 | 47 · 1.000 | 44 · 1.000 |
self_receiver |
46 · 1.000 | 58 · 1.000 | 845 · 1.000 | 89 · 1.000 | 457 · 1.000 | 500 · 0.992 | 4 · 0.750 | — | — | — |
inheritance |
73 · 0.986 | 773 · 0.987 | 10 · 1.000 | — | 60 · 1.000 | 8 · 1.000 | — | — | — | 83 · 1.000 |
return_type |
930 · 1.000 | 19 · 0.947 | 6 · 1.000 | — | 74 · 1.000 | 494 · 0.966 | — | 19 · 1.000 | 20 · 1.000 | — |
package |
— | 6 · 1.000 | — | — | — | 8 · 1.000 | 18 · 0.944 | — | — | 1 322 · 1.000 |
local_scope |
3 · 1.000 | 34 · 1.000 | 72 · 1.000 | — | 3 · 1.000 | 45 · 0.956 | — | — | — | 3 · 1.000 |
bare_name |
— | — | — | — | — | — | — | — | — | — |
Over all ten corpora: 55 158 answers, 385 wrong, 0.9930. By family:
| Rule | answers | wrong | precision |
|---|---|---|---|
package |
1 354 | 1 | 0.9993 |
self_receiver |
1 999 | 5 | 0.9975 |
nothing_declares_it |
9 266 | 24 | 0.9974 |
callers_file |
11 667 | 32 | 0.9973 |
visible_unique |
2 123 | 10 | 0.9953 |
import |
14 905 | 105 | 0.9930 |
inheritance |
1 007 | 11 | 0.9891 |
return_type |
1 562 | 18 | 0.9885 |
local_scope |
160 | 2 | 0.9875 |
written_type |
11 115 | 177 | 0.9841 |
return_type answers three times as often as it did — 494 of those are
Rust's, where it had 44 — and its precision fell from 0.9954 to 0.9885 in the
process. That is the trade the Rust rules made: a rule used eleven times as
much, wrong four times more often, for 427 further calls placed.
bare_name answers nothing here any more. It was Ada's parameterless calls,
and every one of them is now placed by a rule that reads the source — the
package a name belongs to, or the type a receiver is written as.
What this is for. A consumer who needs certainty keeps the families measured
at 1.000 on its language and treats the rest as leads. Before the field,
in_tree read off an import the file writes and in_tree chosen because a
name was unique were the same word.
Two cells read 0.000 and neither is a rule that fails. A family's precision
has to be read over the population it answers on. nothing_declares_it only
ever says outside, so against javac and GNAT — which record only in-tree
calls — every one of its answers is judged against a target the oracle says is
inside. The same family is 1.000 on ky, TSX, ArkTS, Go, Kotlin and Rust. Ada's
import row is the same shape: seven answers, all of them judged against an
in-tree target.
This table was wrong twice before it was right, and both mistakes are worth
knowing. Go read 0.000 everywhere because its oracle compares by CLASS and was
being run through the by-position comparison. And families that answer nothing
read 8 · 1.000 on Rust and 5 · 1.000 on C++, because a site where the
compiler itself says undetermined matches muundo's undetermined whatever
rule did or did not apply — a trivial agreement no family can claim. Both are
excluded now.
One row was a corpus and not a language, and then it was a defect. ky read
31.6 % against excalidraw's 86.0 %, measured by the same checker and the same
script, and the difference was written off as ky being a two-thousand-line
library built on callbacks. Most of it was one shape muundo could not read:
const ky = createInstance(); is what every consumer of that module calls into,
and a module-level constant bound to a call was not a declaration at all. ky now
reads 64.1 %. The gap that remains — a value holding a function, an option
object whose members are callables — is the part that really is the corpus.
The same call in every language: which rules are missing¶
A corpus says how much is missed. It does not say what. scripts/corpus_shapes.py
asks the other question: a call graph has a small number of SHAPES — a receiver
whose type the source writes, a chain through a declared return type, a member
declared on a base type, an overload told apart by what receives its result —
and each exists in every language, spelled differently in each. Where muundo
places a shape in one language and not in another, the difference is never the
language: it is a rule read in one grammar and not in the other.
The matrix names the rule that placed each cell, or MISS:
| shape | Ada | C++ | C# | Go | Java | Python | Rust | TS |
|---|---|---|---|---|---|---|---|---|
written_type |
ok | ok | ok | ok | ok | ok | ok | ok |
return_chain |
ok | ok | ok | ok | ok | ok | ok | ok |
base_member |
ok | ok | ok | — | ok | ok | — | ok |
result_type |
ok | — | — | — | ok | — | — | — |
interface_member |
ok | — | ok | ok | ok | — | ok | ok |
field_type |
ok | — | ok | ok | ok | — | ok | ok |
named_arguments |
ok | — | — | — | ok | ok | — | — |
aliased_namespace |
ok | — | ok | — | ok | ok | ok | — |
Every cell it can ask about is filled. Ten gaps on its first run, all ten closed:
- Go's struct fields were read as
identifierwhere the grammar writesfield_identifier, sotype Holder struct{ part Inner }recorded nothing. - Go's interface methods were not entities at all, so
s.area()on aShapenamed something muundo had never heard of — the same gap Rust's trait methods and Ada's spec subprograms had. import p.deep.Deep;thenDeep.work()builtp.deep.Deep::work, a name no entity carries, and called itoutside.using p.deep;thenDeep.work()looked for the member beside the namespace instead of inside the type. 1 962 of Newtonsoft.Json's calls.- An Ada call written without parentheses lost its prefix.
C.Describeis how Ada calls a parameterless function, and only the leaf survived — so the type ofCwas thrown away before anything could use it, andH.Part.Deep, with three segments, produced no edge at all. - An inherited Ada primitive needs no
with. A file that writeswith Child;callsC.Describewithout ever namingParent, and the visibility check refused it. - C++ emitted no inheritance edge.
struct Child : Parentproduced noExtendsdependency, so the rule every other language has had nothing to walk.
Three of those were in the matrix before they were in any corpus. The Ada pair
is worth 16 calls and three fewer wrong answers on ada-util — base_member and
field_type looked like one gap each and were one cause.
And with --lsp, the same corpus and the same score¶
The same command with a server named. Measured 2026-09-24 in this devcontainer; the time is the whole run, muundo included.
| Language | server | coverage | precision | was | time |
|---|---|---|---|---|---|
| Go | gopls 0.16.2 | 100 % | 1.000 | 80.8 % | 2 min |
| C | clangd | 100 % | 1.000 | 100 % | seconds |
| Java | jdtls (JDK 25) | 98.9 % | 0.9995 | 95.8 % | 15 min |
| Rust | rust-analyzer | 92.3 % | 0.953 | 84.2 % | 1 min |
| ArkTS | typescript-language-server | 91.6 % | 1.000 | 71.2 % | 3 min |
| Kotlin | kotlin-language-server | 90.7 % | 1.000 | 41.5 % | 1 min |
| TSX | typescript-language-server | 90.1 % | 0.993 | 86.0 % | 5 min |
| C++ | clangd | 75.2 % | 0.911 | 50.1 % | 9 min |
| TypeScript | typescript-language-server | 67.6 % | 0.995 | 64.1 % | 1 min |
| Ada | ada_language_server | 99.2 % | 1.000 | 98.5 % | 3 min |
| JavaScript | typescript-language-server | 34.6 % | 1.000 | 86.2 % | 1 min |
| Python | pyright | 67.2 % | 0.787 | 27.4 % | 1 min |
C++'s row moved with the oracle, and only C++'s. clangd read 79.9 % at precision 0.897 against the oracle that missed every call that can throw; it reads 75.2 % at 0.911 against the one that sees them. Rust's and C's rows are unchanged, because a Rust module compiles whole and cJSON's calls do not throw.
Two are not in the table, and neither is a muundo result:
- C# — OmniSharp 2.0.0 starts and answers
initialize, then places 0 before the ask budget runs out: the project has not opened. The provenance says so in the server's own words. - Python —
corpus_calls.pyreads a reviewed corpus of hard cases and has no--lspflag. The click row is scored bycorpus_python.pyinstead, which compares against the RUN. Its--lsp pyrightrow is stale: re-run after the two rules below, pyright added exactly nothing, digit for digit, which is what a server that failed to start looks like. Nobody has confirmed it ran.
The Python oracle RUNS the program. py_runtime_calls.py is a measuring
instrument for this repository's own development, pointed at corpora chosen and
read beforehand, inside a container. muundo analyze executes nothing, ever,
in any language — that is the property the engine is built on, and nothing here
is wired into it.
Python is measured twice, and the two say different things. The PyCG
corpus is a catalogue of hard cases by construction — a fair test of the limit
and an unfair one of ordinary code. scripts/py_runtime_calls.py records what
the interpreter actually called while a project's own test suite ran, and
scripts/corpus_python.py compares muundo against that. It is the strongest
evidence there is for a call that HAPPENED, and it is silent about branches the
run never took, so only the sites it recorded are judged.
It also sees what no reader of the source can: a method found through
__getattr__, an object a test monkeypatched. Those count against muundo, and
they are the honest reason pyright's own answers disagree with the run 501
times — a type checker names what is declared, a run names what happened.
jdtls needs JDK 25. Started on 21 it fails initialize with an OSGi
Require-Capability error, and muundo then reports unavailable and changes
not one edge — which is the right behaviour and looks exactly like a server
that added nothing.
Change evidence is recall / precision over the declarations git says changed:
384 commits and 1 265 declarations in all, with nothing excluded from the
count. Precision on call edges means an edge called in_tree really is one;
honesty means that of the calls it did not place, it said it could not place
them rather than placing them wrongly. 16 832 probes across 22 projects;
every law held.
All thirteen have a compiler to answer for them. Two answer through what
they WRITE rather than through an API: kotlinc keeps a line-number table and
a source name in every class, which javap -c -l -p prints, and GNAT writes
its cross-references into an .ali file, where a reference marked s is a
subprogram call. ArkTS is measured
on the files the TypeScript checker can READ: its declarative ArkUI body —
build() { Column() { Text(this.name) } } — is not TypeScript at all, so the
tree is handed over with .ets renamed to .ts and only the files that parse
with no syntax error are read back. In a HarmonyOS application that is the
models, the services and the utilities; the screens are covered by the probes. Each oracle found
defects on its first run — which is the whole reason for building one — and
the largest was shared by seven languages: new Thing() produced no call edge
at all, in every language that spells construction with a keyword.
YAML is supported and is not in this table: it has no declarations and no calls, so none of the three instruments has a question to ask of it.
What a site is, and where that is not enough¶
A call site is a FILE AND A LINE, which is all debug info gives. A line can write several calls, and where both sides have answers left over that disagree among themselves, no rule pairs them except by order — and order is not evidence. Those are counted apart as ambiguous and judged neither way: 96 of fmt's sites, 57 of clap's, and 0 to 6 everywhere else. The harnesses that read the source through a compiler's own API — TypeScript, Java, C#, Kotlin, ArkTS — put the NAME the source writes in the key as well, and so almost never reach it. Both sides read that name off the same line; it is not a name matched between two vocabularies.
The limit, and its size¶
muundo does not track what a name holds. Two consequences, both measured:
- Overload resolution. C++ and Rust pick between same-named declarations by argument type. What is left of it after the ambiguous sites are set aside is 27 of fmt's 1 819 compiled call sites and 8 of clap's 3 569, counted against muundo rather than excused.
- Value flow. Against the PyCG corpus, muundo resolves 33 % of the edges a value-flow analysis finds. 87 % of what it misses needs to know what a name holds — a parameter called as a function, a method on a receiver whose type is not tracked.
Asking a compiler instead¶
--lsp <command> asks a language server about the calls reading the source
cannot place. On anyhow, against rustc's own debug info, the call sites muundo
places correctly go from 59 to 119 of 120 — in 29 seconds, where the
analysis alone takes under one.
Every language muundo speaks has a server it can drive. One corpus each, counting the calls a server placed that reading the source had not:
| Language | Server | Placed |
|---|---|---|
| C | clangd | 68 on sds |
| C++ | clangd | 1 835 on fmt |
| Rust | rust-analyzer | 815 on anyhow — undetermined 720 → 34 |
| Python | pyright | 3 221 on click |
| TypeScript | typescript-language-server | 1 253 on ky |
| TSX | typescript-language-server | 2 020 on excalidraw |
| ArkTS | typescript-language-server | 783 on OpenHarmony photos |
| JavaScript | typescript-language-server | 1 120 on axios |
| Go | gopls | 573 on uuid — undetermined 505 → 7 |
| Java | jdtls | 590 on gson |
| Ada | ada_language_server | 14 on ada-util |
| Kotlin | kotlin-language-server | answers; okio's Gradle model had not opened after 20 minutes |
| C# | OmniSharp | answers; Newtonsoft.Json had not opened after 20 minutes |
The counts are not comparable across rows: each ran under its own
--lsp-budget, and what a server can add depends on how much the language
lets muundo place by itself.
The command is the whole configuration, arguments included, because that
is how a server is told where the project is —
--lsp "ada_language_server --config als.json". Ada places 14 and not 0
because of that file: without it the server opens no project. And it costs
what muundo otherwise does not: a project the server can open. Where it
cannot be started, the provenance says so in the server's own words and not
one edge changes.
It also changes the snapshot id, because the server's identity joins the
hashed inputs — two runs under different toolchains are two analyses and must
not share an id. A report made this way cannot be replayed by verify on a
machine without the same server, and the verdict says resolver_absent
rather than comparing against a weaker analysis. Every edge a server placed
carries resolved_by.
Neither limit is a defect to fix within this design. undetermined exists to say so
per edge, and it is the honest majority: across six trees and six languages,
in_tree is 8–19 % of call edges, outside 0–25 %, and undetermined the
rest.
Reproduce¶
scripts/corpus_change.py <repo> --commits 25 # any git repository
scripts/corpus_c.py <tree> # C, against clang
scripts/corpus_cpp.py <tree> --include <dir> # C++, against clang
scripts/corpus_rust.py <crate> # Rust, against rustc
scripts/corpus_go.py <module> # Go, against x/tools callgraph
scripts/corpus_ts.py <tree> # TS/JS, against the tsc checker
scripts/corpus_java.py <tree> # Java, against javac
scripts/corpus_csharp.py <tree> # C#, against Roslyn
scripts/corpus_arkts.py <tree> # ArkTS, against tsc
scripts/corpus_kotlin.py <tree> --classes <dir> # Kotlin, against its bytecode
scripts/corpus_ada.py <tree> # Ada, against GNAT's cross-references
scripts/corpus_calls.py PyCG/micro-benchmark/snippets
scripts/corpus_probe.py <tree> --lang <language> [--kind member]
What these numbers do not say¶
A 1.000 over 25 commits of one repository means muundo named the right
declarations there. It is enough to find defects — it found eleven — and it is
not a rate to quote about your code.
A held probe means the machinery answers when a declaration is placed where the language makes it visible. Only a compiler oracle says a project's own calls are all resolved. All thirteen now have one.