Muundo — Design Decisions¶
Deliberate "no"s — things that look like missing features but are refusals on purpose. Each one has a specific reason static analysis cannot do better without either executing user code or building a full language-semantics engine, which Muundo declines to add.
These are not TODOs or deferred work. Do not silently patch them with heuristics. If a real project hits one and needs it fixed, open an issue with a repro so we can implement the proper thing rather than a lossy approximation.
Import resolver — static-only limitations¶
core/src/parser/import_resolver.rs resolves imports per-language at a purely
static level. The following behaviors are out of scope by design.
A refusal that looks like a bug has not been made. Everything below is a
boundary of static analysis, and every one of them used to reach the report as
the same thing: the raw specifier the source wrote. express (a package this
analysis declines to follow), @app/db (an alias nothing answered to) and
./near (already resolved to a file, and the answer discarded) were one shape,
so a reader could not tell a decision from a gap — and asked, reasonably, why
the import was not resolved.
Each import edge now says which: outside for what this analysis does not
follow, unresolved for a name THIS tree claims and nothing answers to,
in_tree with the file for the rest. The refusals below did not change; they
became visible, which is the difference between a design decision and a defect.
TypeScript / JavaScript¶
- Wildcard collisions: when several
pathsentries match one target, the most specific one decides — an exact pattern beats a wildcard, and among wildcards the longest literal prefix wins — and only that pattern's substitutions are tried. If none of them exists, resolution fails rather than falling back to a less specific pattern, which is whattscdoes. (This paragraph used to say the opposite: that every matching pattern was tried in order until a file was found. That behaviour was removed because it invented edgestscwould never produce, and the sentence describing it stayed here.) node_modules/ third-party packages: intentionally out of scope, not a deferral. Structural rules reason about user code; resolving intonode_moduleswould bloat the graph with tens of thousands of library entities for zero value to those rules. Call edges already carry the raw callee name (express.Router,hono.get), which is sufficient for pattern-based rules.extendspackaged in node_modules (e.g."extends": "@tsconfig/node18/tsconfig.json"): we only follow relative/absoluteextendspaths. Package-resolved extends would require traversing node_modules, which we decline per the point above.
Python¶
- Dynamic
__init__.pyre-exports (__all__computed at runtime,globals().update(...), star-imports from namespace packages): impossible to follow statically. We resolvefrom pkg.submod import Xdirectly topkg/submod.py, which works for the vast majority of real code. Star- imports that rely on runtime package structure are not traced. .pthfiles and editable installs: not read. If a project puts its package root behind an editable install only, the resolver will miss it until apyproject.toml/setup.pysurfaces the layout.- Implicit relative imports (Python 2 style): not supported — Python 3
mandates explicit relative (
from .foo import bar), which we handle.
Rust¶
pub usechain depth = 1: we follow one hop throughlib.rs/mod.rs. Deeper hub-chains (mod.rsre-exporting from anothermod.rs) are not chased, to avoid runaway traversal on generated code.- Glob re-exports (
pub use foo::*): not chased. The set of names is only knowable after resolving every item infoo, which inverts the resolver's direction.pub use foo::Baris handled.
Kotlin grammar — depending on tree-sitter-kotlin-ng¶
Since October 5 the crate is a vendored copy, grammars/tree-sitter-kotlin-ng/
(upstream 1.1.0 at commit 3dea6df, +1.2 MB packed), because its scanner
inserts an automatic semicolon before the constructor keyword whenever the
previous token is a nested class's name: class Builder on one line and
internal constructor(…) on the next ended the class body, and the rest of
the file parsed as error nodes. Seven of KotlinPoet's files were cut short that
way — FileSpec.kt kept 257 of its 600 call sites, TypeSpec.kt 510 of
1 061. The patch is one function in src/scanner.c, described in the copy's
README.md: when a ( follows constructor, the semicolon is withheld where
the parenthesised list declares a val or var, or is followed by : and a
supertype — a primary constructor header, which cannot be a class member of
the class just named. Both Kotlin corpora parse with no error node under it;
the upstream crate is still what Cargo.toml names, by path, and the version
lock verifies the same 1.1.0.
Kotlin support (core/src/parser/kotlin.rs) depends on the
tree-sitter-kotlin-ng grammar crate. There is no official Kotlin
tree-sitter grammar (JetBrains ships none), so this is the de-facto
standard: maintained by amaanq under the tree-sitter-grammars org
(a core tree-sitter maintainer), MIT, ~1M downloads, 100+ reverse-deps.
It uses the modern tree-sitter-language 0.1 binding, so it builds
against our pinned tree-sitter 0.25 — the ABI blocker that previously
kept Kotlin out (old tree-sitter-kotlin pinned <0.23) is gone.
Accepted risk (low–moderate): single-maintainer/community governance
(not the official tree-sitter/ org), and a future grammar major could
rename node kinds and break core/src/parser/kotlin.rs.
Containment: version pinned in Cargo.toml and lock-verified at build
(MUUNDO_LOCK_TS_KOTLIN + verify_versions), so upgrades are deliberate.
core/src/parser/kotlin.rs is isolated behind our own node-kind handling; if the
grammar breaks or stalls, Kotlin extraction can be disabled by dropping the
Language::Kotlin dispatch arm without touching any other language. The
failure mode is graceful — unparseable .kt constructs yield ERROR nodes
(fewer entities), and a consumer's test map falls back to its heuristic; no
crash, no cross-language impact.
Containment — the root is proven, the traversal is not handle-pinned¶
Two acts, and only the first is atomic.
The analysis root is canonicalized and proven to sit inside the containment
boundary before anything is opened, and its filesystem identity — (device,
inode) on Unix, the volume serial and file index on Windows — is captured then
and compared again immediately before the walk. A root that was replaced between
those two moments is refused.
The walk itself takes a PATH. ignore::WalkBuilder, and walkdir beneath it,
accept no directory handle to anchor traversal to, so the instant between the
last comparison and the first openat is not covered. Closing it would mean
writing the traversal ourselves against a pinned handle — openat-relative on
Unix, a handle plus NtQueryDirectoryFile on Windows. That is a different
traversal, not a stricter check, and it would replace a well-tested dependency
with our own directory walker for a window that the next paragraph already
makes harmless to the OUTPUT.
What is guaranteed regardless: every file the walk yields is canonicalized and re-checked against the boundary before it is admitted. A replacement mid-walk changes which files are offered; it cannot change which are accepted. No file outside the boundary is read, hashed, or named in a report.
So the precondition is stated rather than assumed: the root must stay put for
the duration of the walk. A root on storage a third party can mutate while the
analysis runs is outside what the traversal half can promise. The engine records
the precondition in resolve_and_contain_root, the published reference repeats
it, and a test fails the build if the two ever stop saying the same thing —
because a guarantee that exists only in the code is one a reader cannot rely on,
and one that exists only in the documentation is worse.
A nested mount stops the walk, and the report says so¶
A mount inside the analysed tree is storage the caller did not name. On a network mount the traversal precondition above does not hold at all, so descending into it would analyse a tree nobody asked for under a guarantee that does not cover it. Discovery therefore stops at the boundary by default.
Stopping is the safe half; stopping silently is not. ignore's own
same_file_system prunes without a word, and a report that describes a subset
of the tree while claiming to describe the tree is worse than one that admits
the gap — a consumer branching on is_complete() would have been told nothing.
So the pruning is done with filter_entry, which prunes AND records: each
boundary reaches skipped_files under its root-relative path, raises a
file_discovery note, and makes is_complete() answer false.
AnalyzeOptions::cross_filesystem analyses through, for a project deliberately
spread across mounts. muundo-server sets it exactly when
MUUNDO_ALLOW_SHARED_WORKSPACE accepts the storage — a deployment that has
accepted a boundary it cannot vouch for has, by the same act, accepted the ones
nested inside it, and would otherwise get reports declaring themselves
incomplete for a line it already agreed to cross.
Per-file containment is untouched in both modes: every file is resolved and checked against the boundary before it is read.
Analysis resource budgets — bounded by default, env-overridable¶
Analysis caps its workload on purpose so a hostile or accidentally huge tree
cannot exhaust memory. collect_source_files() checks file metadata before
reading/parsing: a file larger than max_file_size is skipped (recorded in
skipped_files with an explicit reason and the walk continues), while breaching
max_total_bytes or max_file_count is a hard Error::LimitExceeded —
the server maps it to 413 Payload Too Large, not a generic 500. Parsing
retains decoded source_text only (the raw source_bytes are dropped after the
tree-sitter parse; the BLAKE3 hash is computed before the decode).
The limits live on AnalyzeOptions with named-const defaults in
core/src/analyzer.rs; the server reads the env overrides at startup.
Per-caller path (Python binding / CLI). The env vars above are read by the
server/CLI startup, not by every caller. The Python MuundoAnalyzer
constructor takes the limits as optional keyword args —
MuundoAnalyzer(root, languages=…, ts_stateflow_strategy=…, max_total_bytes=…,
max_file_count=…, max_file_size=…) — each defaulting to the engine const. A host
raises the ceiling for a large monorepo, or lowers it to fail fast in CI. (Before
this, the binding hardcoded AnalyzeOptions::default(), so it silently ignored
MUUNDO_MAX_TOTAL_BYTES / MUUNDO_MAX_FILE_COUNT and a >512 MiB repo just
errored with no recourse from Python.) A byte/file-count breach stays a hard
LimitExceeded on this path too — a global resource ceiling truncates to a
non-deterministic surviving subset, so it is refused (fail-closed), never served
as a partial report + PartialAnalysisNote. Only localized, bounded
partiality (state-flow-pairs truncation, doc-coverage degeneracy, a malformed
tsconfig, a filesystem walk error) degrades-with-note.
| Env var | Default const | Default value |
|---|---|---|
MUUNDO_MAX_FILE_SIZE |
DEFAULT_MAX_FILE_SIZE |
2 MiB |
MUUNDO_MAX_TOTAL_BYTES |
DEFAULT_MAX_TOTAL_BYTES |
512 MiB |
MUUNDO_MAX_FILE_COUNT |
DEFAULT_MAX_FILE_COUNT |
50 000 |
MUUNDO_ANALYZE_MAX_CONCURRENT |
— | 4 (concurrent /analyze requests) |
MUUNDO_ANALYZE_TIMEOUT_SECS |
— | 60 (per-request deadline) |
MUUNDO_MAX_SERVER_BYTES |
DEFAULT_MAX_SERVER_BYTES |
4 GiB (server-wide memory budget) |
MUUNDO_RSS_AMPLIFICATION |
DEFAULT_RSS_AMPLIFICATION |
160 (retained bytes / source byte) |
Server memory admission — weighted, not just a request count¶
max_total_bytes bounds one request's source bytes, but tree-sitter trees plus
the extracted entity/dependency/call-graph tables amplify that many-fold. Measured
peak RSS (core/benches/peak_rss.rs, VmHWM via /proc/self/status) on dense
synthetic input (one call-heavy function per few lines — a near-worst case):
| source | entities | call edges | peak RSS | amplification |
|---|---|---|---|---|
| 1.9 MiB | 32 800 | 32 000 | 298 MiB | ~155× |
| 3.9 MiB | 65 600 | 64 000 | 594 MiB | ~154× |
| 3.9 MiB × 4 concurrent | — | — | 2 370 MiB | ~linear in concurrency |
So a fixed request-count semaphore alone is unsafe: 4 concurrent 512 MiB inputs
would retain ~4 × 80 GiB. The server therefore adds weighted memory admission
(server/src/main.rs): a semaphore holding MUUNDO_MAX_SERVER_BYTES worth of
MiB permits, from which each request reserves
ceil(max_total_bytes × MUUNDO_RSS_AMPLIFICATION / MiB). The permit is moved into
the blocking closure, so a detached (timed-out) analysis keeps holding its share
until it truly winds down. This bounds the sum of concurrent retained
footprints to the global budget regardless of the count cap.
The memory permit covers analysis and JSON serialization (peak = report + encoding), both of which run inside the blocking task — and it is released at blocking-thread exit. It does NOT extend through transmission: once the report is spooled to its temp file, the in-memory report is dropped and RAM holds at most one streamed chunk, so keeping the weighted permit through streaming would reserve memory the response no longer uses (starving admission for no benefit). The spooled bytes are guarded instead by the separate aggregate spool budget (previous section), whose reservation lives inside the response body's stream state and is released when the body is fully sent, times out, or the client disconnects.
The per-request max_total_bytes is reduced at startup to
MUUNDO_MAX_SERVER_BYTES / amplification / max_concurrent when the configured cap
exceeds it (logged as a warning). Safe defaults — 4 GiB budget, 160×
amplification, 4-way concurrency — yield an effective source cap of ≈ 6.4 MiB, so
four concurrent worst-case analyses fit ~4 GiB.
Why divide by the concurrency. Not for safety: one request alone would fit
budget / amplification, and the semaphore already keeps the sum within the
budget. The division buys capacity. Each request reserves its worst case — the
cap times the amplification — because the server learns a tree's real size only
inside the engine, after admission. With a fixed reservation, the largest request
and the number that run at once are one choice: a cap of budget / amplification
would make every request reserve the whole budget, and analyses would run one at
a time whatever MUUNDO_ANALYZE_MAX_CONCURRENT says. We chose predictable
capacity — N worst-case analyses at once, always — over the largest request an
idle server could take. It is the arithmetic operators already do by hand for
PHP-FPM (memory_limit × pm.max_children) or PostgreSQL (work_mem ×
connections), enforced instead of advised. The alternative, not taken: reserve
by each request's measured size, which needs a size-only walk before admission.
An operator who wants larger single requests lowers MUUNDO_ANALYZE_MAX_CONCURRENT
or raises MUUNDO_MAX_SERVER_BYTES.
Operators on larger
hosts raise MUUNDO_MAX_SERVER_BYTES (and may then raise MUUNDO_MAX_TOTAL_BYTES);
those who have measured a gentler amplification on their corpus lower
MUUNDO_RSS_AMPLIFICATION to admit larger inputs. The arithmetic
(resolve_memory_admission) is unit-tested for the fit-within-budget and
never-below-1 invariants; the semaphore behaviour (exactly budget/weight
concurrent holders) has its own test.
One byte budget per analysis — source and configuration together¶
MUUNDO_MAX_TOTAL_BYTES is what the weighted memory semaphore prices a request
by: resolve_memory_admission divides the server budget by that number times an
amplification factor, and admits requests on the result.
Configuration reads were outside it. tsconfig.json / jsconfig.json and their
extends targets had their own aggregate budget, max_total_config_bytes,
defaulted to the SAME value — so an analysis could read 2 × max_total_bytes
while the server had admitted it for one. An extends chain makes that
reachable rather than theoretical: each link is a separate file, read and
retained, and each was charged to the budget nobody was watching.
There is one budget now. ConfigReadLimits::aggregate_budget is the analysis's
own ByteBudget — the one source reads claim, commit and release against — so
source + configuration ≤ max_total_bytes holds by construction and the weight
the server computes covers both.
A configuration that does not fit is skipped, exactly as an oversized one
is: the analysis continues with weaker import resolution rather than failing.
That is the same degradation a caller already sees from the per-file cap, and it
is visible in the resolver's skipped-config count. bench_config_heavy_budget
in core/benches/peak_rss.rs runs a chain of large configs against a
source-sized budget and fails if the analysis stops producing entities.
Spool admission — aggregate temporary-disk budget for streamed responses¶
Each /analyze response is serialized incrementally to a private temp file
(the spool) and streamed from it in bounded chunks. MUUNDO_MAX_ENCODED_REPORT_BYTES
caps ONE response, but N concurrent responses could each claim that allowance —
and the temp filesystem is often tmpfs, i.e. RAM. The server therefore adds a
server-wide spool budget (MUUNDO_MAX_SPOOL_BYTES, default = half the
server memory budget, resolved and clamped at startup — see below): a CAS-guarded
byte pool from which every response's CountingWriter
reserves as the encoder grows the file (never after the fact), released in
full when the spooled response is dropped — fully sent, timed out, client
disconnected, or failed mid-encode. Exhaustion aborts the serialization with a
retryable 503 ("report spool capacity unavailable"), the same capacity class
as the admission gates, never a partial spool. Concurrency behaviour (at most
capacity / size concurrent spools, exact release, no overshoot even under
racing reservations) is unit-tested.
Startup resolution — the spool must not silently defeat the memory ceiling.
A fixed spool default (say, 8 GiB) is dangerous precisely because the temp
filesystem is often tmpfs: those bytes are RAM the weighted memory semaphore does
not account for (the analysis permit is released the moment the report spools
to disk — see the previous section). An 8 GiB spool on a 4 GiB-budget box could
therefore push real process RAM to ~12 GiB while every in-process accounting
number stayed "within budget". resolve_spool_budget (unit-tested, pure) closes
this at startup:
- The default is a fraction of the server budget, not a constant:
MUUNDO_MAX_SERVER_BYTES / SPOOL_BUDGET_SERVER_DIVISOR(÷2 → half). It scales with the operator's configured budget instead of dwarfing it. - A memory-backed TMPDIR shares one combined ceiling with analysis. When
statfsreports tmpfs/ramfs, the spool is capped atserver_budget / 2and main() subtracts the resolved spool from the analysis budget fed toresolve_memory_admission. Soanalysis_RAM + spool_RAM ≤ MUUNDO_MAX_SERVER_BYTESholds by construction. On a real-disk TMPDIR the spool is separate storage and the full server budget stays available to analysis. - The value is clamped to the temp filesystem's real free space (
statvfs, 80% headroom), so a full spool can never ENOSPC the temp dir — RAM or disk.
Clamps only ever lower the value and are logged with the reason. The capacity
probe is best-effort: a probe failure (or a non-Unix target) skips the
filesystem clamp but keeps the server-fraction default. The startup arithmetic
and the combined analysis + buffers + spool ≤ budget invariant are unit-tested
alongside resolve_memory_admission.
Streaming-buffer admission — the third term in that sum¶
The invariant above used to be analysis + spool ≤ budget, and it was not the
whole process. Reading the spool back allocates: each response's reader takes a
batch of READ_BATCH_CHUNKS × BODY_CHUNK_BYTES — 1 MiB — plus the frames still
in its channel. Nothing admitted those bytes.
The reason it matters is that streaming is not bounded by the analysis semaphore. That permit is released when the report finishes spooling, i.e. when the response STARTS; the body then streams for as long as the client takes to read it, which a slow or stalled client controls. N stalled clients therefore held N MiB the accounting did not know about, on the hosts where the ceiling was set precisely because memory is scarce.
The read buffers are now a third CAS-guarded pool, resolved at startup by
resolve_buffer_budget. It has no environment variable of its own: it
follows from MUUNDO_ANALYZE_MAX_CONCURRENT and MUUNDO_MAX_SERVER_BYTES,
and an operator setting it independently of the concurrency it has to serve
could only get the two into contradiction.
- Sized for the configured concurrency —
max_concurrent × (batch + floor)— then capped atMUUNDO_MAX_SERVER_BYTES / BUFFER_BUDGET_SERVER_DIVISOR, so a large concurrency setting cannot quietly claim the whole process. - Always subtracted from the analysis budget, unlike the spool. There is no disk these bytes could live on: they are process RAM on every host.
- A floor per reader of one chunk plus what can remain resident from the previous batch. Below that a response could not stream at all, so a server budget that cannot cover one reader fails startup rather than accepting requests it will refuse.
A reader takes a permit before each batch, for exactly what that batch is about to allocate. Under pressure it halves the batch down to one chunk rather than exceeding the pool; it never waits for space, because the holder it would wait on may be a client that never reads again. A reader that cannot get even one chunk fails the response with a clear message instead of allocating anyway.
The permit travels with the bytes, not with the reader. This is the part that was wrong first: the permit was released when the reader finished the file, and the frames it had already handed to the channel were still resident — held by exactly the stalled client this exists to bound. The accounting read zero while the memory was there.
Each chunk buffer is framed with Bytes::from_owner, and the owner carries that
chunk's share of the batch's admission. The bytes go back to the pool when the
last clone of that frame is dropped, wherever that happens: in the channel, in
the response body, or inside hyper. Cancellation, read errors and client
disconnects all return them on the same path, because it is one Drop.
Entity identity — overloaded methods stay merged (semantics belong to the consumer)¶
qualified_name is file::Scope::…::name, built from purely syntactic
facts (tree-sitter AST). It distinguishes entities by lexical scope (nested
functions/classes no longer collide — see the scope-chain work), but two
overloaded methods in one class — same name, different signatures — share
one qualified_name and are therefore MERGED into a single node (one metrics
entry, one call-graph node).
We deliberately do not discriminate overloads inside muundo. The reason is a boundary decision, not a limitation to "fix later":
- Muundo is a syntactic extractor. It does the ~90 %: every entity with its
line_number/end_line, and (cheaply, if a consumer needs it) aparam_count. These are all facts the AST gives directly. - Overload disambiguation is a semantic concern. Routing a call
foo(x)tofoo(int)vsfoo(String)requires knowing the static type ofx— i.e. a per-language type-checker (arity only covers the different-arity subset). That knowledge lives in the consuming application, not in a tree-sitter extractor. - Forcing it into muundo would be net-negative. Making overload
qualified_names distinct is not a local change: the mapmetrics_by_entityis keyed byqualified_name, so distinct keys immediately break call-edge resolution (by_name["foo"]becomes ambiguous → the edge is lost) unless arity/arg-count is also threaded through the caller attribution, the resolver, and a newEntityfield — a change comparable in scope to the whole scope-chain refactor, for a rare case, and one that still can't resolve same-arity/different-type overloads without types.
The contract, therefore: muundo surfaces the raw syntactic facts (entity
qualified_name, line_number, end_line, and — when the need is real — an
additive param_count); a consumer that cares about overloads groups the
same-qualified_name methods and disambiguates them for its purpose, where
the type/context knowledge already exists. Muundo does not change its identity,
resolution, or metric keying to carry semantics it cannot soundly compute.
Entity.parameters — the map records what a function accepts¶
The report recorded, per call site, every argument a caller passes
(body_features.argument_flows), and nothing about what the callee accepts.
Meanwhile returns_args was documented as "every 0-indexed parameter
position the function directly returns" — the model indexed parameters while
never naming them, so a reader holding returns_args: [1] could not learn which
parameter that is, or even how many there are.
Entity.parameters closes that: the declared names, in source order.
One walk, two outputs. Every extractor already collected the parameter names
to compute returns_args, and threw them away. Rather than add a second walk
that would drift from the first, each extract_<lang>_returns_args now calls
extract_<lang>_parameters and indexes against exactly the list it returns. A
test in each of them pins the pair together.
The receiver follows the language, not a rule of ours. Python omits self
and cls because returns_args always did — the indices count from the first
argument a caller passes. A language that declares its receiver as a parameter
keeps it. What matters is only that the two fields agree, which sharing the walk
guarantees.
Empty means "not recorded". A non-callable has no parameters, and so does a
language whose extractor does not fill the field. The field is omitted from the
JSON when empty, so a report gains bytes only where there is something to say.
Kotlin was the one language with no parameter walk at all — it had no
returns_args either — and now has one written from the grammar.
BodyFeatures.parameter_uses — and whether the function read it¶
Entity.parameters says what a function accepts. On its own that answers half
a question: the defect it was added for was a parameter received and never
read — a cache lookup that took the tenant scope, documented it, and never
looked at it, so one tenant's answer was served to another.
parameter_uses is the other half: one entry per named parameter, with the
lines that read it. An entry with an empty use_lines is a parameter the body
never touches.
Not folded into use_def_chains. Those key on the whole access path —
p.count is the variable, not p — deliberately, so that a member written and
never read still looks that way. Adding parameters there would either invent
variables the code does not have or split a path's reads across two chains.
Seeding them was tried and reverted: it moved chains out of use_def_chains,
and it changed def_line for a parameter later assigned, because the walk
inserts with or_insert.
A member or index access counts as a read. p->count, p.count and p[i]
all put p in an identifier node of its own, so the pass sees it. A parameter
reached only through a field would otherwise read as untouched, and a false
"never read" is worse than no signal — it is the failure mode that killed the
call-graph asymmetry rule proposed before this.
One implementation, nine languages. It needs nothing language-specific: walk the body, collect the lines of identifier nodes whose text is a parameter name. Ada is the exception in form only — it has no separate body node in this grammar, so its specification is skipped by line, or every parameter would count its own declaration as a use.
Measured on the file that motivated it: across 1066 lines, exactly one parameter reports as never read — the defect — and none after the repair.
UseDefChain.def_contains — what a variable was built FROM¶
def_line and use_lines say WHERE a value is written and read. They never say
what it was made OF. A reader asking "which inputs decide this string?" had to
go back to the source, which is the one thing the extraction exists to avoid.
def_contains holds every identifier appearing in the expression the variable
was defined from, at any depth, sorted and de-duplicated. A string assembled
from a parameter, a local and a field of the receiver reports all three.
Both halves of a member access are recorded. collect_*_contained_identifiers
keeps only the object — obj.x contributes obj — because a taint question asks
"does this touch obj?". A provenance question asks the other half: WHICH part
of obj. Recording one and making every caller reconstruct the other is how the
two drift apart, so obj and obj.x both appear.
One mechanism, eleven grammars. parser::contained_with_paths takes the
language's own collector, its dotted-path builder and the node kinds that spell
a member access. A language contributes three constants and gets the behaviour;
nothing about the shape lives in a per-language file.
A destructuring target list shares one source. a, b = f(x) gives both names
the whole right-hand side. Narrowing per target would need to know which
component each name takes, which the grammar does not say — and a wrong
attribution is worse than a wide one, because it reads as precise.
Empty means "not defined from an expression", not "unknown". A parameter
arrives from the caller and has no source expression; so does a declaration
without an initialiser. Those are answers. The cross-language fixture in
core/tests/fixtures/def_contains/ is what keeps an unfilled extractor from
looking like one of them.
Field parity across languages — measured, not assumed¶
Every field the model has is filled by every extractor whose language HAS the construct. That was not true, and the gaps were invisible because an unfilled field and an absent construct look identical: both are empty.
The check is a fixture with one function per language exercising every field at once, run through the binary. Four gaps came out of it, none of which any test would have caught:
| gap | what it meant |
|---|---|
Kotlin had no returns_args |
a Kotlin function handing a parameter back was indistinguishable from one transforming it, so taint stopped at every Kotlin call |
| C and C++ had no regex detection | the same catastrophic pattern was reported in seven languages and invisible in two |
C++ is_async was never set |
a coroutine read as ordinary synchronous code |
Go swallowed_exceptions was always zero |
recover() stops a panic and tells nobody — a catch block that does nothing — and went unrecorded |
Parity means "has the construct", not "has the field". The cells still empty
are languages without the thing: Rust and Go have no exceptions (Result values
and panic/recover instead), Go, Java and Ada have no async keyword, Ada has no
regex in its standard library. Those are not gaps, and filling them would mean
inventing a meaning the language does not have.
Two of the four were node-kind mistakes caught by measuring. Kotlin's return
is return_expression in this grammar, not the jump_expression other Kotlin
grammars use; C++'s std::regex r("x+") puts the pattern in an argument_list
child of the init_declarator, not in a parameters field. Both first attempts
compiled, ran, and found nothing — which is why the fixture is run through the
binary and its output read, rather than trusting that the code looks right.
A verification covers the configuration, or it is not a verification¶
verify answers "does this tree still match this report". It re-read the source
files and nothing else, and for TypeScript that made the answer wrong in a way
nobody could see: the nearest tsconfig.json above an importing file decides
what @app/db points at, so repointing one alias changes every edge that
crossed it while every source byte stays what it was. The old verdict was ok,
about a graph that no longer existed.
So the configuration that resolved the imports is an input, hashed like source.
That closes half of it. The other half is a file that did NOT exist when the
report was made: drop a tsconfig.json into a source directory and it becomes
the nearest one for everything under it, and every file the report listed is
still untouched. A report therefore records, per directory of TypeScript or
JavaScript source, the places the walk looked and found nothing — and verify
checks they are still empty.
The evidence is derived, not trusted. A document that simply dropped those records would otherwise verify again on exactly the tree they exist to catch. So which directories must carry a record follows from the source the report itself lists, and each candidate chain follows from the walk rule — one definition, in the engine, re-derived by the verifier. A record that is missing, duplicated or shortened is refused as unverifiable rather than believed.
And a decision made out of reach is refused, not waved through. Where the
containment boundary is wider than the analysed root — what the HTTP server
does — the governing config can sit above the root, and the report cannot name
it. Nothing under --root can say whether that file still resolves imports the
same way, so verify refuses instead of answering a narrower question than the
one asked. Such a report is verifiable from the root that contains the
configuration, and from nowhere else.
Only a proven containment event is a containment event¶
A guarded open can fail for many reasons and the reader used to report one: a
containment refusal. On the HTTP surface that is a 422 and a page for whoever
owns the deployment, so rm inside the analysed tree — or a file with the wrong
mode — raised a security alert.
The refusals are classified now. ELOOP from an O_NOFOLLOW open is the policy
firing and keeps its code; on Windows the guard returns a typed error, because
there a reparse point, an out-of-root handle and an ordinary ACL denial are all
PermissionDenied and an io::Error cannot tell them apart. A path that no
longer resolves is unresolvable_path, one this process may not read is
unreadable, and neither is a security event: the file is missing from the
report, the report says so, and the analysis carries on.
The same rule decides what a CONFIG candidate is. canonicalize answers "not
found" for two different worlds — nothing is there, and a symlink is there
pointing at nothing — and recording the second as an empty place would fail a
tree nobody touched, the day the link's target appears. The entry is looked at
before it is resolved; a candidate that exists and cannot be used is a
degradation named in partial_analysis, which is not the same claim.
A refusal says which refusal, and never where the server keeps its files¶
Every refusal carries a stable code beside its sentence — invalid_root,
workload_too_large, containment_escape — because a caller that has to match
English prose is a caller who breaks when the prose is reworded. The codes for
what the ENGINE refuses come from the engine, so one failure is called one thing
on the command line, over HTTP and inside a report.
The sentence is path-free by construction, not by review. An engine failure interpolates what it carried — the file a parse died in, the canonical root a walk resolved — and that went out in the HTTP body. Each engine failure now has a fixed public sentence; the interpolated one goes to the log, where the operator is. There is no constructor that builds a refusal without a code, so a refusal added later has to name itself, and nothing on that surface interpolates an error it did not write.
The command line splits the same answer three ways in its exit code: 1 the
answer is no, 2 the invocation is wrong, 3 muundo could not answer. Every
failure used to be 1, so a build gate could not tell "the tree changed" from
"the check is broken" — and the second must never be read as the first.
.h belongs to the language its content says¶
C and C++ share the extension and the table has to give it one language, so it
gives it C. --languages cpp therefore discovered no .h at all: a C++ project
whose headers are named that way was analysed without its declarations.
Adding .h to both extension lists is the obvious fix and the wrong one — it
puts C++ headers in a C-only run, which is the same defect facing the other way.
Discovery admits the file under either filter, because at that point nothing has
read it; the content settles the question during the parse, and the filter is
applied again to the answer. Naming either language never means "and some of the
other one".
Two limits on a report, because they bound two different things¶
MUUNDO_MAX_REPORT_BYTES is how long the document may be.
MUUNDO_MAX_REPORT_NODES is what it may expand into, charged per node and per
byte of every string kept. They were one number — the raw input was capped at
the structural allowance — so the smaller governed both: every report over
64 MiB was refused while the documented byte ceiling said 512 MiB, and the
refusals included reports muundo had just written, because analyze indents its
output and whitespace is document without being structure.
The raw cap is still needed, and for one reason: a string is charged in
visit_str, which serde calls after it has built the value, so the peak
transient allocation is bounded by whatever bounds one token. That is the
document limit — no token is longer than the document, and the document is what
the operator sized.
What the analysis decided is published, not only what it scored¶
Three answers were computed and dropped, and each one had a consumer who could not get it back without re-implementing the engine.
The import resolution fed the metric graph — cycles, fan-in, fan-out — and the report published the raw specifier, so a reader could not tell a package this analysis declines to follow from a name nothing answered to. The documentation verdict, per entity, fed one term of the fragility score, so a coverage figure raised the question "which ones?" and the report had no answer. And the cycle detector returns the FILES that form each cycle, of which the report kept the count: "2 dependency cycle(s) detected".
The rule, stated once: an answer this engine computes for its own arithmetic is published beside the arithmetic. A score is a summary of something, and a consumer who trusts the score will eventually need the something.
The counterpart is that a field nobody fills is removed rather than kept as a
promise. entities[].description and entities[].dependencies were in the wire
contract and were empty in every language on every codebase measured; a reader
had no way to tell "not documented" from "muundo does not extract that", which
is the same defect facing the other way.
Two scores, because the first one measures mostly documentation¶
fragility weighs file fan-out and import cycles at 60 % and
1 - doc_coverage at 40 %. Measured across eleven of our own crates, the
structural half contributes between 0.007 and 0.06, so 88 % to 100 % of every
score is the documentation term. On sawabona-core, 0.202 is 0.185 of
undocumented code and 0.017 of everything else. On pimatika-pro-core the
structural part is exactly 0.000.
Two of its thresholds cannot fire. FAN_OUT_NORMALIZER divides by 20 a
file-to-file fan-out that ranges from 0.00 to 3.00 in practice, so that term is
capped at a seventh of its range; and the god-object risk factor triggers above
a fan-out of 20 where the largest measured anywhere is 11. Cyclomatic complexity
is not an input at all: six functions were taken from 36, 35, 28, 28, 26 and 26
down to 11–19 and the score moved by 0.005.
So a second index was added rather than the first one rebalanced.
fragility is a published field with consumers, and a score that quietly
starts meaning something else is worse than a score that is wrong in a known
way. A reader comparing two reports would have no way to tell which
definition each was written under.
risk_index weighs complexity 0.35, coupling 0.25, documentation 0.25 and
cycles 0.15, and publishes all four components individually — which is what lets
"0.31" be read as the four numbers it came from. Its normalisers are set from
the measured ranges, not guessed: coupling saturates at a fan-out of 4, cycles
at 3 (a count, because one cycle in a hundred files is one defect and dividing
it by a hundred says it is nothing), the worst function at complexity 40, and
the share of functions above complexity 10 at 5 %.
It is None, never a fabricated number, when its inputs were not measured — the
fragility fast plan skips cyclomatic metrics deliberately, and doc coverage
can be switched off. Absent says nobody looked, the same distinction
doc_coverage makes.
core/tests/report_arithmetic.rs re-derives every component and the combined
score from published fields alone.
Production scores beside whole-tree scores, because tests are most of a tree¶
Run on its own source, muundo scored 47 % documented and a risk_index of
0.53. Its production code was 72 % documented. The gap was the test suite: 62 %
of the entities were tests, and a test whose name is already a sentence carries
no doc comment. The documentation term measured the tests more than the code.
Documenting a thousand tests would have moved the number and told nobody anything. So the report now says which entities are not production code, and repeats the two scores without them.
One rule decides "not production", and it already existed. The entry-point
classifier skipped tests, specs, fixtures and examples by path, and Rust inline
tests by attribute, so that a test route would not seed reachability. The
production scores read the same verdict through one function,
reachability::non_production_reason. A second copy of "what is a test" would
drift, and the scores would then disagree with the reachable set about the same
function.
The whole-tree fields keep their meaning, for the reason fragility kept
its own: a published score that quietly changes definition is worse than one
wrong in a known way. production is a second answer beside them.
Coupling and cycles are not recomputed. They are measured between files, and an import cycle that runs through a test file is still a cycle in the tree.
The reason is published, not only the verdict: path or
test_attribute. A reader who disagrees with a classification can see which
rule made it.
A callee is a name, and an expression is not one¶
Every extractor takes the callee from the source text of the expression being
called. For a method chain that text is the whole receiver, newlines included —
so v.iter().map(…).collect arrived in a field that names things as four lines
of code. Thirteen per cent of the call edges of a real Rust tree were shaped
like that.
Three consequences, and none of them looked like a defect. Fan-out, hotspot
scores and the reachable set counted edges naming entities that do not exist.
The same text landed in argument_flows, a third of the report's bytes. And a
chain that merely mentions :: — qn.split("::").nth — is read by the
qualified-name rule as <file>::<name>, so verify demanded a file called
qn.split(" and refused the whole report: analyze produced documents its
own verify would not accept, on every Rust project it was pointed at.
The rule is now one function, applied AFTER resolution — the resolver reads the
receiver out of the raw text (PathBuf::from("/tmp").join gives it the hint
PathBuf, which the file's use bindings expand into
std::path::PathBuf::join), so cutting the expression down first would take
away what it works from. What resolution could not turn into a name is cut to
the segment after its last top-level separator, with type arguments dropped.
What still cannot be read as a name — an immediately-invoked closure — is handed
back as written, because inventing a name puts something in the graph that names
nothing.
No is_sanitizer flag — sanitiser knowledge belongs to the taint-tracker¶
The same boundary retired a field that used to violate it. muundo once emitted
Entity.is_sanitizer: bool, set by matching the entity name against a list of
sanitiser verbs. This was a semantic judgment wearing a syntactic costume,
and it was wrong on two counts:
- muundo cannot justify it. Whether a function actually clears taint — and
for which sink family (
html.escapeneutralises an XSS context but not a SQL sink;normpathneeds a following containment check) — is domain knowledge applied where taint flows: at the call site, in the consumer's taint tracker. muundo classifies definitions by name and has no call-site type or sink context, so a userdef cleanwas indistinguishable frombleach.clean. - A context-free bool, read downstream, suppresses findings it can't warrant. A consumer that trusts the flag clears taint the flag can't justify — a false negative, the dangerous direction.
The sink-aware value already lives in the consumer (a name → {sink families} map keyed on the callee tail), which satisfies the "carry the sink context" guarantee at the layer that has the context. muundo's bool was therefore redundant and unsound, so it is gone — muundo emits only the syntactic facts (entities, calls, edges). Recognising a sanitiser, and scoping what it clears, is the consumer's job. Same rule as overloads: muundo does not carry semantics it cannot soundly compute.
A preprocessor line is blanked before C# is parsed¶
#if HAVE_DATE_TIME_OFFSET written inside a method body ends the class node in
the C# grammar. Every declaration after it in the file is then read as a free
function: 67 of Newtonsoft.Json's methods lost the class they belong to, and
reader.Read() on a JsonTextReader found no Read under that class.
The directive LINES are replaced by spaces — never the code between them — so
every byte offset, line number and span after the directive is exactly what it
was. A directive declares nothing, so nothing is lost. #if/#else keeps both
branches, which is not valid C#; it was not valid before either, and the count
of call sites muundo finds went up, not down.
The type a receiver is written with, and the class that declares the member¶
The type the source writes for a name is the answer only when that type
declares the member. JsonObject may not declare toString — the call reaches
JsonElement, one class up — so where the written type declares nothing by
that name, the Extends edges are walked and the ARGUMENT COUNT picks the
member.
Only where the written type declares nothing. A declaration's parameter
count is not what it accepts: C# params and a defaulted parameter both take
fewer. Preferring a base class's exact count over a declared member sent 697 of
Newtonsoft.Json's calls to a method the source never names.
Two files, one type: partial class¶
C# writes partial class JValue in JValue.cs and again in
JValue.Async.cs. muundo indexes each file's half as a type of its own, which
is right for everything except a construction: new JValue(v) saw two
candidates and named neither. Where every candidate is a type of the same name,
a construction takes the half that declares the constructor. 2 326 of
Newtonsoft.Json's calls.
And var writes its type after the =. var reader = new JsonTextReader(sr)
names its type as plainly as JsonTextReader reader = … does, on the other
side of the equals sign. Reading only the left side left those receivers with
no type at all.
How many arguments the call passes¶
Which overload a C++ call reaches is decided by its argument TYPES, which this engine does not read — and that was taken to mean it could read nothing about the call. It can read HOW MANY there are.
fmt's chrono.h declares one write, taking five parameters of which two have
defaults, and writes write<Char>(out_, upper) a dozen times. Two arguments
cannot reach a declaration that needs four, so the call is another write
entirely, one file away. muundo named the near one because it was the only
write in the caller's file.
So every call site's argument count and every declaration's parameter list are
read, in ONE tree walk that no grammar needs code for: a call is a node holding
an argument list, a declaration is a node holding a parameter list, and every
tree-sitter grammar spells those with argument and parameter in the node's
name. What is read is three numbers — how many are declared, how few may be
passed, how many at most — because a default may be left out and a variadic
takes any number.
Where nothing fits, the call stays unplaced. A declaration that cannot take the arguments written is not a weaker answer, it is the wrong one.
Four things are deliberately not counted, each of which made a call look like it passed the wrong number:
int f(void)takes nothing — that is C's empty list, not one parameter;self,&selfandclsare receivers: Python and Rust declare them and no call site passes them;- a TRAILING comma closes nothing —
listOf(a, b,)passes two; write<Char>(out, v)has two lists whose node kind saysargument, and the template arguments are not the call's.
JavaScript does not check the count. A missing argument is undefined and
an extra one reaches arguments, so neither bound means anything there and
neither is applied to the JavaScript family.
An implicit receiver is still a receiver, and it is this¶
A bare call reaches a member only of the caller's OWN type, of one that type
extends or implements, or of a type enclosing it. fmt's chrono.h declares
duration_formatter::write, private to a class the caller is not in, and a
bare write(out_, name) written inside tm_writer bound to it because it was
the only write in the file that took two arguments.
Three things are not members reached that way, and each restored calls the rule had taken: a CONSTRUCTION names a type; a STATIC IMPORT is the caller saying in the source that it reaches the member without a receiver; and a FIELD INITIALISER is written at file level, so the extractors attribute it to the file's own scope and not to the class that holds it.
The receiver is a call, and the call says what it returns¶
JsonParser.parseString(json).getAsJsonArray() writes its receiver as a call.
Nothing in the source says what that expression IS — except the declaration the
first call reaches, which writes its return type. Reading it placed 1 131 of
the 1 375 calls gson wrote that muundo could not name.
It is not inference. The return type is written on the declaration, exactly as a parameter's type is written on the parameter, and the rule that reads one now reads the other. Where the declaration writes no return type, where two overloads return different types, or where the type it returns declares no member of that name, nothing is added.
It runs as a second pass, after every edge has been resolved once, because what the first link reached is what the second link is asked about. Four passes, so a chain of four is followed a link per pass.
A chain written down the page is one chain. new GsonBuilder() on one line
and .create() three below: each link is recorded on the line its own name is
on, so the first link is looked for at or above the second's line, nearest
first.
A construction returns what it builds. new GsonBuilder().create() is the
same chain with its first link written differently.
Go writes the type of everything, and had never been asked¶
Every parameter, every method receiver, every struct field and every x :=
T{}. The rule that reads a written type existed for six languages and not for
Go, whose calls are overwhelmingly recv.M(). Adding it took cobra from 383
calls placed to 675, against Go's own call graph, with nothing wrong.
A planted probe must be a program the language accepts¶
Twice in one day the probes failed, and both times they were right to: they had planted code that does not compile.
First a declaration with no parameters under a call site passing two. Then, once
chains resolved, BOTH links of one chain — parseString(json).getAsJsonArray()
became muundo_probe_2(json).muundo_probe_1(), where the planted first link
returns int and the second is then asked for on a type the tree no longer
says it is on. muundo answers undetermined and is right.
A site whose receiver is another chosen site is now left alone, and the chain it belongs to is read across lines, because that is how a fluent chain is written.
An Ada subprogram is scoped by its PACKAGE, and the package is the whole path¶
Util.Log.Loggers.Create (…) names a package of three segments. The receiver
hint every other language shares keeps only the last one — Loggers — which
scopes nothing in an Ada tree, where a declaration sits under its full dotted
path. 399 of ada-util's calls.
And a prefix that is an OBJECT names a package too. Log.Error (…) where
the source wrote Log : Util.Log.Loggers.Logger: Ada does not scope a
subprogram by a type, so the type mark is read for the package that declares
it — the dotted prefix when the mark is written in full, and otherwise the
package the type itself is declared in. 658 more.
A spec declares, a body implements, and both are read. .ads writes
function Curl_Easy_Strerror (…) return Chars_Ptr; followed by pragma Import
and there is no body anywhere. Reading only bodies left those calls naming
something muundo had never heard of. Reading both makes every Ada name appear
twice, so the resolver collapses the pair: the BODY is the answer where there
is one, because it is what runs and where a reader wants to land.
Together: 600 calls placed on ada-util becomes 875, and precision rises from 0.990 to 0.994.
A call whose receiver and name this tree never declares has LEFT it¶
Console.WriteLine(text) is not something muundo could not decide about.
Neither Console nor WriteLine is declared here, so the call left — and
answering undetermined made a real gap and an external call read the same. A
reader asking what leaves the codebase got nothing back.
Only where the call WRITES A RECEIVER. A bare name may reach anything the
language makes visible without writing it down: a base class's member, a used
package, an overload set. An unknown bare name therefore says only that this
analysis did not find it. Read as outside, it cost fmt 76 wrong answers,
every one of them a bare call.
Against tsc on ky: 3 643 more calls answered and precision rises from 0.994 to 0.998. Against Roslyn on Newtonsoft.Json: 1 126 more.
A trait's method is a declaration, in Rust as in Ada¶
fn into_resettable(self) -> Resettable<T>; inside a trait has no body
there, so nothing was extracted — and val.into_resettable().into_option()
could not follow its chain, because the first link named something muundo had
never heard of.
And the receiver's type is written in both spellings Rust has. impl Trait
names it directly; <T: Trait> writes it twice — val is a T, a T is an
IntoResettable — and only the second names a declaration, so the first is
rewritten through it.
rustc's debug info names the MONOMORPHISED impl and muundo names the trait's declaration: 18 of clap's 1 604 judged calls disagree that way, and the declaration is the answer a reader of the source wants. Placed calls go from 1 356 to 1 568.
Ada writes short names for long packages, and reaches through records¶
Three shapes, each worth a piece of what Ada could not name.
package Bool_Prop renames Util.Properties.Basic.Boolean_Property; and
package Mgr is new Util.Files.Rolling.Protected_Manager (…); both give a
local name to a package that exists nowhere in the tree under that spelling. A
generic instantiation is followed as a renaming of the generic: the instance is
a different package in the language, and every subprogram it offers is DECLARED
in the generic, which is where a reader wants to land. The chain is followed up
to four hops, because an instantiation of an instantiation is ordinary.
The renaming IS the visibility. The file that wrote it needs no with for
what it points at — Util.Properties.Discrete appears in no context clause and
is reached through Bool_Prop alone.
A nested package is recorded under its own name. package Protected_Manager
is new … inside Util.Files.Rolling scopes its subprograms as
Protected_Manager::Openlog, while the path that reaches it from elsewhere is
the whole dotted one, so both spellings are tried.
Self.File.Openlog (…) reaches a primitive of whatever the COMPONENT File
is. Reading only the segment before the subprogram name picks a local variable
of the same name three lines above the call, so the whole path is resolved: the
first segment through the type its declaration writes, each one after it as a
component of the step before.
And a type a spec only NAMES is a declaration. type Logger is tagged
private; is what every client writes, and the full declaration is hidden below
private — often in another file. Reading only the full one meant Log :
Logger named a type muundo had never heard of, so it could not say which
package Log.Info (…) reaches.
Together, and with the package path and spec declarations recorded above: 600 calls placed on ada-util becomes 968, precision 0.990 -> 0.991.
Ada finds a subprogram through its FIRST ARGUMENT, and withs a unit's parents¶
Initialize (Factory, "samples/") and Factory.Initialize ("samples/") are the
same call written two ways: Ada reaches a primitive operation through the type
of its first parameter, and the prefixed spelling is sugar for the other.
muundo read the prefix and not the argument, so half of how Ada is written went
unanswered — 385 of ada-util's calls are bare. Positional arguments only:
Q (Path => Factory) names a formal, and which formal is first is not what the
source writes there.
with Util.Streams.Files; makes Util.Streams visible too, and Util
with it. Ada withs a child unit's parents along with it, and recording only the
leaf of each context clause left 154 calls written against a parent that
appears in no clause.
An instantiation declares a subprogram. procedure Free is new
Ada.Unchecked_Deallocation (…); is the only declaration of Free anywhere,
and reading only bodies and specs missed every one. 143 calls.
A type mark is written with as much of the path as the file needs. Log :
constant Loggers.Logger where the package is Util.Log.Loggers, because the
file already said with Util.Log.Loggers — so packages are also indexed by
their last segment. And Print_Stream'Class names Print_Stream: an attribute
is written with a tick, and reading the last word of the declaration gave
Class, a type nothing declares.
Over one day, and with the package path, the renamings and the spec declarations recorded above: 600 calls placed on ada-util becomes 1 269, and precision rises from 0.990 to 0.991. Ada moves from the bottom of the table to above C++ and level with C#.
An Ada body sees its spec's context clause¶
mapping.adb writes one with and its mapping.ads writes five. Ada compiles
the body against both, and reading only the body's left 280 calls looking at a
package that appears in no clause of the file they are written in. It is the
single largest thing Ada was missing, and it is one rule of the language.
Every type declaration names a type. Only enums and records became
entities, so type Output_Stream is new … with private; — a derived type, which
is how Ada writes most of its hierarchies — declared nothing muundo could see,
and Stream : in out Output_Stream then named a type it had never heard of.
A derived type keeps its parent's name, so two packages declare
Output_Stream and neither is wrong: the one the caller sits in is the one it
means.
An instantiation can be a compilation unit. package Util.Strings.Transforms
is new Util.Texts.Transforms (…); binds a whole dotted path, not a segment of
one, so the alias index is asked for the whole prefix before its first segment.
Together with everything above, over one day: 600 calls placed on ada-util
becomes 1 722, precision 0.990 -> 0.994. Ada goes from the bottom of the table
to third, behind C and Java. With ada_language_server it reaches 88.8 %.
An edge says WHICH RULE placed it¶
resolution says what became of a callee: in_tree, outside, unresolved,
undetermined. It does not say how that was decided — and the second question
is what tells a reader whether to trust the first. in_tree read off an
import the file writes is not the same evidence as in_tree chosen because a
name happened to be unique in the tree, and a report that says only in_tree
cannot be filtered.
Every edge now carries placed_by: the FAMILY of rule that placed it, twelve
of them, spelled as a stable token the way resolution is. rebind_callee —
the one funnel every decision passes through — requires it, so a new rule
cannot be added without naming itself.
And each family's precision is measured, not asserted.
scripts/corpus_score.py <language> <corpus> <muundo> - --only <family> keeps
one family's answers and reads every other placement as undetermined, which
is how the table in docs/MEASUREMENTS.md is produced. It is not uniform: on
Newtonsoft.Json the written-type rule answers 3 554 times at 0.957 while the
import rule answers 3 691 times at 0.997, and every one of ada-util's ten wrong
answers comes from two families out of twelve.
Why this and not a confidence number. A score would be invented; a family name is a fact about what was read, and its precision is measured against a compiler. A consumer who needs certainty keeps the families measured at 1.000 and treats the rest as leads — which is a decision it can make, and could not before.
A family's precision must be read over the population it answers on.
nothing_declares_it only ever says outside, so on gson — where javac
records only in-tree calls — its two answers are both judged against a target
the oracle says is inside, and it reads 0.000. The same family is 1.000 on ky
and 0.989 on Newtonsoft.Json. That trap is why the table names the corpus in
every column.
Ada's declarations that are not bodies¶
Four more, each the only declaration of its name anywhere.
function To_Upper_Case (…) renames TR.To_Upper_Case; is how a package offers
what an inner instantiation computes. with function To_Object (From : in T)
is <>; is a generic's FORMAL subprogram, and every call to it inside the
generic names it. with package IO is new Util.Commands.IO (<>); is a formal
PACKAGE, and IO.Put (…) inside the generic reaches the generic it is new of.
And a derived type inherits its parent's primitives. type Output_Stream is
new Util.Serialize.IO.Output_Stream with private; declares nothing called
Write, and Stream.Write (…) on it reaches one declared in the parent's
package and nowhere else — so the search walks up the derivation chain, eight
rungs, before giving up.
A first argument is a path too. To_Object (From.First) passes a record
component, and reading the first identifier saw From and lost what the call
is about — the same rule that resolves Self.File.Openlog resolves this.
On ada-util: 1 722 calls placed becomes 1 784, precision 0.995. Over one day, 600 -> 1 784.
Translating a fragment to find the rule that was missing¶
An idea worth recording because it worked, and not the way it was meant to.
The proposal: where muundo cannot resolve a call in language A, translate the fragment to a language B it reads better, resolve it there, and carry the answer back. Tested on two of ada-util's remaining misses, written twice and analysed twice.
The first resolved in both. The second — L : constant Logger := Create
("log"); against two Creates that differ only by result type — came back
undetermined in Ada and in_tree in the Java version.
That is a real result, and its cause is not that Java is better. The Java
version carried a static import and the Ada version carried L : constant
Logger. The same evidence, written differently, and muundo was reading one and
not the other. Ada lets two subprograms differ by their result type alone, so
the declaration that receives the result is the only thing telling them apart —
and it sits three characters from the call.
So the rule was added rather than the translation: a bare Ada call whose result initialises a declaration of known type is the one returning that type.
Translation is the wrong vehicle and a good probe. It cannot create
information the source does not carry, and reading a fragment through a model
would trade a measured answer for an unmeasurable one — the property
placed_by exists to protect. But as a way of asking why does B manage what A
does not, it points at a missing rule every time the answer is "it doesn't,
the evidence was there".
It also needed one repair to work at all: Ada calls its parameter list
formal_part where every other grammar spells it with parameter, so the walk
that reads return types started from a single parameter and never saw the
return that follows the list.
The same call in every language: an instrument for finding missing rules¶
A corpus says how much is missed. It does not say what.
scripts/corpus_shapes.py asks the other question. A call graph has a small
number of SHAPES — a receiver whose type the source writes, a chain through a
declared return type, a member declared on a base type or an interface, an
overload told apart by what receives its result, a namespace named by an import
— and each exists in every language, spelled differently in each. The script
writes each shape once per language, runs muundo on it, and prints which RULE
placed it or MISS.
Where muundo places a shape in one language and not another, the difference is never the language. It is a rule read in one grammar and not in the other, and the matrix names which one.
Ten gaps on the first run, five closed the same day: Go's struct fields
(field_identifier, not identifier), Go's interface methods (no entity at
all), a member of a type named by an import, a member of a type inside a
usinged namespace, and the matching in the harness itself, which compared
qualified names case-sensitively and read every Ada cell as a miss.
It finds what no corpus can. using p.deep; then Deep.work() is 1 962 of
Newtonsoft.Json's calls and it was invisible in the totals — one shape among
thousands, indistinguishable from the receivers muundo genuinely cannot type.
Written on its own, beside the Java spelling that worked, it took one look.
An Ada call written without parentheses keeps its prefix¶
C.Describe is how Ada calls a parameterless function, and muundo kept only
the leaf — throwing away the type of C before any rule could use it. Three
segments, H.Part.Deep, produced no edge at all: the reader handled two and
gave up.
And an inherited primitive needs no with. A file that writes with
Child; calls C.Describe without ever naming Parent, because the derived
type carries its parent's operations. Only the first rung of the derivation
walk is asked to be visible, and it is — the file declared a variable of that
type.
C++ says what a class extends¶
struct Child : Parent produced no Extends dependency, so the inheritance
rule every other language has had nothing to walk and c.describe() came back
undetermined. Every base in the clause becomes one edge, public and private
alike: what a call reaches does not depend on who may write it.
Together, on ada-util: 1 790 calls placed becomes 1 806, and the wrong answers fall from 9 to 6.
Ada's visibility is not a list of imports¶
Nine rules, all found the same way: the calls muundo missed on ada-util were
filed by what they looked like (scripts/corpus_misses.py), then grouped by the
declaration GNAT says each one reaches. Each group named one rule. Together they
take Ada from 1 806 calls placed to 1 995 of 2 098, and the wrong answers from 6
to 2.
Each rule reads something the source wrote. None counts candidates.
A body and its spec name packages together. Falling back to the spec only
when the body named none let one package UBO renames Util.Beans.Objects; in a
body hide every instantiation its spec declares. Ada puts the instantiation in
the spec and the calls in the body, so this was the single largest group.
A type has as many parents as it writes. is new Print_Stream and
IO.Output_Stream with private derives from one type and implements an
interface; an interface derives the same way, is limited interface and
Util.Streams.Output_Stream. And a private extension — is new P with private —
is its own thing in the grammar, so reading only derived_type_definition
missed every type whose parent a package hides, which is how a library writes
them. The derivations now form a tree and the nearest rung is tried first.
A child unit sees what its parent declares. package body
Util.Beans.Objects.Datasets calls To_Object and writes no context clause at
all: being the child is the visibility.
And it inherits its parents' context clauses. Util.Files.Filters calls
Util.Strings.Index while neither its spec nor its body writes with
Util.Strings — Util.Files does. The unit a file holds is its basename, so
dropping the last -segment names the parent wherever in the tree it sits.
A body sees its spec's file-scope declarations. A generic's formal subprograms are written just before the package, under no package name, and every rule that reads one file or one package walked past them.
An instantiation is found tree-wide when one name answers. A spec declares
package Boolean_Property is new Util.Properties.Discrete (Boolean); and
another unit writes use Util.Properties.Basic; and then Boolean_Property.Get
(…). Exactly one alias may carry the leaf; two and the rule says nothing.
A generic package's formals are declared just outside it. generic with
procedure Put (…) is <>; package IO is … puts Put in the enclosing unit, and
IO.Put (…) reaches it.
A formal package names itself without a field. with package IO is new
Util.Commands.IO (<>); writes IO as a plain identifier and the grammar gives
it no name field, so the whole declaration was skipped and every call through
the formal reached nothing.
Two bodies for one spec is a portable unit. Util.Systems.IO has an
os-unix body and an os-windows one and a build compiles exactly one; asking
for a single body found two and answered nothing. The declaration both share is
the spec.
A nested package is never named in a context clause. You write with
Util.Files.Rolling; and the Protected_Manager inside it comes with the unit,
so the question is whether the unit that DECLARES the subprogram was withed —
and the file it is declared in is what names that unit. Compared without case,
because the unit comes off a file name and GNAT writes those in lower case.
With ada_language_server the same corpus reads 97.3 %, up from 88.8 % before
these rules: the server and the rules agree about different calls.
Eight more, and Ada answers 98.5 % with nothing wrong¶
The same instrument, run again on what was left. Four of the eight were DECLARATIONS muundo never read — and a call whose declaration does not exist cannot be placed by any rule, so no amount of resolution would have helped.
A protected entry is a declaration. entry Enqueue (Item : in
Element_Type); inside a protected type is how Ada writes an operation that may
block. Nothing was recorded for one.
A protected type and a task type are types. Into.Buffer.Enqueue (Item)
reaches that entry through a component whose type is protected type
Protected_Fifo, and with nothing recorded under that name the package
declaring the entry could not be found. Worth 49 calls on its own.
An expression function is a declaration. function Is_File_Excluded (R :
Filter_Result) return Boolean is (…); is Ada 2012's one-line form, and every
call to one looked for something that was not there.
A library-level instantiation writes its whole path. procedure
Util.Encoders.KDF.PBKDF2_HMAC_SHA256 is new …; is a compilation unit with no
package around it, so the prefix is the package and the last segment is the
name a caller writes.
The other four are visibility and types.
A use clause may sit in any declarative part. util-http-headers.adb
writes use Util.Strings; inside the function that needs it, and the reader
descended only into the top of the file.
A subtype is its base type for every operation. subtype Ref is IR.Ref;
inside a generic is how Ada re-exports a type from the instantiation it wraps,
and a call on a Ref reaches a primitive two packages away — reachable only by
following that one line.
A record may inherit the component a call reaches through. type
Encoding_Stream is new Encoding.Encoder_Stream with null record; declares no
component at all, and Stream.Transform names one of the type it derives from.
A package path may be written relative to an ancestor unit. Inside
Util.Log.Formatters, Strings.Formats.Format (…) names
Util.Strings.Formats: Util is visible by being an ancestor, so the path
starts one segment in. The COMPLETED path has to exist as written — matching it
by its last segment instead would let Main.Other find a package Other the
file never withed, which is the visibility rule this completion is not allowed
to go round. That mistake was caught by a test already in the suite.
On ada-util: 1 995 calls placed becomes 2 088 of 2 119, precision 1.000,
and 99.2 % with ada_language_server.
Rust writes the type of a local nowhere, and says it anyway¶
Eight rules, found the same way as Ada's: the calls muundo missed on clap were filed by what they looked like, then grouped by the declaration rustc says each one reaches. Half the misses were one shape — a method called on a name whose type Rust infers — and every one of them had the answer written a line or two above.
Nothing is inferred. Each rule reads something the source wrote.
A name holds what filled it. let styled = StyledStr::new(); and then
styled.push_styled(…) twenty lines below: the call at the end of the
initialising expression says what the name is. The same fact types if let
Some(action) = …, for arg in … and a closure's one parameter, because all
four write the expression down beside the name.
A method that hands its receiver back is not the end of a chain.
self.get_num_args().expect(MSG).min_values() calls min_values on what
get_num_args returned; expect is declared in no tree here, and stopping at
it ended the chain one link short. The list is short and closed — unwrap,
expect, clone, as_ref, iter and a dozen more.
A return type's first type ARGUMENT is recorded beside its outer name.
-> Option<ValueRange> returns an Option, which no tree declares, and the
call written on the result lands on the ValueRange inside it. Both are kept
and the member being looked for decides which answers — so the same fact serves
impl Iterator<Item = &Arg>, where the type argument is the element a loop
binds.
And so is a written name's. subs: Vec<Sub> says both the container and
the element, and for sc in &mut self.subs binds the second.
-> Self names the type the declaration sits in. Rust's builders are
written that way, and a type called Self is declared nowhere: every call
chained onto one came back undetermined. Worth 162 calls on clap alone.
Self::default().id(id) is the same word in the receiver, and the same answer.
A chain is walked segment by segment. Reading the whole receiver at once
made new Builder()\n.with()\n.create() answer Builder — the first ( in
the text — instead of create. Cutting at the last . outside any brackets
and stepping left is what makes a chain written down the page one chain.
A return type is not cut at the = inside a generic. impl Iterator<Item =
&Arg> was read as Item, because the scan for the body's {, ; or =
counted the one inside the angle brackets — and -> ends with a >, so the
depth had to know that too.
A path written from the crate root stays in the crate. crate::Error::
invalid_utf8(…) and super::ValueParser::bool() both prove the target is in
this crate, and Rust scopes a declaration by the type its impl is on rather
than by the module path a call writes — so the segments in between matched
nothing, and the import rules read the first of them as another crate and
answered outside. That is a wrong answer, not a vague one, and it was 166 of
clap's calls.
Against rustc's debug info on clap:
1 367 -> 1 838 placed of 2 182, 65.1 % -> 84.2 %
return_type answers 44 -> 494, and its precision rose from 0.932 to 0.966
with rust-analyzer 92.3 %
every probe holds (67 of 67)
Java gained 239 calls from the same work, because the chain walk is shared: 93.7 % -> 95.9 %, precision unchanged at 0.9996.
CommonJS is a module system, and muundo read half of one¶
require and module.exports bind names exactly as import and export do.
Reading only the second pair left every call through the first naming nothing —
652 of express's calls are express(), a name bound by a require.
Three things, and all three are needed before one call resolves:
var x = require('m')bindsx, andvar {a, b: c} = require('m')binds each name, renaming included. Until this,requireproduced a dependency between two files and no binding at all.exports = module.exports = createApplication;names what the module hands back. However many assignments are chained, the name is the one at the end.- A module may hand another's on.
module.exports = require('./lib/express');is a whole file whose content is "ask that one", and express's entry point is two of them deep — the name arequireof it binds is declared in neither.
Against the TypeScript checker on express: 300 calls placed becomes 964 of 1 119, precision 0.999. TSX and ArkTS are unchanged; ky is ESM and does not move.
A report says what the source wrote, where resolution replaced it¶
CallEdge.wrote — added with schema 0.7.0, absent on almost every edge.
express() reaches createApplication: the callee published is the
declaration, and the token on that line is gone. Anything that reads a report
back against the source needs both — a reviewer, a diff, and the harness that
scores this engine, which keys each site on the written token and had 654 of
express's calls silently leave its population the moment they started resolving
correctly.
A constant bound to a call is a declaration, and it carries a type¶
Four shapes of TypeScript, all of them written down in the source, all of them walked past. Measured against the checker on ky: 1 007 calls placed becomes 2 042 of 3 187, precision 0.995.
A module-level const x = f(); is a declaration a call can reach. const ky
= createInstance(); is what every consumer of that module calls into, and the
checker names that line when asked. Module level only, because a local inside a
function is part of no module's surface — ky declares one of the same name three
lines into createInstance. And not a name bound to a require: there the
declaration a call reaches is the other module's export, and an entity here
would shadow the import rule and answer the wrong file.
What such a name HOLDS follows the import across files. ky.get(url) is
written in a test and ky is bound in source/index.ts, so the binding that
says what it holds is in neither the caller's file nor its own scope. The import
leads to the file and the exported name; the binding there names the call that
filled it; that call's return type is what the member is looked for on.
An arrow function is named by what it is bound to. const make = (): Inst =>
… writes = between the name and the parameter list, and the reader scanned
backwards from the (, found the sign and stopped — so every arrow in
TypeScript and JavaScript reported no name, and its return type and its arity
were recorded against nothing.
A function TYPE writes its result after an arrow. get: (u: string) => Prom
is how an interface declares a callable member, and the => was being read as
the start of a body — so the type was cut off before it was reached.
And a parameter is not a parameter list. TypeScript spells one
required_parameter, which contains the word the reader matches on, so
head(u: string): Prom was read once from its list and again from u — the
second reading said the declaration returns head. Two answers for one name is
no answer, and every interface member lost its return type. This one was
invisible until the name fallback above made the second reading produce a name.
Java loses 10 calls to the same changes and keeps its precision: 95.9 % -> 95.8 %. Ada, Rust, C, Kotlin, ArkTS and TSX are unchanged.
A receiver path is walked one field at a time, in TypeScript too¶
this.menuContext.broadCast.emit('x') reaches a member of whatever broadCast
is, and broadCast is a field of MenuContext declared in a third file. A
per-file table of written types answers the first step and nothing after it.
So TypeScript now keeps the index Ada keeps for record components: (the type,
the field) -> the field's type, gathered from every class and interface in the
tree. this is the type the caller's own declaration sits in, and every segment
after the first is a field of the type before it.
On the ArkTS corpus: 2 392 calls placed becomes 2 416, precision still 1.000.
It is not why ArkTS lagged. That corpus's misses were classified, and the
biggest group is not a shape at all: 137 of 444 are minified code — a bundle
where the call is bu and the declaration is ku, one letter each. Another
sixty are calls on a holder of static helpers (MathUtils.rectToPoints) whose
declarations are top-level functions. Neither has anything to do with the rules
that lifted ky, which is why ArkTS gained nine calls from them and not nine
hundred.
new T names the type, in every language, and C++ pays for it¶
Seven languages name the type a construction builds, and one test holds them to
it. C++ is the one where that costs something measurable: new leveldb_t on a
plain struct with no constructor compiles to an allocation and nothing else, so
clang records no in-tree call and muundo records one. 29 of leveldb's 159
remaining disagreements.
The edge is still worth reporting. new leveldb_t IS a dependency of that file
on that struct — it is what fan-out, hotspots and the reachable set are asked
about — and answering differently for C++ than for Java or TypeScript, on the
same shape of source, is what the shape matrix exists to prevent. So the rule
stays and the cost is written down here rather than discovered again.
What is NOT deliberate, and is the next thing to fix: mutex_.AssertHeld()
where mutex_ is a member whose type a header declares comes back outside —
23 of those 159. outside is a claim, not a shrug, and it is the wrong one.
Four shapes every published C++ library writes¶
C++ read 50.1 % on fmt and 38.7 % on leveldb. Four extraction gaps, none of them subtle once named, and all four are in code any C++ library ships.
A visibility macro hides a whole class. class LEVELDB_EXPORT Status { … };
is how a library with a published ABI declares a type, and no grammar can know
LEVELDB_EXPORT is an attribute: tree-sitter reads the declaration as a
FUNCTION whose type is a class called LEVELDB_EXPORT. The class, its methods
and every call to them disappeared. A pre-parse pass blanks an identifier in
CAPITALS between class/struct/union and the type's own name, space for
space, so every line and column in the report still points where it did.
const T& and T* lose the name. C++ passes an object by reference or by
pointer, and the name then sits inside a reference_declarator or a
pointer_declarator rather than beside the type — so a parameter written either
way had no recorded type, and every call on it was unplaced. Reading only
Thing t covered the one form C++ code hardly uses.
A lock built on the stack is a call. MutexLock l(&mutex_); runs a
constructor, and it is how C++ takes a lock, opens a file or wraps a buffer —
RAII is the language. Thing t; stays out, because nothing is passed and a
default constructor may not exist; Slice s(a, b); stays out too, because the
grammar reads it as a function declaration and nothing in the file says which it
is.
A type declared in a header is visible. The cross-unit rule consulted a list of function names — C's rule, where a function must be declared before use — and a class is not in it, so every construction of a type from another file was refused. In C++ that is nearly all of them; the file the type is written in is the evidence.
What was tried and taken back: exempting a class's METHODS the same way cost 194 calls on leveldb. Too many candidates became visible at once, so the rules that need one answer got none. Visibility was not what held those back.
C++ (fmt): 50.1 % -> 54.0 %, precision 0.955 -> 0.968 C++ (leveldb): 38.7 % -> 66.1 %, precision 0.865 -> 0.914 C, Rust, Ada, Java, Go, Kotlin, TypeScript: unchanged
A dynamic import binds a name, and a real codebase showed it¶
const {authHeaders} = await import('./util'); is how modern TypeScript splits
a bundle. It binds a name exactly as require('./util') and import … from
do — and muundo read neither the binding nor the module.
It was found by running muundo over a 462-file application and comparing it with
another tool: 27 calls to authHeaders, 2 placed. The declaration was there.
Every call was extracted. Only the join was missing, and the reason was one
spelling of an import.
Two details cost a try each. The call sits inside an await_expression, so the
name it binds is two levels up, not one. And import is not a function: a
module specifier must not become an edge to something called import.
TSX (excalidraw): 86.0 % -> 86.9 %, precision 0.993 -> 0.999 TypeScript (ky): 64.1 % -> 67.5 % JavaScript, ArkTS: unchanged
Saying a call left the tree needs the receiver the source wrote¶
Two passes run after resolution: one cuts an unresolved callee down to a name, and one says what became of every call. They ran in that order, and the first took away what the second needs.
x.read()?.unwrap_or(1) arrives at resolution as its whole text. The rule that
proves a call LEFT the tree requires a written receiver — a bare name may reach
anything a language makes visible, and reading an unknown bare name as outside
once cost fmt 76 wrong answers. After the cut, every tail of every chain was a
bare word, so the rule refused it and the call got no answer at all.
Stamping now runs first. On jagora-worker that is 48 % of calls carrying an
answer, and 90 % — the same 70 in-tree edges, and 124 calls that went from "no
answer" to a proven outside.
Two more things fell out of the swap. Go gained 43 in-tree calls, because a
resolved name that is not name-like was being cut before anything could match it
against an entity — and so did 69 of fmt's, where resolution lands on
formatter<T, Char>::parse. That name is now left whole: cutting it left an
in_tree edge naming no entity, which is the one thing that pass exists to
prevent.
Go (cobra): 82.1 % -> 86.5 %, precision 1.000 ArkTS (photos): 71.2 % -> 73.9 %, precision 1.000 C++ (fmt): 54.0 % -> 54.5 % C++ (leveldb): 66.1 % -> 65.8 %, precision 0.914 -> 0.896 Rust, Java, Ada, C, TypeScript: unchanged or a tenth of a point better
leveldb pays 30 wrong answers for 40 more correct outside ones. In C the same
rule also demands a DOOR — an include that leaves the tree — and the tree-wide
version has no such guard. That is where to look if this precision matters more
than the answered share.
Python: two ordinary lines, and 32 points¶
Python was the lowest row on both tables, and the two measures agreed about it. What they were pointing at was not hard.
import pkg then pkg.helper(). The module resolves to
pkg/__init__.py, which declares nothing and writes from .util import helper
— so looking for helper there found none and left pkg::helper, a name
nothing carries. 1 453 of click's own calls, and it is the most ordinary line
in the language. muundo already knew how to follow a package that gathers its
parts; this rule simply was not asking it to.
len(x) reaches the interpreter. The rule that proves a call left the tree
refuses a BARE name on purpose, because a bare name may reach anything a
language makes visible. Python's built-ins are the exception that can be proved:
the list is closed, and a tree that defines its own len declares one — so the
two conditions together are evidence and not a guess. 1 380 calls in one real
application, len, float, str, int and isinstance being the first five.
Against click's own run — what the interpreter actually called:
863 calls placed becomes 2 186 of 3 681: 27.4 % -> 59.4 % precision 0.926 -> 0.804
The precision cost is real and it is the point of the metric. Those 533
disagreements were 2 218 sites with no answer at all before; muundo now answers
them and gets about one in five wrong. The comparison that matters: pyright,
against this same run, scores 0.787 — so muundo alone is now within eight points
of a type checker's coverage and slightly more precise than one. Part of what
remains is the oracle's nature, because a run sees a method found through
__getattr__ and an object a test monkeypatched, and no reader of source can.
The share of ALL of click's calls carrying any answer: 46 % -> 77 %.
The language answers for some names, and that is provable¶
The built-in rule was written for Python and is not about Python. len(x) in
Python and Some(x) in Rust are the same shape: a bare call to a name the
language itself carries, which no tree can reach and none of them declares. So
the rule now takes its list from the caller's language — PYTHON_BUILT_INS or
RUST_PRELUDE — and the two conditions that make it proof are unchanged: the
name is on the closed list, and the tree declares no such name.
Rust's list is the variant constructors first, because Some and Ok are calls
and clap writes 938 of them, then the macros every file uses. A macro is written
with a !; the extractor has already dropped it by the time this rule sees the
name.
clap, share of ALL its calls carrying any answer: 52.0 % -> 55.3 % coverage against rustc unchanged at 84.4 %, precision 0.968
What it does not fix. unwrap is called 1 383 times in clap and stays
unanswered, because the tree declares one — MatchesError::unwrap, a method on
one type. The set of declared names is flat, so one declaration anywhere blocks
the proof everywhere. Narrowing it to the receiver's type is the next lever and
it is a different rule.
Kotlin: what a val holds, and what a file-scope function belongs to¶
Kotlin sat at 58.0 % while Java, on the same JVM and mostly the same shapes, was at 95.8 %. It had never been classified. Two rules closed most of the gap.
val b = Thing.builder(). Kotlin's parser did not fill value_bindings,
the fact that says what a local name holds — Rust and TypeScript both filled it
already. Filling it makes b.add(1) and b.build() reach the nested Builder,
which is how every builder in the language is written.
xs.toImmutableList(). An extension function is declared at file scope and
called on a receiver, so looking for it among the receiver's members finds
nothing. A basename declared exactly once at file scope is the only thing that
call can reach.
KotlinPoet: 58.0 % -> 76.0 %, precision 0.986 -> 0.989 klaxon: 41.5 % -> 45.0 %, precision 1.000 -> 0.947
The scope functions are per language now. x.apply { … } returns x, so a
chain reads through it — the same reason Rust's HANDS_IT_BACK exists. The table
cannot be shared: TypeScript's apply is Function.prototype.apply and returns
what the function returned. Kotlin's list is apply, also, takeIf,
takeUnless, requireNotNull, checkNotNull, and deliberately not let or
run, which return the lambda's result. On KotlinPoet it is worth one call —
kept because it is correct and costs nothing, not because it earned its place.
A chain written across lines had lost its receiver¶
A long method chain is formatted one link per line. Every extractor records the text of the expression being called, so the first link arrived as
self.names
.iter
— a receiver, a newline, and a method. Nothing downstream reads a receiver
through that, and the edge was cut down to the bare word iter, which every
receiver rule refuses on purpose: a bare name may reach anything a language
makes visible. Half of clap's calls carrying no answer at all were that
shape.
join_chains_written_across_lines runs before resolution and takes out the
spaces that separate the links. Only those. A space between two words is
load-bearing — new ArrayList is two words — and gluing them cost 3 278 of
gson's resolved edges in the first version of this pass, because a construction
stopped looking like one. So a run of whitespace with a word character on both
sides becomes one space and every other run disappears. A comment between two
links stops the join entirely: joining would make the comment part of the path.
The receiver's type is written, and this tree does not declare it¶
groups: Vec<Id>, then self.groups.extend(…). The method belongs to Vec,
Vec is not here, so the call left the tree. That is proof of the same kind as
the prelude table — not "nothing was found", but "the type it belongs to is
somebody else's".
It is the answer left_the_tree cannot give, because that rule refuses as soon
as the BASENAME is declared anywhere in the tree. clap writes its own extend on
FlatSet and its own iter on three types, which left 373 calls on a plain
Vec with no answer at all.
A type PARAMETER is excluded: fn f<T>(val: T) writes T, which is declared
nowhere and means the opposite — anything at all. The test is a name of more than
two characters carrying a lowercase letter, which Vec, String and HashMap
pass and T, U, K and IO do not.
Allowed per language, because the answer is not the same everywhere.
| where | what it did |
|---|---|
| Java, gson | 86.7 % -> 92.9 % answered, coverage unchanged at 95.8 %, precision 0.9996 -> 0.9986 |
| Go, cobra | coverage 86.5 % -> 87.7 % at precision 1.000 |
| ArkTS, photos | coverage 73.9 % -> 77.6 %, precision 1.000 -> 0.9992 |
| Rust, clap | 52.0 % -> 57.4 % answered, coverage 84.3 % -> 84.8 % |
| C++, fmt | REFUSED — precision 0.969 -> 0.848 for four tenths of a point of coverage. A template's receiver is written as a parameter and the method IS in the tree, so 86 answers became wrong to gain nothing |
| Ada, ada-util | REFUSED — 4 wrong answers for 1.2 points of answered share, against a precision of 1.000 this engine publishes for Ada |
| C | REFUSED — no receiver types to read, so it could only add a way to be wrong |
C# is measured too, now. The .NET SDK is not in the container image; it
installs into /var/tmp/dotnet with Microsoft's own script and the Roslyn
checker scripts/csharp_callgraph builds against the 8.0 runtime. See
"Measuring C#" below. With the rule, Newtonsoft.Json answers 69.5 % of its
calls against 66.0 % without, at identical coverage (71.9 %) and a precision of
0.978 against 0.985. Three and a half points of a reader's picture for 89 more
wrong answers out of 19 234 judged — taken, and published so the trade is
visible.
A scope the source wrote, and no import to follow¶
Command::new(name) in a file that imports no Command. The source WROTE the
scope, so the only open question is which Command it means — and where the tree
declares exactly one member of that name under a scope of that name, there is
nothing to choose between. A second declaration refuses the answer, which the
test pins.
Only where the separator is ::. In Rust, C++ and Ada that spelling says the
segment before it is a path and never a receiver. x.bar() in Java writes a
local and a type the same way, so allowing it there would turn evidence into a
guess.
This is how one crate of a workspace reaches another, and the case that made it
necessary has no import at all: clap's root crate publishes its API with
pub use clap_builder::*;, a glob re-export of a whole other crate. clap's
tests, benches and examples write clap::Command::new and clap::Arg::new 668
times, and each one is the ROOT of a builder chain — so every later link
(.arg(), .short(), .about()) failed with it.
clap: 4 930 -> 6 559 in-tree edges, 57.4 % -> 64.6 % answered
coverage against rustc 84.8 % -> 85.3 %, precision 0.967 -> 0.961
Go cobra 87.7 % -> 88.6 % coverage at precision 1.000
The gain is mostly where the oracle cannot see it, because
scripts/corpus_rust.py builds one package (clap_builder) and these calls are
in other crates' test and example targets. So ten of the 1 198 new edges were
read by hand against their source lines — clap::Arg::new("flag") reaching
arg.rs::Arg::new, .short('a') reaching Arg::short — and all ten were right.
That is a sample, and it is said here as a sample.
Measuring C¶
The container image has no .NET runtime, which is why this row was stale for weeks. It does not need one installed system-wide:
curl -sSL https://dot.net/v1/dotnet-install.sh -o /tmp/dotnet-install.sh
bash /tmp/dotnet-install.sh --channel 9.0 --install-dir /var/tmp/dotnet
bash /tmp/dotnet-install.sh --channel 8.0 --runtime dotnet --install-dir /var/tmp/dotnet
export PATH=/var/tmp/dotnet:$PATH DOTNET_ROOT=/var/tmp/dotnet
cd scripts/csharp_callgraph && dotnet build -c Release -o /var/tmp/muundo/ls/csgraph
python3 scripts/corpus_score.py csharp /var/tmp/muundo/corpus/newtonsoft ./target/release/muundo
Both runtimes. The SDK channel gives 9.0; scripts/csharp_callgraph targets
8.0 and fails with a dotnet-core-applaunch URL rather than a plain message when
only 9.0 is present. That URL is the missing runtime, not a broken build.
Python: what a name holds¶
Python writes no types, so the value a name holds is the only evidence there is
for a call on that name. The Python parser filled no value bindings at all,
while Rust, TypeScript and Kotlin all did — so a = MyClass() then a.func(),
the most ordinary pair of lines in the language, reached nothing.
python::collect_bindings records both forms, because they answer different
questions.
A CALL says what type the name holds. a = MyClass() makes a.func() reach
MyClass.func. It reuses the fact and the consumer the other three languages
already had; nothing new was needed below it.
A plain NAME says the name IS that function. handler = do_thing then
handler() calls do_thing. That needed a rule of its own, because the existing
consumer reads a binding as "a type", and here the binding is a function. It runs
after every rule that looks for a declaration actually named handler, so a real
function of that name wins over an alias to one. Three shapes are bound:
a = b = f, a, b = f, g paired by position, and the plain single name. A
starred target (a, *b = …) is left alone — the positions after the star mean
something else.
A parameter holds what its callers pass. is_binary(stream) inside a
function whose every caller passes one function for that position reaches it.
Every caller, and a plain name from each: one that passes a lambda or an
expression refuses the answer outright, and two different functions refuse it
too. click's _force_correct_text_stream is the case that shows the refusal
working — its two callers pass _is_binary_reader and _is_binary_writer, so
the parameter holds either and muundo says nothing.
Python only, and that is measured. Elsewhere the parameter rule paid nothing and cost: excalidraw gained no coverage and answered 12 more calls wrongly, fmt one. A callback in TypeScript comes from too many places for one name to be the answer.
PyCG corpus: 35.2 % -> 53.0 % coverage, precision 0.989 -> 0.984
click, its own run: 59.4 % -> 63.7 % coverage, precision 0.804 -> 0.809
click, answered: 77.5 % -> 80.8 %
Where the four steps landed: the bindings took 35.2 % to 44.8 %, the alias to
49.4 %, the parameter rule to 50.6 %, and dropping PyCG's nine <builtin>
expectations — which no in-tree edge can ever match — to 53.0 %.
What is left, and why. PyCG's dicts and lists categories hold a function
in a container and call an element (a = [f, g] then a[0]()): 25 calls, and it
needs element tracking. mro and the rest of classes are a method found
through inheritance or bound to an attribute (self.child = self.func2). And a
__getattr__ lookup cannot be answered by reading source at all — click reaches
detach that way.
Five holes in the surfaces, and the one review comment that was wrong¶
A review of the MCP server, the HTTP server and the snapshot store on 2026-09-27. Four of the five findings held as stated; the fifth was right about the problem and wrong about the fix. None of them touches call resolution — every language was re-measured afterwards and every figure is unchanged.
A root the boundary never saw. Every analysis tool hands its root to the
engine with containment_root set, and the engine refuses a file outside it. Two
MCP tools do not go through the engine at all: verify replays a report, and
query kind=freshness re-hashes a manifest. Both took the caller's root
verbatim. With MUUNDO_WORKSPACE_BASE=/srv/app, a root of /etc had its files
hashed and the answer named which ones differed — an existence and digest oracle
over anything the process can read, which is the exact thing the module's own
documentation says the boundary prevents. tool_query did not even receive the
limits. The HTTP server has always confined both paths, so the fix is that check,
in the other surface, canonicalized before the comparison so a symlink inside the
boundary cannot point out. The refusal names neither the boundary nor whether the
path exists: two sentences would make the refusal itself an oracle.
A wider base the caller chose. verify takes a workspace_base — the
containment root a report produced by a server records, wider than the analysed
tree. Taken from the caller it is the same hole by another name. Confined, it is
the boundary and nothing else; a caller naming anything different is refused
rather than quietly narrowed.
A replay bigger than its reservation. A verification takes the memory permit
an analysis takes, and the replay IS an analysis — but both surfaces passed
VerifyLimits::default(), so the replay ran up to 8 GiB under a reservation
sized for the surface's own per-analysis cap. VerifyLimits now carries replay
ceilings of its own and each surface fills them from its admission.
Where the review was wrong. It also asked for the replay to be REFUSED when
the report's recorded ceilings exceed the surface's. A report records the
ceilings it ran under, not the bytes it read: a two-file tree records the same
512 MiB as a two-million-file one. Implemented as written, it refused
verification of a two-file tree, and a_large_config_appearing_is_a_changed_tree_not_an_exhausted_budget
caught it. The clamp alone achieves what the finding was about — the replay
cannot read past the reservation — and a replay that does cross it stops with a
workload error the verifier owns.
An id taken on trust. Store::save used report.snapshot_id as the storage
key when the document carried one. The id is a digest of the analysis's inputs,
so it is an address the content derives; snapshot save --report accepts a
document from anywhere, and one could be filed under any address, including an
address already taken — where reused: true was answered without looking, so the
caller was told their report was kept while a different one stayed and answered
every later question. The id is computed now, always; a claimed one that differs
is invalid_report. And because the id digests the INPUTS and not the findings,
two reports can carry one id and different conclusions: that is an integrity
error naming the fields that differ.
Hashing that could not be stopped. freshness and the verification hashing
loops took no cancellation token, so a timed-out or abandoned request left a
worker hashing a tree with the permit its caller had reserved — repeat it and the
surface runs out of workers. The token is checked inside the read loop, because
one file can be the whole workload. On the HTTP surface freshness now takes the
analysis permits and the cancel-on-drop lifecycle, with the permits moved INTO
the blocking closure so capacity is released when the thread exits rather than
when the request does. The other query kinds still read an index and answer in
milliseconds with no permit, which is why the path is split.
An index file read whole. load_index did fs::read and then
from_slice, with none of the ceilings every other report entry point goes
through, and it never checked that the index names the snapshot asked for — one
that names another answered every question about the wrong tree. It streams under
a cap now and validates the id, falling back to rebuilding from the report, which
is read under the engine's own limits. The byte cap is what bounds the structure
here: SnapshotIndex has no recursive field, so its node count cannot exceed its
size.
Four follow-ups on those five, and a number that was measured¶
A second pass over the same surfaces on 2026-09-27, all four confirmed. Two of them are consequences of the first round's own fixes.
One pair of ceilings, not two. An analysis and a verification's replay are the
same walk over the same tree on the same server, so a ceiling that binds one and
not the other is a hole. The MCP's replay inherited the verification defaults —
200 000 files — and took its per-file cap from the AGGREGATE byte cap, where an
ordinary analyze on that same server stopped at 50 000 files and 2 MiB. Limits
carries both numbers now and base_options writes them out rather than inheriting
them, because verify_limits reads the same two fields: left implicit, the two
paths took their ceilings from different places and drifted without a line saying
so. Hashing keeps the published manifest ceiling — it streams a file and holds
none of it, so a manifest longer than a walk may be re-read.
A ceiling of ours must not read as a no. A file over the replay's per-file
cap is SKIPPED by the engine, not refused, so the replay finished, its
skipped_files differed from the report's, and the verdict came out
verified: false — the tree blamed for our ceiling. Reachable only because the
first round lowered that cap on the servers. It is provable before the replay
runs: every path in file_hashes was READ by the analysis that produced the
report (a skipped file carries no digest), so one of them larger than what the
replay may open cannot be read again. The size is counted while hashing, where the
budget already decrements by exactly the bytes read, and the refusal names the
file and says the report is not at fault.
Three answers at an id that is taken, not two. reused: true was returned for
an object that could not be loaded at all — an interrupted save leaves one — so
the caller was told their report was kept while the store held something
unopenable, and every later question about that snapshot failed. The first round's
own comment said such a file was "a file to overwrite" and the code did not
overwrite it. It does now, and answers reused: false.
An index question reserves what it parses. It answers from a file the server wrote, so it runs no analysis and took no permit at all — while parsing that file into memory, with a cap of 512 MiB fixed in the core. Several questions at once each held a copy, so a host given a small budget could be asked the same question eight times. The cap is now the per-analysis source ceiling admission already resolved, so it falls with the host's budget, and each question reserves weighted units from the same semaphore, held until the thread exits so it covers the parse, the traversal and the encoding.
The ratio is measured. INDEX_RESIDENT_RATIO is 2, from one measurement: an
index of muundo's own core/src is 6 607 098 bytes of JSON and 6 269 440 bytes
resident once parsed — 0.95x, because the shape is mostly strings and a string
costs what its bytes cost. The reader streams, so the text is never held beside
the structure. Two copies are reserved: the parsed index, and the answer built
and encoded from it.
C#: the namespaces that enclose the caller¶
C# resolves an unqualified type name through the namespaces that ENCLOSE the
caller's own, innermost outward. A file whose namespace is
Newtonsoft.Json.Tests.Serialization sees Newtonsoft.Json.JsonConvert with no
using at all, because Newtonsoft.Json is an ancestor of its own namespace.
Nothing here read that. The file holding 1 246 calls into JsonConvert.cs
imports Newtonsoft.Json.Linq, .Converters, .Serialization and .Utilities
— and never Newtonsoft.Json itself. It does not have to. Three rules already
looked at a namespace: the one the CALL writes out, the one an import names, and
the one a using names. The caller's own was the fourth, and it is the one a
test file relies on.
Newtonsoft.Json: 71.9 % -> 85.5 % coverage, precision 0.978 -> 0.981
answered: 69.5 % -> 76.1 %
Precision went UP, which says the answers are the ones the compiler names.
C# only, and that is the language rather than a measurement. Java and Kotlin
packages are flat: a.b.C is invisible from a.b.d without an import, so the
same walk there would answer from a namespace the compiler never looks in.
The walk goes outward and nowhere else. A sibling namespace is not an ancestor, and a test pins that it stays out of scope.
One name under one scope, in two files¶
A candidate set can hold declarations with DIFFERENT qualified names: the same member name under one scope, declared in two files. That is a C# partial class, and Newtonsoft.Json splits its readers and writers into a synchronous and an async half exactly that way. Asking for a single candidate refused between them.
one_of is the tie-break, used wherever a scope, a file or an import has already
narrowed a call to a set: the count the call passes picks one, and where two take
the same count nothing is claimed. It adds a tie-break and never a search —
preferring a base class's exact count over a declaration the type itself makes is
the mistake member_fitting_the_call documents, and it cost 697 calls when it was
tried.
Worth 23 more named calls on gson for one wrong (95.8 % -> 96.0 %), and it is what makes a partial type reachable at all.
What it is not. An overload set does not need it: two overloads of
Serialize carry ONE qualified name here, so the set has one member and the
ordinary path answers. Counting declarations rather than names made overloads look
like the dominant cause of C#'s misses — 3 370 of 5 105 — and they were not the
cause at all.
An abstract class is a class¶
tree-sitter gives abstract class X {} a node of its own, and the TypeScript
extractor listed only the ordinary one. So an abstract class was no entity at
all, and its members came out named after the FILE instead of after their class
— which made a static call on it reach nothing.
ArkTS writes export abstract class MathUtils, and that file is the single
biggest target in the HarmonyOS corpus measured here. Five places had to learn
the node: the shared scope list, the entity extractor, the inheritance reader, the
owning-type walk and the list of nodes that declare a name.
ArkTS photos: 77.7 % -> 82.2 % coverage, precision 0.9985 -> 0.9993
TSX excalidraw: 86.88 % -> 86.91 %
#field is a field¶
The same shape of omission one level down. A private field's name node is a
private_property_identifier, which neither the typed-name collector nor the
field collector accepted — so a class's private fields had no written type and
this.#options.throwHttpErrors(…) arrived as the bare word throwHttpErrors.
The name is kept AS WRITTEN, # and all, because that is what the callee text
the resolver reads contains. An optional call through it (?.()) resolves for
free once the field is known.
ky writes this.#… on 251 lines of one file, and the fix is worth one edge
on the four trees measured — ky's remaining gap is elsewhere (below). It is here
because this.#x.m() was unreadable, not because of what it scored.
super(…) reaches the base class¶
It is a bare call named super, and nothing declares super: every rule for a
bare name looked for one and found nothing, so a constructor delegation carried
no answer at all. super.method() was already handled — the receiver names this
tree — but super(…) names no member, only the base.
The constructor's spelling is per language, which is why it is a table and not a
name: TypeScript writes constructor, Python __init__, and the JVM, C# and C++
repeat the type's own name. Rust and Go have neither a base class nor super. One
base and one constructor under it, or nothing is claimed — multiple inheritance in
C++ makes super meaningless anyway.
Worth 14 calls on the ArkTS corpus and 4 on excalidraw.
What is left in ky, and why it is not cheap¶
ky sits at 67.5 % and its remaining misses are one shape: 531 of 1 128 are
text and json, expected at source/types/ResponsePromise.ts. Its tests call
ky('', {…}).text() and ky.get(url).json(), where ky is a value typed by a
CALLABLE INTERFACE — const ky: KyInstance = createInstance() — and text is a
method signature on the interface the call signature returns.
Placing those needs TypeScript's call-signature resolution: what ky(…) returns
is written on the type of ky, not on any function declaration. That is a
different mechanism from everything above, and it is named here so the number is
not mistaken for a gap nobody has looked at.
C++: a template type is a type, and a cast is not a call¶
C scores 100 % and C++ scored 54.5 %, which is a gap no amount of "C++ is harder" explains on its own. Two of it were real bugs.
first_of_kind reads DIRECT children. plain& p; puts a type_identifier
right there and worked; buffer<char_type>& buffer_; puts a template_type
where the name should be and was skipped entirely — so the field had no recorded
type and every call on it was unplaced. In C++ that is most fields.
bare_type_name now descends into a template_type the way it already did for
every other language's generic wrapper, and the C++ collector accepts the node.
A cast is written like a call and is not one. static_cast<char_type>(ch)
parses as a call whose function is named static_cast, so every cast in a C++
tree produced an edge to a function nothing declares. Worse than useless: the
comparison pairs answers per POSITION, and fmt writes 265 casts on lines that
also hold a real call, so the false edge took the place the real one needed. The
four casts, sizeof, alignof, typeid and decltype transfer control nowhere,
which is the only thing a call edge means.
fmt: 54.5 % -> 57.9 %
And then the ceiling, which is not the resolver. fmt has 750 syntax errors
across 39 files against leveldb's 116 across 28. class buffer in core.h is
recorded as lines 1759-1773 because the parser stopped at the first constructor it
could not read; push_back is at 1829 and is therefore no entity at all. Three
macro shapes were masked on a rewritten copy — the namespace macro, the SFINAE
macro in a template parameter list, the pragma macro at namespace scope — and 750
became 658. The rest is spread through the grammar's handling of template
metaprogramming rather than concentrated anywhere a mask could reach. The numbers
are in docs/MEASUREMENTS.md; what the engine already did right is report it, in
every report, per file.
One guard earned its keep. Teaching the TypeScript family about
abstract_class_declaration broke
cyclomatic_scope_chain_equals_the_parser_scope_kinds immediately: the metrics'
scope list and the parser's must be the SAME set, or metrics silently stop joining
their entities in any file that has one. It is the second list that had to learn
the node, and nothing but that test would have said so.
A qualified name is not an overload¶
Src/Newtonsoft.Json/JsonConvert.cs::JsonConvert::SerializeObject is eight
declarations, at lines 547, 563, 577, 596, 617, 639, 659 and 682. muundo places a
call on it and cannot say which. Protocol v2 stopped letting the oracle pick one on
the engine's behalf — that was credit for precision the engine does not have — and
the measurement now counts those apart instead of hiding them in either direction:
C# Newtonsoft.Json: 6 371 missed, 5 077 of them overload-only (80 %)
Java gson: 3 015 missed, 2 577 of them overload-only (85 %)
Kotlin KotlinPoet: 448 missed, 251 of them overload-only
C++ fmt: 2 025 missed, 379 of them overload-only
So overload_unnamed is a MISS — a call graph that cannot name the declaration has
not placed the call — and it is reported, because it is a different failure from
answering nothing: the type and the method name are right. DECLINED in the Rust
harness holds it beside undetermined and unresolved so the three cannot drift
apart, and scripts/test_corpus_protocol.py pins that it stays a miss.
This is the largest single lever left in the engine. Every language with
overloading pays it, and no rule in the resolver can fix it: the report's qualified
names would have to carry the arity or the signature, which changes
REPORT_SCHEMA_VERSION and everything that reads a callee. Named here so the next
piece of work starts from it rather than from another resolution rule.
C# reads its enclosing namespaces for a BARE name too¶
The rule added for JsonConvert.SerializeObject was wired into the qualified path
only. A bare new JsonTextReader(sr) in Newtonsoft.Json.FuzzTests reaches
Newtonsoft.Json.JsonTextReader by the same walk — that file imports System and
System.IO and nothing else.
Newtonsoft.Json: 59.6 % -> 65.9 %, precision 0.987 -> 0.989
Precision went up again, and the unplaced · bare bucket fell from 2 186 to 693.
A file can publish a name it does not declare¶
export { renderApp as render }; — the file's render IS its own renderApp. An
importer asked the file for render, the file declared no such thing, and the call
reached nothing. excalidraw's test helper is exported that way and called 266
times: TSX 85.5 % -> 87.1 %.
Two shapes, and the presence of a from decides which:
export { a as b }publishes a LOCAL name under another. That is a third map onImportBindings—exported_as— becausesymbolsmeans "this name lives somewhere else" and this means "this name is mine, under another spelling".through_reexportsconsults it: the name changes and the file does not.export { a as b } from "./c"republishes ANOTHER file's name, which is the same thing as importing it and exporting it — so it goes insymbols, where that already means what it needs to. It did not resolve either, and a test now pins both.
And a file whose only binding is what it PUBLISHES was left out of the index.
The condition that decided whether to keep a file's bindings asked about symbols
and receivers; a file that imports nothing and exports an alias has both empty,
so the alias was collected and never consulted. That is the whole bug in one line,
and it is the kind that a new field always has: two places to add it, and only one
is where the reader looks.
Why the overload lever was measured and not taken¶
The overload gap is 80 % of C#'s misses, so it was the obvious next piece of work. Measured first:
qualified_name+ the declaration's ARITY separates only 161 of 323 overload sets.JsonConvert.SerializeObjecthas eight declarations at arities 1, 2, 2, 3, 2, 3, 3, 4 — so a call passing two or three arguments, which is most of them, stays ambiguous.- Separating the rest needs the TYPES of the arguments at the call site. That is type inference, and it is the reason the oracle needs a compiler.
So the reachable part is a few hundred calls for a REPORT_SCHEMA_VERSION bump and
a change to everything that reads a callee. Left undone, deliberately, and the
number stays visible in its own column so nobody mistakes it for a gap nobody
looked at.
What the same measurement said instead: separating the overload misses from the rest re-ranks the work by REAL resolution misses — excalidraw 2 198, C# 1 294, ky 1 219 — and the TypeScript family led it.
A type can be called¶
type KyInstance = { (url: string, options?: Options): ResponsePromise; get: …; }
— a TypeScript object type may declare a CALL SIGNATURE. ky('url') matches it,
the signature has no name, so nothing was reached and every call chained onto the
result stopped there. ky's own tests are written that way throughout.
ky: 62.5 % -> 70.4 % coverage, precision 0.9946 -> 0.9952
The return type is read off the TYPE, not through types_returned_by. That
helper answers "constructing this yields itself" for a type, which is right for
new Thing() and wrong here: a call signature yields what IT returns.
Three hops, because a value gets its type three ways. Written for it
(declare const c: Client) is the easy one. ky's own shape is the other: ky('')
reaches source/index.ts::ky, which is const ky = createInstance(); — a const
declares no return type, so the chain died on the first link. What is left was
already recorded and just never joined up: what filled the const, what THAT returns
(KyInstance), and what calling a KyInstance yields. The binding is read in the
file that declares the const, not in the caller's, and a file that binds the name
twice refuses rather than guesses.
What is still out of reach there. ky's json is json: { … } — a nested
object type, not a function-typed property — so it is no entity and 206 calls to it
cannot resolve whatever the chain does. text is a plain () => Promise<string>
and it is the 325 this bought.
A C# extension method, and the half that makes it safe¶
"{0}".FormatWith(provider, arg) reaches a static method of a static class whose
FIRST parameter is this string format — so the receiver is a string and the
declaration is nowhere near it. No rule that looks inside the receiver's type can
find it. Newtonsoft.Json calls that one 333 times.
BOTH HALVES OR NOTHING. The tree says which names ARE extensions (the this
on a first parameter, which the parser now collects with the type it extends), and
the receiver's TYPE says which of them a call can reach.
Matching the name alone was tried and measured first: 378 calls went to the wrong
declaration for 138 right ones, precision 0.989 -> 0.961. The reason is in the
corpus: LinqBridge.cs REIMPLEMENTS LINQ, so ToList, Select and Where are all
declared there as extensions while nearly every call to them reaches the framework.
A name being an extension somewhere says nothing about what this call reaches.
Then the rule was moved after every receiver-type rule, in case the order was the problem. The numbers came back identical to the character — which is how one learns it was the rule and not its position.
With the type matched: a string LITERAL answers for itself, because
"{0}".FormatWith(…) writes the type in its first character, and that is most of
these calls. Otherwise the type the source wrote for the name.
Newtonsoft.Json: 405 `FormatWith` edges placed
the right METHOD: 92.1 % -> 94.3 %, for 36 more wrong edges
the right OVERLOAD: 65.9 % -> 66.3 %, because `FormatWith` has five
That last line is what the two columns are for: the calls are placed on the right method, and which of its five overloads is a question nothing here can answer.
A wrapped arrow belongs to the field that holds it¶
private beforeUnload = withBatchedUpdates((event) => { this.load(); });
The arrow is an ARGUMENT of a call, and the call is the field's value — so nothing
above the arrow is a function and the walk for the enclosing callable ran out.
Every call in there came back owned by <module>. Two costs, and only the first
shows in a coverage figure: this.load() could not find its class, and fan-out and
the hotspot scores counted the work against the file instead of against
beforeUnload. It is the wrapper idiom — throttle, debounce — and class-based
React is built out of it.
excalidraw: 87.1 % -> 87.5 %, precision unchanged at 0.994
An arrow still borrows no name of its own. expression_callable_name refuses an
inline callback on purpose, so the walk can reach a real enclosing function first;
this only answers once that walk has found nothing at all.
Two conditions, each for its own reason, and a guard found both. The first
version named any binding, and
the_call_graph_of_this_tree_joins_to_its_own_entities failed on five edges whose
caller named no declaration — in muundo's own scripts:
const skip = new Set(["node_modules", …]);crosses NO anonymous function, so the fallback must not fire at all. It now requires having actually crossed one.const owned = new Set(files.map((f) => resolve(f)));holds a Set, not a function. So only a CLASS FIELD is borrowed: the entity extractor records one whatever its value is, which keeps every caller a declaration the report carries. A localconstgives no such guarantee, and filingresolveunderownedwould make fan-in count against something nobody declared.
excalidraw kept the whole gain after both conditions, which is what says they cost nothing: it came from class fields throughout.
And it must be asked of the arrow, not of the call. Starting the search at the
call site named syncableElements — the local const on the line — instead of
beforeUnload. The owner is whatever holds the OUTERMOST anonymous function the
walk passed through.
A dot in a file name is not a dot between modules¶
import { AppCursor } from "./App.cursor";
excalidraw cuts its 14 000-line component into App.cursor.ts, App.viewport.ts,
App.pan.ts and a dozen more, and keeps App.tsx beside them. Two readers took
the .cursor for a file extension, and each broke something different:
- the import resolver called
Path::with_extension, which REPLACES — so./App.cursorwent looking forApp.ts, a real file two lines away. The comment above that loop said "try appending each extension"; the code did not. - the import index split the specifier on its dots and filed
cursoras "a module this file imports".this.cursor.set(…)then looked like a call on that module, an arm that runs BEFORE the one reading the field's declared type — sopublic cursor: AppCursorwas never consulted.
The second is the one that cost: 94 calls on this.cursor and this.viewport
alone, and it fires for any field whose name matches the tail of a file name the
same file imports.
excalidraw: 87.5 % -> 88.5 %, precision 0.9936 -> 0.9937
photos (ArkTS): one more call named
the eleven other corpora: unchanged, to the call
A dotted MODULE path is still split. os.path, java.util.List, Ada's
Util.Log: there the dots separate names, and the leaf is the last one. The test
is whether the target is a PATH — if it holds a /, its last segment is one
name, dots and all.
Appending comes before replacing, not instead of it. TypeScript's ESM
spelling asks for ./foo.js when the file on disk is foo.ts, so replacing is
still needed — it is simply the second question, which is also the order tsc
asks them in.
An object of helpers keeps the name it is bound to¶
const Regex = {
build: (regex: string): RegExp => new RegExp(regex, "u"),
join: (...parts: RegExp[]): string => parts.map((x) => x.source).join(""),
};
A module of helpers written as an object rather than a class. Each function was
named after its key alone, so the entity was textWrapping.ts::build while every
call writes Regex.build(…) — and the receiver Regex scoped nothing of that
name. 66 calls in that one excalidraw file reached nothing.
The decision was already taken one level down: nested keys are kept, so
{ tenants: { list: … }, workspaces: { list: … } } gives two names and not one.
This carries the same walk one step further, to the name the whole object is
bound to. Only a const X = … with a plain identifier: a destructuring pattern
binds several names and none of them is the object.
excalidraw: 88.5 % -> 88.8 %, precision unchanged at 0.9937
the thirteen other corpora: unchanged, to the call
A name may now carry its own path, and every index asks for the short one.
Regex::build is what the entity is called — one value builds both the entity
and the caller of its calls, which is why the path lives in the name. The
lookups the resolver starts from want build, the word a call writes after the
receiver, so each entity is indexed under its name AND under the last segment of
its qualified name when the two differ.
A bare build(…) still reaches no property. A callable bound to a key is a
method, and in the JavaScript family a bare call never reaches one — it needs
the receiver in the source, exactly as a class member does.
A barrel republishes what it lists, and sometimes it lists nothing¶
A folder's index.ts that publishes its neighbours' names, so an importer writes
one path instead of five. Two spellings were not followed, and both are the
common ones:
export { getNormalizedZoom } from "./normalize"; // no `as`
export * from "./normalize"; // no names at all
The first was read only for its alias, so export { work as doWork } from
'./deep' was followed and the plain form was not. The second lists nothing, so
which republished file holds the declaration is knowable only by looking: each
one is asked, and only a file that really declares the name answers. A barrel
that republishes four modules gives no clue on its own, and a guess there would
be a wrong edge rather than a missing one.
excalidraw: 88.8 % -> 89.5 %, and the right METHOD passes 90 % (90.1 %)
ky: 70.4 % -> 74.1 %, precision 0.9952 -> 0.9955
the twelve other corpora: unchanged, to the call
ky is the larger jump for its size: it publishes its own surface through one
index.ts, and its tests import everything from there. 122 calls, and not one
new wrong answer.
A property assigned a function is under its object¶
function abort(what) { … } // line 849
Module.abort = function () { … } // line 3845
Two declarations in one file, and muundo called both woff2-bindings.ts::abort.
One qualified name over two declarations means a bare abort(…) can choose
neither, and the report merges their fan-in. 99 calls in excalidraw's bundled
woff2 decoder, which is generated code and writes every export that way.
The assigned property now carries its object, exactly as an object literal's keys
carry theirs: Module::abort. this.updater = … keeps only updater, because
the class it is written in already scopes it, and anything else on the left — a
call's result, an index — still gives the property alone.
excalidraw: 89.5 % -> 90.2 %, and the right METHOD 90.8 %
express: 86.1 % -> 86.3 %
the eleven other corpora: unchanged, to the call
What it costs, measured. 125 more calls named, and 9 more wrong: 2 of them in excalidraw, and in ky 2 wrong targets plus 3 answers claimed in-tree that are not. ky's are one shape, and it is worth naming because the rule did not invent it —
response.clone = () => { … }; // line 697, inside one test
response is a local of that test. Twenty-seven lines earlier a DIFFERENT test
writes response.clone() on its own local, and it now reaches the assignment
made in the other one. The receiver rule matches a scope segment by name, and a
variable's name is reused across a file in a way a class's is not. Precision:
excalidraw 0.9935 -> 0.9932, ky 0.9955 -> 0.9942.
An indexed access type names a member, so read it as one¶
const LayersFieldset = ({ renderAction }: {
renderAction: ActionManager["renderAction"];
}) => <fieldset>{renderAction("sendToBack")}</fieldset>;
ActionManager["renderAction"] is the source naming a declaration — the type and
the member, both written. It is how a method reaches a component that must not
know the object it came from, and 77 of excalidraw's calls are one. The rule runs
FIRST among the bare-name rules, because a file of its own may well declare a
function of that name and the annotation is the stronger evidence.
A destructured parameter needed reading at all. Its annotation is ONE object type covering several names, so read as a single type it had none — and every name it binds had none either. The keys now come from the pattern and the types from the object type, matched by name, so a name the annotation does not mention gets nothing rather than its neighbour's type.
excalidraw: 90.2 % -> 90.8 %, right method 91.4 %, precision 0.9932 -> 0.9931
the thirteen other corpora: unchanged, to the call
T[keyof T] and T[number] name no single member and are refused: the index has
to be a string literal.
T | null is still T¶
A nullable union is how TypeScript writes "this may be absent", and it says the type as plainly as the bare form. Read as "no type at all", every call on such a name was lost. Two REAL members name two possible declarations and are refused.
excalidraw: 15 168 -> 15 170 calls named, precision unchanged
Two, not six hundred. 596 annotations in excalidraw are a nullable union and
almost all of them name DATA — string | null, AppState | null. The change is
right and costs nothing; the figure is what it is, and stating 596 as if it were
the gain would be the mistake this line exists to avoid.
Why an indexed-access interface member is NOT a declaration¶
export interface CollabAPI {
isCollaborating: () => boolean; // a declaration
getUsername: CollabInstance["getUsername"]; // not one
}
Eleven of excalidraw's twelve CollabAPI members are written the second way, and
only the twelfth is a declaration muundo reports. Accepting the others looks like
the obvious completion of "an interface's members are declarations".
It was tried and measured: not one more call named, and excalidraw's precision
fell from 0.9931 to 0.9914. Whether the member an indexed access names is
callable is not knowable from that file, so theme: AppState["theme"] became a
member too — a few hundred entities crowding the name index, for nothing.
Reverted; the reason is written on declares_a_function.
The head of a receiver path may BE the type¶
LocalData.fileStorage.getFiles(fileIds)
type_of_the_path walks a receiver one field at a time, and it started by asking
the calling file for the written type of the first name. Here there is no name to
ask about: LocalData is the class, and fileStorage is its static field. The
head is now looked up among the names the tree declares when no local one has a
type — in that order, so a local shadowing a class keeps its own answer.
excalidraw: 15 170 -> 15 176 calls named, precision unchanged
Six, because most of excalidraw's field-path misses start somewhere else. The walk itself is the check: it continues only while each segment is a declared field of the type before it.
A browser global carries its type, and an inline type gets a name¶
// App.tsx
declare global {
interface Window {
h: { app: InstanceType<typeof App>; scene: Scene; store: Store };
}
}
// every test file
const { h } = window;
h.app.actionManager.executeAction(action);
How a TypeScript project declares a browser global. excalidraw's tests reach the running editor through this one, and 214 of its calls were that shape — the single largest block left. Three things were missing and the chain needed all three:
- a member of
interface Windowis a global. That is what givesha type at all; nothing else in a test file says what it holds. - an inline object type has no name, so it is given the path of the member
that holds it:
Window::h.appis then a member of THAT and not ofWindow— read as a member ofWindowit would answer a bareappjust as readily.::cannot appear in a TypeScript type name, so a synthetic path never collides with one the source wrote. InstanceType<typeof App>isApp. The generic walk took the first child and answeredInstanceType, a name nothing declares. Only that one utility is read this way, by its own name.
The walk itself was already there — type_of_the_path steps one field at a time
— and it stops at the first segment that is not a declared field of the type
before it.
excalidraw: 90.9 % -> 91.9 %, right method 92.5 %, precision unchanged at 0.9931
the thirteen other corpora: unchanged, to the call
171 calls named and not one new wrong answer, which is what a rule reading only declared types should cost.
Python unpacks what it assigns, and a call sees the nearest write¶
a, (b, (c, d)) = func1, (func2, (func3, func4))
a, *rest, c = func1, func2, func3, func4
f, g = c, d = func1, func2
A dispatch table written as an assignment, which is how Python writes one. The
pairing read a single flat pattern_list against a single flat expression_list,
so a nested target bound nothing at all, and it refused a starred target rather
than get it wrong. Three things it now gets right:
- nesting, by pairing recursively;
- a starred target:
*resttakes whatever is left over, so the names AFTER it are counted from the END —cis the last value and not the second. The starred name itself holds a list, which is not a name and stays unanswered; - where a chain ends:
f, g = c, d = func1, func2puts an assignment on the right, and what fills the outer targets is what fills the inner ones. The inner targets are bound when the walk reaches that assignment itself.
A mismatch binds nothing. Three names against two values is a program that raises at run time, and a guess there is an edge nobody wrote.
And which write the call sees. A name bound twice in one scope holds what the
nearest line ABOVE the call put there. binding_at chose the narrowest region and
then the FIRST match, so a = b = func1 … a = b = func2 … a() reached
func1 — a write that call never saw. Narrowest region first, latest line second.
PyCG: 49.6 % -> 53.2 %, precision 0.9843 -> 0.9926
the thirteen other corpora: unchanged, to the call
Nine calls on a 252-edge corpus, and one fewer wrong answer: the nearest-write rule removed a claimed edge as well as adding real ones.
Python reads a dispatch table¶
d = {"a": func1, 1: func2} a = [func1, func2, func3]
d["a"]() a[0]()
How Python writes dispatch when it does not write a class. The extractor keeps the
container's name and drops the subscript, so each of these arrived at the resolver
as the bare word d and reached nothing — two whole categories of the PyCG corpus.
The key is read back out of the callee text the resolver already has, so the
container and the slot need no table of their own. str:a and int:1 are kept
apart, because d = {1: func1, "1": func2} is a real line and d[1] reaches the
first of them.
A slot may be written more than once — the literal, then d["a"] = other, or
b = [None] then b[0] = func4 — and the call reads the NEAREST write above it.
PyCG: 53.2 % -> 57.5 %, precision 0.9926 -> 0.9932
And a starred target holds a list. a, *rest, c = f1, f2, f3, f4 puts f2
and f3 in rest, read back as rest[0] and rest[1] — the one target that
holds a container rather than a name.
Three refusals, and one of them is measured. A computed key (d[which]) names
no slot. A slot filled with a call's RESULT holds what the function returned, not
the function. And a container something called a METHOD on has stopped being
written down: d.update({"a": func2}) puts a different function in that slot, and
answering with the literal was the one wrong edge this rule added. Reading what
update did is a different analysis; knowing that something did is enough to stop
answering. That guard cost no named call at all.
A name bound to a bound method¶
a = MyClass()
b = a.func
b()
a.func() written out resolves already. Given a name first it did not: b = a.func
is not a call, so nothing was placed on that line and no rule that reads a receiver
could see one. The chain stage now rewrites such a callee to the expression the name
holds and carries on — keeping the CALL's own line, so it still reads what was
placed above the call.
PyCG: 57.5 % -> 59.1 %, precision 0.9932 -> 0.9933
Only where the object is a plain name. make().func is an expression this reader
does not follow, and half-following one is worse than not.
The same line appears inside unpacking — c, (d, e) = a.func1, (a.func2, a.func3)
— which is how the corpus writes it and where three of the four gained calls are.
A for loop is two calls, and the source writes neither¶
for i in c:
Python runs c.__iter__() and then c.__next__() until it raises, and a class
may define both. muundo emitted no call at all, so a user's own __iter__,
reached by a loop and nothing else, looked dead to every reachability question
the report answers — which is a worse defect than the missing coverage.
The two calls are written as <the iterable>.__iter__ and .__next__, so every
rule that resolves a receiver already resolves them: a construction, a name bound
to one, a field whose type the source writes. No new resolution rule — a call the
reader was not emitting.
PyCG: 59.1 % -> 61.5 %, precision 0.9933 -> 0.9936
Only where the iterable is written on ONE line and holds no comma: a tuple of iterables is not what the loop iterates, and a receiver spanning lines is text the resolver cannot read back.
What it still does not reach. for i in Cls(): i() — what i holds is what
__next__ RETURNS, and a Python return func writes no type. Same for a
generator's yield. Both need value flow through a return, which is a different
mechanism from reading a declared type.
A factory hands back a function, and Python says so with no type¶
def make():
return handler
a = make()
a()
The name holds the RESULT of a call, and the result is a function because the one
it called hands one back BY NAME. Nothing in the calling file says what a is,
and Python has no return type to declare it with.
Where the rule runs is the point. Not with the other rules for a bare name,
but at the stage where the call that FILLED the name has already been placed. That
is what lets it work for a method — b = a.func1() reaches a declaration a bare
func1 could never reach, because a bare call may not reach a member — and for a
for loop's variable, which is filled by the __next__ the loop calls:
for i in Cls(): i().
The returned name is looked for in the file that declares the FACTORY, not the
caller's: from lib import make then a = make() reaches a name written over
there.
PyCG: 61.5 % -> 63.5 %, precision 0.9936 -> 0.9938
return self.other counts too. It hands back a bound method, and that name is
what the caller calls — which is how three of the corpus's class cases are
written, super_class_return among them, where the method is the base class's. Only
self: any other object is a value this reader does not follow. The method is then
looked for by name in the declaring file, and two classes with a method of that
name make it ambiguous, so nothing is answered.
What a returned EXPRESSION says is nothing. return 1 + 1 and
return make() hand back a value this cannot name.
A C++ operator is a declaration¶
auto operator()(T value) -> bool { … }
How fmt writes every visitor and how any C++ tree writes an iterator. muundo
declared none of them: the node naming one is an operator_name, and the walk
that unwraps a declarator to its innermost identifier went straight past it, so
the whole definition became no entity.
What it cost is the model, not the score. 417 of the source positions fmt's
compiler names as a call target had nothing declared there, so no resolution rule
could ever have answered for them — and every operator a class defines looked dead
to reachability, the same defect as Python's unemitted for loop. Measured on fmt
with scripts/corpus_ceiling.py:
positions nothing was declared at: 547 -> 130
the coverage it could reach: 80.2 % -> 94.9 %
the coverage it does reach: 28.8 % observed, unchanged
precision: 0.9857 -> 0.9869
Saying the second line and not the third would be the mistake this entry exists to avoid. 202 declarations appeared in fmt and not one more call was named, because naming the target and reaching it are two different questions.
A conversion operator is left out. operator basic_string_view<Char>() const
is an operator_cast, whose own text carries its parameter list and qualifiers;
naming a declaration after that would put brackets and a const in the name. Six
positions in fmt.
Why C++'s remaining gap is not a parse ceiling¶
That was the standing explanation and it is wrong. format.h carries ONE error
region — 4 624 lines, the whole include guard — and that looked like the answer
until the region's own children were counted: 332 of them, parsed, including
every function_definition and template_declaration inside. walk_tree visits
the children of an ERROR node like any others, so those declarations were read.
The message "code under those regions was NOT read" overstates what is lost.
Two hypotheses were tried and measured away: expanding FMT_BEGIN_NAMESPACE made
the region BIGGER (4 619 lines), and blanking it, or the template specialization
beside it, left the region exactly where it was.
What the 1 858 reachable misses actually are, from scripts/corpus_ceiling.py and
the site breakdown: overload and instantiation SETS. One source line —
return vis(value_.int_value); — is expected to reach two different operator()
declarations, because clang's IR is read across 18 translation units and each
instantiation is its own function. muundo can name one declaration for one call.
And the receivers are template parameters (vis, handler), so there is no
written type to read: the compiler knows them by instantiating, which is the line
this tool does not cross.
An operator used as syntax is a call, and tree-sitter always said so¶
*out++ = '.';
Three operators on one line — operator*, operator++, operator= — and
tree-sitter gives every one of them: a pointer_expression holding an
update_expression holding an identifier. The C++ reader only ever turned a
call_expression into an edge, so 357 of the positions fmt's compiler names as a
call target had no edge written towards them at all. Nothing was missing from
the grammar; the reader was not asking.
*it is now it.operator*, ++it is it.operator++, a[i] is a.operator[] —
written on the receiver so every rule that resolves one already resolves these,
the same shape as a Python for loop's two calls.
An operator nothing overloads is not a call. *ptr on a raw pointer and i++
on an int are what the language does. Those edges resolve to nothing and are
dropped, so a plain C file gains no edge at all and the report carries only the
overloads.
leveldb: 50.8 % -> 51.3 % observed, precision 0.9445 -> 0.9450
fmt: 28.8 % observed, UNCHANGED
sds, linenoise (plain C): not one edge added
the twelve other corpora: unchanged, to the call
fmt gains nothing, and the reason is the whole story of C++ here. Its
iterators are template parameters — OutputIt out_; — so the receiver of every
*out++ has no type any source writes, and two operator calls resolve in the
entire library against leveldb's thirty. The compiler knows those types by
instantiating. That is the line this tool does not cross, and no amount of reader
work moves it.
Where the syntax tree stops, and the bill that comes with it¶
The tree is DROPPED in the parse worker. PrematFacts exists for exactly that
reason — "materialized WHILE the tree is hot", so the large tree can be freed
before the serial resolution passes run. Everything after parsing therefore works
on TEXT: a callee is a string, and the resolver re-parses it.
That is a memory decision, not an accident, and it is the right one. What it costs is worth writing down, because the bill came due three times in one day.
The count, in core/src/analyzer.rs: 63 string-splitting sites, and 13
functions totalling about 190 lines whose whole job is to re-read a callee —
callee_basename, callee_scope_hint, split_at_the_last_dot,
bottom_of_the_receiver, name_of_the_call, receiver_expression,
without_new, import_leaf, a_subscript_written, and the four that join a
chain written across lines.
The rule that follows from it. A fact that needs STRUCTURE is extracted in the
parse worker, as a PrematFacts entry, from the node that holds it. Splitting text
in the analyzer is only defensible where the text was never structured — a
qualified name the resolver itself built. Four things done right on the same day
say what that looks like: slot_key reads a string/integer node,
member_named_by_a_lookup reads a lookup_type, the_only_type_in_a_union reads
a union_type, collect_container_slots reads a dictionary.
Four places where the tree was not asked, and what each cost.
- The C++ reader never asked for an operator.
pointer_expression,update_expression,subscript_expression— tree-sitter gives all three, and onlycall_expressionwas read. 357 of fmt's expected call positions had no edge written towards them. This was blamed on a parse ceiling for weeks. import_leafsplit a module specifier on its dots. The parser held the specifier; the analyzer took./App.cursorapart and filedcursoras an imported module. excalidraw gained a point when it stopped.join_chains_written_across_linesrebuilds a chain from WHITESPACE. The tree knows the chain's shape. Compacting the source text gluednew ArrayListintonewArrayListand lost 3 278 of gson's resolved edges untiljoins_two_wordswas added to undo the damage.declarator_qualifierchooses text over the AST in writing: "Text-based split … is robust to the grammar's nesting direction". Sometimes true, and the reason it is stated is that it needs one.
A round trip is the smell. The C++ reader writes it.operator* from a node,
and the resolver splits that string back apart. Python's dispatch rule writes
d["a"] from a subscript node and a_subscript_written re-splits it on the
bracket. Both work, and both mean the structure was thrown away and rebuilt. Where
that round trip is unavoidable — the callee has to cross the tree's lifetime as
text — say so; where it is not, put the fact in the worker.
A callee is built from the tree, not from its source span¶
The first repair of the audit above. Nine extractors recorded node_text — the
raw source span — so a chain formatted across lines arrived with its newlines and
its indentation inside it, and a pass in the analyzer took them back out by
SCANNING CHARACTERS. That pass needed a string-literal state machine and a comment
check to be safe, and it still glued new ArrayList into newArrayList and cost
gson 3 278 resolved edges before a guard was added to undo the damage.
parser::called_expression_text builds the callee from the parts the tree gives.
One rule and one exception:
- a node is split into its children only when they cover every byte of it apart
from whitespace. A raw string's delimiters belong to no child — splitting
br#"{"version":1}"#.to_veclost thebr#"and the"#and produced a callee starting with a brace. That is what the exception is for, and finding it out cost a round of measurement; - where two WORDS would run together, one space goes back. That was always the right rule; the state machine around it was not needed.
And an eight-line net stays, where the pass was seventy-nine. One path still
hands over raw source: a call written inside a MACRO, where the grammar gives a
flat token_tree and no expression to take apart. ok!(self\n .parse(…)) in
clap is one, and without the net the receiver was lost and the call with it — one
in-tree call, measured, which is how the net earned its place.
all fourteen corpora: identical, to the call
malformed callees in muundo's own tree: 4 -> 3
lines of character-scanning: 79 -> 8
No score moved, and that is the whole result. A refactor that claimed one would be the thing to distrust.
The September 29 audit (dispatch conclusions superseded below)¶
A C++ declarator's qualifier is read off its nodes. void ns::Foo::bar() {}
files the method under ns::Foo, and that used to be the declarator's full text
with the leaf name stripped off its end and the :: trimmed. The grammar nests a
qualified name to the right — qualified_identifier(scope: ns, name:
qualified_identifier(scope: Foo, name: bar)) — so the qualifier is every scope
down the nest.
No input was found where the two disagree on real C++: the leaf is at the end, so
stripping it is sound. The reason to prefer the tree is not a bug, it is three
string operations to recover something already held, and ::bar — a qualified name
whose scope node does not exist — now says so structurally instead of by an empty
string.
Dispatch keys come from the tree, but text is not an occurrence identity. The original callee-text index conflated distinct literal values after whitespace normalization. Its file-wide slot table also conflated lexical bindings and used line numbers as instruction order. These conclusions were disproved by executable counterexamples; the replacement contract is recorded below.
And the eight that stay, with a contract. callee_basename,
callee_scope_hint, split_at_the_last_dot, bottom_of_the_receiver,
name_of_the_call, receiver_expression, without_new, import_leaf. These
currently inspect rewritten names as text:
the resolver rewrites e.callee as it goes — rebind_callee puts a qualified
name there, follow_the_chains does it again up to four times — and then asks these
helpers the same questions about the name it just invented. Original parser facts do not describe that new name. This does not preclude
a structured resolved-symbol representation alongside the original call site.
So what stays gets a contract. All eight had zero direct tests, and two of the
three defects found that day were in exactly this family: import_leaf taking
./App.cursor apart on the dot, and a callee built by scanning characters gluing
new ArrayList into one word. Both cases are now assertions, along with the one
that shows why the language has to be passed in: CodeBlock.builder().apply(action)
bottoms out at builder in Kotlin, where apply hands the receiver back, and at
apply everywhere else.
all fourteen corpora: identical, to the call
Dispatch proofs and measurement scope (September 30)¶
Literal nodes retain their source bytes, including whitespace, before expression formatting. Python keys use typed constant values: bytes differ from Unicode, while hexadecimal and decimal spellings of the same integer agree. Unsupported keys invalidate the literal table rather than receive an approximate identity.
Python dispatch facts are collected per lexical scope, in tree/source order. Assignments replace facts; unknown slot writes, container aliases, receiver calls, escaping arguments and captures invalidate them. Branches join agreeing facts; loops invalidate possible effects before and after their body. This is not a general reaching-definition analysis. Each read keeps its byte span, and the public CallEdge now carries an optional one-based byte column. Conflicting occurrences with the same public identity still abstain.
Resolved method declarations take priority over library name shortcuts. Rust's
shortcut list is used only for Rust; a declared unwrap, clone or Kotlin apply
can have a different return type and must not be skipped.
corpus_calls.py measures caller/callee pairs and explicitly says that it does
not observe site identity. corpus_call_sites.py expected.json report.json adds
a reviewed manifest of internal calls, each with file_path, line_number, caller,
callee, and optionally column. It preserves multiplicity, reports false in-tree
answers and misses, and declares unjudged calls and ambiguous legacy line groups.
Column-aware expectations require matching report columns; columnless answers
cannot pay for them. Legacy manifests still group by line and caller, and cannot
claim expression-level precision. Overlapping legacy and column-aware
expectations on one line/caller are rejected. The two measurement
units must not be combined into one percentage.
Regression coverage is in core/tests/dispatch_proofs.rs and
scripts/test_corpus_call_sites.py, alongside the existing positive dispatch and
callee-formatting tests. Corpus equality is evidence only within that corpus's
unit and scope, not a proof of correctness at every source occurrence.
The global whitespace fallback described in the September 29 entry was removed. Rust macro fragments already have a parsed expression tree; their callees now use the same literal-preserving builder, as do Kotlin navigation expressions. A regression checks both the resolved receiver and the exact raw string value inside a multiline macro call.
Measured on the same 119 PyCG cases before and after this change: 160 correct caller/callee pairs out of 252 expectations, with one existing false pair in both runs. This is unchanged pair coverage (63.49%), not a site-correctness claim. The executable counterexamples from the review go from nine wrong target claims to none: nine of their fourteen calls resolve correctly and five abstain.
What the counter-review changed, verified against running programs¶
The September 30 counter-review is right and the September 29 closure was not
justified. Nine minimal programs showed muundo naming a target the program does
not call — false in_tree answers, not silence. Every one is fixed, and five were
re-checked here against what the language actually does rather than against a
test:
| program | what it really does | what muundo said |
|---|---|---|
make().unwrap().finish() (Java, javac + java) |
prints B.finish |
B::finish |
d = {"k": a}; d["k"](); d["k"] = b; d["k"]() |
a then b |
a then b |
d = {"a b": a, "a b": b} read twice |
a then b |
a then b |
| a table built in another function, parameter passed in | b |
abstains |
e = d; e["k"] = b; d["k"]() |
b |
abstains |
The last two are the right answer: neither is provable without interprocedural analysis or alias tracking, so saying nothing is what a report should do.
Three things the review's own patch needed before it could ship, each found by a guard or a measurement, not by reading:
- It failed the house rules.
collect_subscript_readsmeasured 27,string_key26,compare_call_sites12 andassign11, against a ceiling of 10 with one exception already spent. Split into named pieces — the escape table is data, the scope walk is five functions — and eight things carried no doc comment. - It lost a TypeScript chain. The new "a resolved declaration outranks the
library shortcut" check returned from the whole function on its first answer,
so
const client = createInstance();thenclient('u').text()reached nothing: a const declares no return type, and the two fallbacks below had been skipped. - And it refused too much. Refusing the chain whenever the tree declares SOME
filter,maporitercost 13 of clap's answers. The check now decides only WHERE THE CHAIN STARTS, and only for a name the shortcut would actually step past; everything below keeps its own answer.
What the fix costs, measured. clap 79.0 % → 78.9 %: one in-tree pair, from nine sites that used to give one of several expected targets and now abstain, against one fewer wrong-class answer. Precision unchanged at 0.9751. The thirteen other corpora are identical to the call.
a declared `unwrap`, `clone` or Kotlin `apply` is not the standard library's
That is the trade, and it is the one the review asks for: a target claimed certain and wrong is worse than an abstention.
The two things the counter-review left open¶
Both were named as not corrected, and both are corrected now. Neither moves a corpus figure, and both are checked against what the interpreter does.
A line is not an occurrence: CallEdge carries a column¶
d = {"k": a}; d["k"](); d["k"] = b; d["k"]()
Python calls a then b. Both calls were on one line, spelled the same, so the
dispatch table's two records landed on one key and the site abstained — twice.
CallEdge now has column: Option<u32>, the 1-based column of the call's OWN
NAME, read off the same token as line_number so the pair points at one place.
Seventeen of the twenty-seven edge constructions record it; the rest are synthetic
edges and sites whose line was computed rather than read off a node, and there the
field is absent.
The dispatch index is keyed on it, and the merge-on-conflict rule stays as the net under it: where two records still meet on one key, the site abstains rather than take whichever the map kept.
A new field in a published record. resolution, resolved_by, placed_by
and wrote were each added this way — #[serde(default)], absent when unset — and
the reason to add this one is the same as wrote's: anything reading a report back
against the source needs to find the call.
A branch is joined, not a barrier¶
if, for, while, try and with all used to CLEAR what was proved. So
if x: d = {"k": a} followed by d["k"]() said nothing, and so did a table built
and read inside a with block — which runs exactly once.
withis sequential. Its body runs once; it is read in order like any other block.ifis JOINED over its arms. A slot keeps its value only where every path agrees. Both arms writing the same function proves it; arms that disagree prove nothing. With noelse, the path that SKIPS the branch is one of the paths, which is whyd = {"k": a}thenif x: d = {"k": b}proves nothing. Every condition runs before the branch is decided, anelif's among them, because it is evaluated on the path that skips its own body too.- A loop forgets possible effects on both sides of its body. This includes assignments, alias escapes, method receivers and call arguments. A later iteration cannot reuse an earlier proof invalidated by these effects. Read-only loops retain facts. The original implementation invalidated only before the body and could export an assignment from a loop that never ran; this was fixed on October 1 with executable counterexamples.
-
tryandmatchremain barriers. Their internal reads are unproven. Path exits (break,continue,return,raise,yield) stop proving for the rest of the lexical scope, conservatively including sibling branches.all fourteen corpora: identical, to the call
This is still not a general reaching-definition analysis, and the words matter:
a join over the arms written at one if is not the same as following a value
through a break, a continue, an exception or a match guard. The explicit
barriers above trade recall for avoiding a fabricated target.
Executable rule checks (October 1)¶
scripts/rule_proofs.py compiles and runs twenty small programs. Builds and
executions must succeed within a bounded time, and stdout must match a reviewed
trace before Muundo answers are judged. Generated files stay outside the analyzed
source tree. These checks fail on bad execution as well as false in-tree targets;
abstention remains a measured outcome, not a fabricated answer.
The harness still judges reviewed target sets per line, not individual columns.
Its trace assertions validate each concrete program run; they do not derive a
complete dynamic call graph or prove unexecuted paths. Column-sensitive manifest
comparison is provided separately by corpus_call_sites.py.
Reference declarators and imported receiver heads (October 1)¶
C++ reference declarators store their wrapped declarator as a named child rather
than in a declarator field. Name extraction and qualifier extraction now share
that traversal. A definition such as Box& Box::other() must create
Box::other, attribute its body calls to that method, and resolve b.other();
recovering only the leaf name silently turns it into a free function.
An import-path leaf is only a weak hint. A receiver such as
h.app.refreshEditorInterface() does not refer to import App from './app'
just because it contains app. The import fallback now checks the receiver
head, leaving field-type resolution available. Tests cover this collision and a
direct namespace import that must continue to resolve.
The direct CommonJS form const utils = require('./utils'); utils.etag() now
uses explicit export properties. exports.etag = actual,
module.exports.etag = actual, inline callables and literal export objects carry
the identity emitted by the parser. Private homonyms do not satisfy this proof.
Replacement, unknown property writes, alias detachment and unsupported export
control flow remove facts rather than retaining an obsolete target.
Forwarding module.exports = require(...) chains use a visited-file set instead
of an arbitrary four-hop cutoff; cycles terminate without inventing a target.
The executable harness includes a renamed receiver through eight forwarding files.
These import bindings are still file-wide. A parameter, nested declaration or
reassignment conflicting with a receiver prevents this new proof, as does a local
replacement of require. This is conservative scope handling, not a claim of full
JavaScript dataflow. Existing default calls and destructured requires retain their
separate resolution paths.
Imported CommonJS objects can be mutated (October 4)¶
The explicit export map describes the module's exported object before a consumer changes it. A property write, deletion or update on the consumer's receiver must invalidate that map for this binding. Copying the receiver into another binding or passing it as an argument also invalidates it, because those references may mutate the same object. A constant binding does not make its object immutable.
The file-wide binding model currently abstains for affected receiver names; it does not claim flow-sensitive resolution of their replacement functions. Unrelated object mutations preserve the import proof. Four executable Node counterexamples demonstrate that the replacement runs, and the integration tests also cover computed writes, assigned aliases, deletion and updates.
A symbol table per file, and what it is NOT asked (October 4)¶
Every fact the resolver read about a name was keyed on a FILE and the name:
the names a require bound, the names a parameter shadowed, the regions a
let held in. A parameter called utils in one function therefore blocked
the module-level utils for the whole file, and a bare call to a name that
was a parameter of the function around it could reach a function of that
name three files away, because nothing between the call and the tree said
what the name WAS at that position.
core/src/parser/symbols answers that question. Built in the parse worker
while the tree is hot, from one shared walk and one grammar per language
family: the scopes of a file, the names each scope binds and how (a
function, a class, a parameter, a local, an import), what each binding's
right-hand side wrote as a [Value], the later writes to it, and what
disturbed the object it holds. A lookup walks from the call's own position
outward. What a name HOLDS at a position is decided by one rule in every
language: a write dominates a use when its scope encloses the use's; a
write in a branch the use is not in proves nothing unless it agrees with
the dominating one, or every arm of a complete alternation settles the
same value; a write later in a loop the use is in, or anywhere in a
closure, makes the answer unknown. That is the rule the Python dispatch
proofs were written to, made the same for every grammar, with fifteen
tests on snippets pinning it.
Two lessons from the first measurement, both about what the table must not be asked.
The first version replaced the per-language value-binding collectors
with the table's [Value::Call] and read clap 79.0 % -> 68.4 %. The
collectors know which call in self.args().iter().next() a name's value
ends at, because iter and next hand their receiver back; the table
knows the called expression as written. That is knowledge about CHAINS,
and the table is about SCOPES. Both stay: the table says which binding a
name is at a call, the collectors say what that binding was filled with.
Conflating them cost ten points, and a gate requiring the two to agree
cost forty-two calls more, because a region is a pair of lines and a
scope is not.
The second placed the scope's verdict — "this name is a parameter here"
— BEFORE the rules that read the caller's own file, and it cost
excalidraw 71 calls and gson 4: renderAction: ActionManager["renderAction"]
is a parameter whose annotation names the declaration, and delegate()
in a Java class is a member, parameter of that name or not. The scope's
verdict now comes after every rule that answers from the file's own
evidence and before the one that searches the whole tree for a homonym,
which is the one guess it exists to refuse.
What the resolver reads from the table, measured on the ten pinned
corpora against 08c96fa: eight identical to the call, excalidraw
+7 named, fmt one false outside fewer, nothing lower anywhere, and the
twenty executable programs 20 right, 8 abstained, 0 wrong.
- The CommonJS receiver proof asks whether the receiver AT THE CALL is the
binding the file's
requiremade, undisturbed, with norequireof the scope's own — instead of a file-wide set of names that conflict. - A receiver the scope binds to a value skips the three rules that would read it as a module. The rules that read its written type keep their answer.
- A bare name that is a parameter, or a local holding nothing the file
wrote, is
undeterminedrather than the tree-wide homonym.
What is next is the other direction: the resolver's rules still begin from the callee's TEXT and ask the table on the way; the rewrite begins from the table's answer about the callee's head and asks the type rules after. That is the step the per-language collectors are retired in, one language at a time, each measured.
The head of the call first, and the chain asks the table where the collectors are silent (October 4)¶
resolve_one now asks the symbol table ONCE what the callee's head is —
the bare name, or the one name a receiver is written as — and hands that
answer to every rule as a [Head]: a declaration the scope binds, an
import, a parameter, a local with what it holds, nothing, or not asked.
The two places that each looked the name up on their own read the same
answer now, and the measurement of the restructure alone was ten corpora
identical to the call.
The rule that paid is the chain's start. types_the_binding_holds read
the per-language collectors, and the TypeScript collector records a
const filled by a call at module level only. Inside a function —
const child = parent.extend() in a test — nothing was recorded and the
chain never started. The chain now asks the table SECOND, where the
collector is silent: a local holding a call, a construction, a name or a
path, as the same record the collectors produce. Second, because the
collectors know chains — foo().unwrap() ends at foo for them, at
unwrap for the table — and a return type is written on the first.
Measured against the symbol-table commit: ky 74.1 % -> 81.2 %,
233 more named, precision 0.9942 -> 0.9947; cobra 82.8 -> 84.1 %; fmt
28.8 -> 29.1 %; gson, clap and excalidraw a few each; nothing lower, no
false edge added, no false outside added; the twenty programs
unchanged. Two false edges this would have introduced, each caught by a
corpus and written into the table:
- An awaited call yields what the promise resolves to.
const r = await ky.get(url)is aResponse, not theResponsePromisethe call returns; read as the call it sent 25 of ky'sr.text()to the wrong declaration.awaitis opaque to the table. - A loop variable holds an element, not the iterable. The Java
grammar bound
for (Factory f : factories)tofactories, whose written type isList<…>, foreign — and seven of gson's calls onfbecame a falseoutside. The same shape in TypeScript'sthisreturn type —squash(): this— read as a type the tree knows nothing of and did the same to three of excalidraw's;thisis the current type, asSelfalready was.
What the head does not yet decide: which DECLARATION a bare call reaches when the scope binds one. The misses say that is not where the calls are — ky's remaining half are callbacks held in options objects and parameters, which need the callable-parameter model, not a rule.
The overload a call reaches, named beside the qualified name (October 4)¶
The record above says why the overload lever was measured and not taken:
separating overloads inside qualified_name changes everything that reads
a callee, and arity alone separates 161 of Newtonsoft's 323 sets. Both
hold. What was not tried is the ADDITIVE answer: the qualified name stays
the identity every consumer keys on, and the edge gains
declaration_line, the line of the one overload the call fits, absent
where fewer or more than one do. A reader that ignores it reads the report
it always read; a reader that wants the declaration has it where the source
says which.
What decides it, in order, and what refuses nothing. The parser records,
in the same walk that counts arguments, each declaration's parameter types
as written and each call's argument SHAPES: a plain name, a literal with
the type a literal of that spelling has in that language, a construction or
a cast with its type, or unknown. After resolution, an in-tree edge whose
callee names several declarations keeps the ones the count fits, then
drops the ones a written argument type contradicts. A plain name's type is
the one its own declaration wrote, two lines up; a type variable, Object
or an unknown parameter type refuses nothing; and TWO CLASS TYPES NEVER
REFUSE EACH OTHER, because JsonObject may extend JsonElement and the
spelling does not say. Only a closed type — a primitive, its box, a string,
null — can contradict, by a table of what boxing, widening and the
string supertypes allow.
Three things the corpora corrected before this shipped.
- The first version compared class types by name and read gson 83.5 %
with 41 false edges where there had been 2:
gson.toJson(obj)with aJsonObjectchosetoJson(Object)overtoJson(JsonElement). The open rule gives up 5 points of that recall and all 39 of the false edges. - A
catch (IOException e)is spelled like a parameter list, nothing names it, and the name search climbed to the enclosing method: every method with a handler was recorded a second time with the handler's parameter as its own.names_a_parameter_listrefuses it now, which also tightens the arities every rule already read. - A string literal is a
Stringin Java and a&strin Rust, andStr::from("Commands")had been givenFrom<String>. Literal types are per language, Rust's numeric literals have none, and a lifetime is not part of a type's spelling.
And one correction to the oracle. The javac harness reported, for
roundTrip(gson, gson, PROTO) where PROTO's generated class is not in
the tree, the FIRST overload — javac's pick for an erroneous argument type,
reported as the truth. The oracle now reports no target for a call with an
argument whose type did not compile; 211 of gson's sites leave the
judgment, and MEASUREMENTS.md says so beside the number.
Measured on the ten pinned corpora, the two with overloads on the
corrected oracle: gson 73.3 % -> 78.9 % (8 061 -> 8 677 of 10 998
named, overload_unnamed 2 494 -> 1 888, false edges 2 -> 2, precision
0.9998); fmt 29.1 % -> 31.4 % observed (overload_unnamed 391 -> 298,
false edges 11 -> 11). The eight others identical to the call; the twenty
programs 20 right, 8 abstained, 0 wrong. What is left of gson's overload
gap — 1 888 calls — is arguments that are calls or expressions, and
overloads that differ only by class types; the first needs the return
types the chain already reads, the second needs the hierarchy, and both
are evidence this engine has in part and will read next.
A call on a parameter names the parameter (October 4)¶
onChange() inside a component that takes onChange reached nothing:
undetermined, which is true and less than the file says. The TypeScript
checker answers the same call with the PARAMETER's own declaration, because
the signature it resolves — the parameter's function type — has no body,
and the code that runs is whatever the caller passed. Half of what ky and
excalidraw still missed was this shape, and the measurement record named it
as outside the model: an entity for every parameter would be a different
product and different numbers everywhere.
The answer that costs no entity: a resolution of its own. parameter, with
declaration_line at the parameter, placed by the family parameter. No
node is invented, the metric graph draws no edge to it, and a reader of the
report learns what it could not before — that this call invokes a value the
function was handed. The symbol table decides it: the head of a bare call
is a parameter of a scope around the call.
Claimed only where the checker would claim the same, which the corpora settled in three rounds, each a false edge against tsc until it was written down:
- A parameter typed by inference alone — a destructured
({ generateQRCodeSVG }) =>under a dynamic import — resolves, for the checker, to the function the import carries. So only a parameter whose type the source WROTE is claimed, and the table carries the annotation ([Value::Annotated]). Plain JavaScript writes none and is unchanged. func: typeof resizeFrameOverElementnames a declared function, and the checker names it. Atypeoftype is not claimed.- A destructured prop declared
onExportImage: AppClassProperties["onExportImage"]resolves, for the checker, to that class property's arrow. An indexed access type is not claimed where the member's declared type is readable. Five of excalidraw's calls are this shape under a type alias whose members the field index records without their index, and those five are the cost published below.
Measured on the ten pinned corpora: excalidraw 92.2 % -> 93.5 % (+225 named, false edges 107 -> 112, precision 0.9935 -> 0.9933); ky 81.2 % -> 81.5 % (+9, no false edge); the eight others identical. The five are convention, not fabrication: muundo says the call reaches the parameter it does reach, and the checker follows the parameter's declared type to an implementation. They are counted as wrong all the same, because the oracle is the oracle.
Proved by programs, not only by corpora. scripts/rule_proofs.py now
runs Java and Rust, and an expectation may demand an overload by
@line. Four cases are these two entries': a Java overload told by what
the call passes (an int, a String, two arguments, a String variable;
a Sub where a Base overload also exists must stay unnamed), a Rust
"literal" that is a str and not a String, an awaited call whose
value is not the promise get returned, and a call on a typed callback
parameter that reaches the parameter and never a or b, whichever
caller passed them. The harness ran the programs, read what they printed,
and judged the report: 28 right, 9 abstained, 0 wrong.
A call on a local names the local, and a chain's call is filed where its name is (October 4)¶
const api = ky.create(); api(url) resolves to api, line 11, as
resolution: local. The call runs whatever api holds: a value of the
type create declares it returns, whose call signature no declaration
names. The TypeScript checker answers the same — the local's own
declaration — and so did muundo for a parameter already; a local is the
same shape with the value written one line up instead of one frame up.
for (const hook of hooks) hook() and let done: () => void are the two
other spellings of it.
Only where the checker would answer the local, and measured three times
to find where that is. The first version claimed every local filled by a
call and cost excalidraw 139 false edges: const throttled =
throttleRAF(fn) is a call that returns an INLINE function, and the checker
names that function, not the local. So a call-filled local is claimed only
where the call was placed and the declaration it reached returns a type this
tree declares. A name destructured from a call — const { t } = useI18n() —
is not the local either: it is the member t of what useI18n returns, and
the chain reads it there ([Value::Member] in the table, and
a_member_of_what_a_call_returned in the second pass). Destructured from an
await, or a loop variable over a literal of several functions, it is
undetermined: the value's origin is unreadable, or it is one of several.
The line of a call is its name's line, and await x.json<T>() had lost
it. tree-sitter-typescript parses that call as one whose function is the
AWAIT of the member, and the line walk stopped at the await, on the
chain's first line, three lines above the name. 76 of ky's calls were on the
wrong line for that alone. The walk now steps through await, ! and
parentheses. And json: { <J>(): Promise<J>; } — an object type made only of
call signatures, which is how TypeScript writes an overloaded property — is
a function the way text: () => Promise<string> is, so ResponsePromise::json
is an entity; so is a property whose union type has a function arm,
throwHttpErrors: boolean | ((status: number) => boolean). TypeScript calls
carry a column now, like every other language's, so a binding written
earlier on the call's own line is seen.
super() under a base class that writes no constructor reaches the
class. The language supplies the constructor; the class is the declaration
there is, and the one the checker names.
Measured on the ten pinned corpora: ky 81.5 % -> 93.1 % (+378 named, false edges 14 -> 14, precision 0.9948 -> 0.9954); excalidraw 93.5 % -> 93.7 % (+42, false edges 112 -> 112); the eight others identical.
An element has the type its collection names (October 4)¶
for _, c := range cmd.Commands() then c.IsAvailableCommand() reaches
Command::IsAvailableCommand. 46 of cobra's 87 remaining misses were a
call on a range variable, on an indexed element (cmds[0].Name()), on a
local filled by one (first := cmds[0]), or on a &Command{}. The table
binds a range variable to [Value::Element] of what it ranges over and an
index expression to an element of its operand; the resolver reads an element
through its collection's binding, because the collection's type is already
read as its element wherever a type is read — bare_type_name steps through
a slice, an array, a channel, a variadic and a map's value as it did a
pointer. A receiver written with a subscript (cmds[0].Name) is resolved
with the subscript dropped and kept in wrote; where the collection's type
then declares no such member the answer is undetermined and never
outside, because the collection's members say nothing about its element's.
A struct's fields are indexed by their struct (Go had recorded them as
file-wide typed names only), so y := c.commands and c.helpCommand.Name()
read the field's type the way TypeScript's this.a.b.c() does, and a local
holding a plain name follows that name to what it held, a few aliases at
most.
And a subscript on a pointer is not an operator[]. keys[i].data()
with const Slice* keys had an edge to Slice::operator[], which the C++
reader drew for every a[i] whose a named a type with one; it now looks
for a pointer or array declarator of that name in the function and the
class around the operand first. Two of leveldb's false edges.
Measured on the ten pinned corpora: cobra 84.1 % -> 90.9 % (+37
named, false edges 0 -> 0, precision 1.000); leveldb 51.2 % -> 51.9 %
observed (false edges 35 -> 35 after the operator[] correction, precision
0.937 -> 0.938); ky and excalidraw +1 and +3 named; the six others identical.
The executable proofs read 37 right, 9 abstained, 0 wrong, with a Go case
run by go run.
The overload is named from what the tree writes (October 4)¶
1 450 of gson's 1 888 unnamed overloads are named, and one false edge
became zero. The first version read a literal, a construction and a plain
name, and left two class types alone. What the corpus had left was
gson.fromJson(json, Foo.class), 593 times, gson.toJson(x) where x is
declared something, writer.value(5.0), and arguments that are calls. So:
- An argument that is a call has the type the declaration the call
reached returns, read off the resolved edge at that position.
<T> T same(T t)returns a type variable, which says nothing, and the argument stays unknown. X.classis aClass,thisis the caller's class,JsonNull.INSTANCEis the static field's written type (Java fields are indexed by their class now, as Go's and TypeScript's are), a lambda is a function with no name, which no closed type, noObjectand no class this tree declares takes.- Two class types are judged by the hierarchy the tree WRITES. Both
declared: the parameter's type must be the argument's or one of its written
bases, transitively. Argument declared, parameter not: only where the
argument's ancestry reaches a type the tree does not declare. Argument not
declared, parameter declared: never — nothing outside this tree extends a
type inside it, which is how
CollectionrefusesJsonElementand leavesObject. Neither declared: unknown, and nothing refused. Only in Java, C# and Kotlin: C++ templates relate types at instantiation, and applying the rule to fmt drew four false edges before it was confined. - Of several fitting overloads, the most specific is named, as Java names
it:
Class<T>overTypefor aClass,doubleoverNumberfor a5.0,g(Sub)overg(Base)for aSubthe tree declares aSub. Only where every argument's type is known: the first version applied it with an argument unknown and namedtoJson(JsonElement)for 54 arrays, since the narrower overload is only narrower if it applies at all. - A Java
@interfaceis an interface and its members are methods, soannotation.value()reachesSince::valueinstead of being counted as a call that left the tree. Ten of gson's thirteen falseoutsideclaims.
Two things measured and refused. A single-letter type name is not a
type variable on the argument side: gson's tests declare classes A, B
and C, and reading them as "anything" named toJson(JsonElement) for
every new A(...). And a C++ template parameter is left spelled as written:
spelling it as "takes anything" made 38 of fmt's overload sets ambiguous
that the written spelling told apart (count_digits(uint64_t) against
count_digits<BITS, UInt>(UInt), which needs an explicit BITS and is not
viable for a bare call). Java, C# and Kotlin type parameters ARE spelled ?,
because there a generic method's T is anything within its bound.
Measured on the ten pinned corpora: gson 78.9 % -> 92.7 %
(overload_unnamed 1 888 -> 438, false edges 2 -> 1, false outside 13 -> 3,
precision 0.9999); the nine others identical to the call. What is left of
gson's gap is an argument whose type is a field of another object, a
ternary, an arithmetic expression, or a method whose overloads differ by a
type the JDK relates and this table does not list.
What clap's false edges were, and a quadratic walk in the facts pass (October 4)¶
Twenty-three of clap's forty-six false edges were the instrument's.
rustc compiles one package. clap_builder calling clap_lex::RawArgs::new
reaches an external symbol with no body in its IR, and corpus_rust.py
read every such symbol as "outside" — but clap's clap_lex crate's lib.rs is in the
tree, and muundo naming the declaration there is right. The oracle now
reads the crate a mangled symbol starts in (…Cs11YjcoaHbUt_8clap_lex…,
or the legacy _ZN8clap_lex…) and, where that crate is a Cargo.toml
of the same tree, leaves the site unjudged: the compiler cannot place such
a call and the tree can. Twenty-three sites, 0.975 -> 0.987 for the oracle
alone.
The other half, four rules. a.settings.set(…) with a: &Arg reached
AppFlags::set, because Rust's fields were typed file-wide by name and
Command.settings won; a struct's fields are indexed by their struct now, as
Go's and Java's are, and type_of_the_path reads them. let args =
Vec::new(); args.push(r) reached MKeyMap::push through the file-wide
args field: a field typed for the whole file no longer types a local or a
parameter the scope binds inside a function. self.get_subcommands().find(|s|
…) reached Command::find, because -> impl Iterator<Item = &Command>
hands a Command to the chain and find is a member of Command; a first
fix took the element type away from iterators and cost 59 named calls on
clap and 900 on excalidraw, since a loop over the iterator and the closure's
parameter ARE the element. The rule is on the member instead: find,
filter, last, push and the other names an iterator or a collection
has of its own, written on one, reach nothing of the element. And
Inner::from_static_ref under mod inner { … } in the same file was
"outside", the resolver knowing no crate called inner; a use whose head
the file binds as a module is read in that file.
And the facts pass had a quadratic walk, found by the server's own
cancellation test. type_parameters_in_scope, written the same day,
scanned the children of every ancestor of every parameter list — and a
module of 26 000 functions is an ancestor of each. 700 million nodes for
one file; the test that cancels mid-parse ran 500 seconds and failed. The
walk reads the type_parameters field now, and scans children only of a
node with few. The symbol table had its own: an assignment asked whether its
name was already bound by reading every binding (2.6 billion comparisons on
that file), and a lookup read every binding and every scope. A set during
the build and indexes after it: 290 seconds to 8.5 for that file.
Measured on the ten pinned corpora: clap false edges 46 -> 22
(precision 0.975 -> 0.988, false outside 9 -> 8), coverage unchanged at
79.0 %; excalidraw's false outside 27 -> 26; the eight others identical.
Four shapes PyCG taught, and three corpora measured again (October 5)¶
PyCG 63.5 % -> 71.6 %, false pairs 1 -> 1. The micro-benchmark is 119 programs, each a shape; four of them were read:
func()()runs whatfunchanded back by name, andfunc()()()what that did. The callee is written as the inner call, andwhat_calling_hands_backreads it from the inside out, three levels at most, throughhands_back_by_qn— the same record the factory rule reads fora = make(); a().a.func1()()the same way, under its own name.- A method the class inherits, in the order Python resolves it.
b = B(); b.func()withB(A)reached nothing: the chain found nofunconBand stopped. It asks the bases now, and in the C3 linearisation: forD(B, C)overB(A),C(A), that isD, B, C, A, andC.funcwins over theAboth share — a depth-first walk namedA.func, which is the one false pair the first version added. In a language with one base the linearisation is the chain of bases, so Java and TypeScript read the same way. - A keyword argument fills the parameter it names.
func(func2, c=func4, b=func3)is recorded asfunc2,c=func4,b=func3; the rule that reads what callers pass for a parameter looks forb=before the position. - A lambda is an entity.
x = lambda n: n + 1declares a function with no name; it is<lambdaN>, N its place among the file's lambdas in source order, which is PyCG's own spelling. A call onxreaches the lambda bound on that line; a lambda passed as an argument is recorded under that name, so the parameter it fills reaches it. The keywordlambdais a node of the same kind as the expression and is not counted. A lambda has no docstring to write and is not assessed for documentation.
The pair count is 243, not 252. Nine expectations named a module the case
does not hold: ext.function under from ext import function, with no
ext.py beside it. No in-tree edge can answer that and outside is right,
which is the ground <builtin> was dropped on; they are dropped the same
way, in expected_edges.
Three corpora measured again. dotnet 8 from Ubuntu's archive, GNAT 13
the same way, and a kotlinc 2.0.21 assembled from Maven Central's
kotlin-compiler-embeddable and its five dependencies (the proxy refuses
GitHub's release files). Newtonsoft.Json 66.3 % -> 80.6 % on 19 378
in-tree calls — the overload work, overload_unnamed 5 414 -> 2 565 —
precision 0.979; klaxon 67.1 % on 307 calls, compiled without its Jackson
module; ada-util 98.0 % observed, 192 of 223 units built, precision
1.000. All three are pinned in scripts/corpora.json with their toolchain
written beside them. KotlinPoet is a multiplatform module, which kotlinc
compiles only when told so (-Xmulti-platform -Xcommon-sources=…): 38.9 %
on 874 calls, precision 0.929, the weakest row of a language with a working
oracle, and the Kotlin reader's next piece of work. photos is on gitee, which
the network policy does not reach, and stays unmeasured and says why.
Status s; is a call (October 5)¶
The C++ reader left an object declared without arguments out, on the
ground that "nothing is passed and a default constructor may not exist".
Where the class writes one it does exist and it runs, and clang records the
call on that line: 173 of leveldb's call sites were Status s;, Slice
key;, WriteBatch batch;. A declaration whose type is a named type and
whose declarator is a plain identifier is a construction now, as a
declaration with an argument list was; a pointer, a reference and an
array declare no object and stay out. Slice s(a, b); stays out too, the
grammar reading it as a function declaration.
And the constructor rather than the class. BlockHandle handle; had
resolved to the class BlockHandle, whose span the instrument accepts for a
constructor declared inside it; leveldb defines BlockHandle::BlockHandle()
below the class, and clang names that definition. A construction resolved to
a class reaches the class's one constructor where C++ declares exactly one.
C++ only, measured: the same rule in Java moved new Foo(…) off the class
the overload pass names constructors from, and cost gson 176 calls before it
was confined.
Measured on the ten pinned corpora: leveldb 51.9 % -> 53.5 %
observed (+29 named, false edges 35 -> 35, precision 0.938 -> 0.940); fmt and
the eight others identical. What is left of fmt is parsing, as published, and
detail::max_value<To>() reaching a template declared in two headers under
one scope — two qualified names, which this identity model does not merge.
The Kotlin reader, taught by KotlinPoet (October 5)¶
KotlinPoet was the weakest row with a working oracle, 38.9 %, and the misses sorted into four piles: 212 sites with no edge at all, 151 undetermined, 122 the right method with the overload unnamed, 35 placed outside the tree. The first pile was the grammar (above). The rest were the reader, and each rule below was measured on both Kotlin corpora before it stayed; the numbers are KotlinPoet's, klaxon moved 67.1 % → 68.6 % and its precision 0.958 → 1.000.
A class's fields hold a type (core/src/parser/kotlin.rs,
collect_fields): the one written after the name, val lexer: Lexer, or the
one constructed into it, val lexer = Lexer(reader). Kotlin was the one
language of the nine with a fields table that filled none, so lexer.nextToken()
inside the class reached nothing.
vararg and a default value are the parameter's. The grammar flattens
vararg args: Any? into a parameter_modifiers node and the parameter, and
nonWrapping: Boolean = false into the parameter and the expression — three
named children for two parameters, so add(format, vararg args) counted three
and took exactly three, and 182 calls fitted no declaration. The arity and
parameter-type readers in core/src/parser/mod.rs now read the list as the
grammar writes it. A spread reaches only a variadic: builder(name, type,
*modifiers) names the vararg overload and not the Iterable one beside it —
a new ArgumentShape::Spread, read in Kotlin, JavaScript, Python and Go. And
the fixed-arity overloads come first, Java's and Kotlin's phases:
emitCode("fun ") reaches emitCode(s: String) and not emitCode(format,
vararg args). Only when every argument's type is known: applied blind, the
rule sent addModifiers(KModifier.OVERRIDE) to the Iterable overload, 17
wrong lines, and was gated. A trailing lambda is an argument:
buildCodeString(codeWriter) { … } passes two, which this grammar writes as a
call_expression wrapping the call that holds the parenthesised list; and a
function type's own () is not the declaration's parameter list, which had
given buildCodeString(builderAction: CodeWriter.() -> Unit) a second arity of
zero. A construction passed as an argument has its type:
emitCode(CodeBlock.of(…)) was already typed by what of returns, and
emitCode(CodeBlock()) now by the class it reaches.
An extension function is reached through its receiver (kotlin_extension_function
in core/src/analyzer.rs). The rule that existed took the one top-level
function of the name, and KotlinPoet declares toImmutableList three times:
expect in commonMain, actual in jvmMain and actual in nonJvmMain.
The expect half is set aside from every name index — nothing compiles it,
and the oracle names the actual. Two extensions of one name in one file,
Char.isJavaIdentifierStart() beside CodePoint.isJavaIdentifierStart(),
share a qualified name and are told apart by the receiver's written type, the
edge then carrying the declaration's line, which the overload pass keeps where
the arguments decide nothing. Two actuals are told apart by the Gradle source
set: a caller under src/jvmMain/ reaches the one under src/jvmMain/. A
member extension, private fun Map<…>.generateImports() declared inside
CodeWriter, is reachable from inside CodeWriter and nowhere else. A
receiver whose type the tree does not declare — builder.members: List<CodeBlock>
— no longer sends the call outside before the extensions were asked: 25 false
outside on toImmutableList alone. And a receiver that is a foreign type
named as such, Character.isLowerCase(code), reaches no extension here: the
first version of the rule claimed it, two false edges.
The Builder a factory returns is the one nested beside it. Eleven classes
nest a Builder, and AnnotationSpec.builder(x).build() reached none of the
eleven builds. The type a factory names in its return is looked for first
under the factory's own owner, for a chain and for a value the factory filled
(val builder = AnnotationSpec.builder(x)). A type's own member outranks a
member of a class nested in it: the per-segment scope index files
CodeBlock::Builder::isEmpty under CodeBlock too, so codeBlock.isEmpty()
saw two and named neither; a direct-scope index answers first. A field
typed differently by two classes of one bare name says nothing: kdoc is a
CodeBlock.Builder in one Builder and a CodeBlock in another, and the
shared (type, field) table had kept whichever file came last.
!alreadyEscaped() is a call to alreadyEscaped. The grammar parses it as
(!alreadyEscaped)(), a call whose callee is a unary expression, and the
extractor had dropped it: 9 of the sites with no edge. for (spec in specs)
types spec where specs was written Iterable<AnnotationSpec>, for the
overload pass, which reads written types only.
The oracle had two faults of its own, found because muundo's edges were
right where it said they were wrong. Its javap reader matched a member with
^ [^ ], which consumed the first letter of one written with no modifier —
com.squareup.kotlinpoet.AnnotationSpec(…, DefaultConstructorMarker);, the
synthetic constructor kotlinc writes for a class with default arguments — so
om.squareup.kotlinpoet.AnnotationSpec was never the class, and every call to
such a constructor was "outside": 17 false edges charged to muundo on
AnnotationSpec(this), FileSpec(this), LambdaTypeName(…). A lookahead now.
And it counted the one call each isOneOf$default bridge makes from the
declaration's own line, a site no source wrote: skipped, and a call to a
$default bridge is the function it stands beside (scripts/corpus_kotlin.py).
Both corpora re-pinned: 857 and 299 in-tree calls.
Measured, in order, on KotlinPoet: grammar 38.9 % → 50.9 %; fields,
arity, extensions by receiver, expect set aside, nested Builder 50.9 % →
62.6 %; direct members, trailing lambda, fixed-arity first, typed loop
variables 62.6 % → 68.2 % (and precision 0.960 → 0.936, the blind phase
rule, which the gate restored); extension fallbacks, source sets, member
extensions 68.2 % → 74.8 %; the oracle's two faults 74.8 % → 74.0 % on 857
calls with precision 0.960 → 0.989; !f(), spreads, the unary callee and the
field table 74.0 % → 79.2 %, precision 0.991, one false outside. The
proof is core/tests/the_shapes_kotlin_taught.rs.
What is left, by the audit: calls on an implicit receiver inside
apply { }, buildCodeBlock { } and with(x) { } — add("\n") and
build() written bare inside a lambda whose receiver is a CodeBlock.Builder
(34 sites); overloads whose argument is a generic Class<*> or KClass<*>
against a class whose written ancestry reaches a type outside the tree, which
the hierarchy rule rightly refuses to rank (builder, addAnnotation, 27);
val builder = builder(x).tag<T>(y) through a generic tag returning T
(17); and a data class's get/put delegated to a map by by map, which
javap places on the class line and no source declares (klaxon, 25).
(October 9: the first and the last are answered — see "A Kotlin lambda's receiver" and "Kotlin, continued" below. The two in between are open.)
The C# reader, taught by Newtonsoft.Json's 337 false edges (October 5)¶
The false edges sorted into shapes, and each shape was a rule the reader lacked. Every rule below was measured on Newtonsoft.Json and on gson before it stayed; the figures are Newtonsoft.Json's, 80.6 % → 90.5 %, precision 0.979 → 0.992, false edges 337 → 148, overloads left unnamed 2 565 → 872. gson moved 93.6 % → 93.9 % on the same rules.
params JsonConverter[] converters is one parameter. tree-sitter-c-sharp
writes it straight into the parameter list as a type and a name with no
parameter around them, so SerializeObject(object, params JsonConverter[])
counted three parameters and took exactly three. The arity reader in
core/src/parser/mod.rs reads the pair as the one variadic parameter, whose
element type is what an argument is compared against. The this receiver
of an extension method is never passed: "…".FormatWith(provider, a, b)
passes three arguments to a declaration that names four, and the receiver is
excused the way Python's self is. An extension that cannot take the call
is not claimed: date.Trim('"') is string.Trim(char), where the tree's
only Trim extension takes a start and a length.
A call that writes type arguments reaches a generic declaration.
DeserializeObject<Foo>(json, converter) and DeserializeObject(json,
converter) differ in nothing but the <Foo>, and the overload pass could not
see it: the parser now records how many type arguments each call writes and
how many type parameters each declaration has, and a declaration with the
wrong count is out. A call that writes none says nothing — the language
infers.
An explicit interface implementation is reached through the interface
alone. void ICollection<JToken>.Add(JToken item) beside Add(object?
content) won every c.Add(x) whose argument fitted JToken better, and no
call on a JContainer can reach it. The C# reader marks it (explicit_interface
in the entity's metadata) and the overload pass sets it aside where anything
else remains — after counting the set, so the one declaration left is still
named.
Types the reader lost on the way. A byte[] data was a byte, so
WriteValue(data) reached WriteValue(byte): the typed-name tables of C# and
Java keep an array's brackets, and the overload pass alone sees them, every
other rule stripping them again. A decimal? value was a decimal, so
_textWriter.WriteValue(value) reached the override taking decimal and not
the decimal? overload: the mark is kept the same way, a nullable value does
not pass to the value's parameter, and a nullable reference is its reference.
1.1m is a decimal. Formatting.Indented passed as an argument had no type,
C# having no fields table at all: core/src/parser/csharp.rs now records a
class's fields and properties with their written types, and an enum's members
as values of the enum, and a C# member access is shaped as the static field
it is, as a Java field access was.
A class whose only member of the name cannot take the call is not the
answer. compositeExpression.IsMatch(o1, o1) on a class whose IsMatch
takes three reached that member, 52 calls, where the base class's
two-parameter IsMatch was the one called. The sole candidate under a
receiver's type is now asked what it accepts — its params and its defaults
counted — and refused when it cannot, so the rules that walk the bases
answer; those rules look the type up as the one TYPE of that name, not the
one declaration, since a C# class and its constructors share the name.
And the base's normal form over the type's expanded one: parsed.WriteTo(
bsonWriter) on a JObject is JToken.WriteTo(JsonWriter) and not
JObject.WriteTo(JsonWriter, params JsonConverter[]) with an empty array, as
C# removes the expanded forms once a normal form applies. An earlier attempt
at "the base's exact count" sent 697 calls wrong because it counted a default
as a mismatch; this one asks what each declaration accepts.
The overload set spans the hierarchy. writer.WriteValue((char?)null)
on a JsonTextWriter reaches the WriteValue(char?) its base declares, not
the WriteValue(object?) it overrides, and no set keyed by one qualified
name could say so. For Java, C# and Kotlin the pass now gathers the type's
own declarations and the bases' that no nearer signature hides, nearest
first, and rewrites the edge's callee to the base's when the base's wins.
With an argument the reader cannot type, the type's own overloads are asked
first and keep their answer: measured the other way round, 136 calls that
the override named rightly came back unnamed. Accessibility is part of
the set: s1.Equals(s2) from a test reaches Equals(object) and not the
protected bool Equals(NamingStrategy) beside it. Fixed-arity overloads
come first wherever one certainly applies: FormatWith(provider, object?
arg0) takes (CultureInfo.InvariantCulture, x) whatever the culture's
type is, which named 330 of the calls the "every type known" gate had
left; addModifiers(Iterable<KModifier>) may not take KModifier.OVERRIDE,
and the vararg beside it stays.
What a scope's members are. t.Add(1) on a JProperty reached the
Add of the JPropertyList nested in it, because the per-segment index
files a nested class's members under the outer class too. A type this tree
declares answers for its DIRECT members or not at all, and the rules that
walk its bases take over; a scope that is not a declared type — a namespace,
a module — keeps the loose index. JsonConvert.DeserializeObject<T>(json)
from a file that usings the namespace of a test class called
DeserializeObject reached that class by import: a member of a type the tree
declares is never an imported symbol.
What is left, by the audit: WriteValue on a receiver typed by a
property the reader cannot follow (92 undetermined); Children() on an
indexed receiver, rss["channel"]["item"], where the indexer's declared type
is not read (12); a generic method called without type arguments where the
type parameter appears only in the return, which C# resolves to the
non-generic overload and this pass leaves open (Values, 5); the 73
calls Roslyn places in metadata that the tree also declares —
Type.IsAssignableFrom, string.Trim — because the project reimplements them
for other targets under #if, and one target was compiled; and 3 402 sites
with no edge, which are new StringReader(…) and the other constructions of
framework types the reader does not record as calls.
Rust's derived traits (October 5)¶
x.clone() on an Id declared with #[derive(Clone)] reaches Id, and
no function the source writes: rustc files the derived clone on the
#[derive] attribute's line, and 124 of clap's 458 in-tree misses were calls
to a derived clone, default, cmp or eq. The Rust reader records the
traits each struct, enum and union derives (collect_derives in
core/src/parser/rust_lang.rs); a method whose name a derivable trait
writes — clone, default, eq, cmp, partial_cmp, hash, fmt —
on a receiver whose type derives it and declares no such method itself
reaches the type, placed by inheritance. The receiver is the caller's own
type for Self::default(), a type named as itself for ValueHint::default(),
the type a chain or a binding yields, or the type written for a NAME THE
SCOPE BINDS — not a field typed for the whole file: if let Some(long_version)
= self.long_version.as_deref() shadows an Option<Str> with a &str, and
the first version of the rule claimed Str::clone for it.
The oracle's line moved to the item. rustc's debug info names the
attribute's line for a derived impl, and muundo's struct begins on the line
below the attributes: scripts/corpus_rust.py now moves an in-tree target
that sits on a #[…] or doc-comment line down to the item it decorates —
the position normalised, the expectation kept. clap 79.0 % → 80.1 %,
precision 0.988 → 0.988, false edges 22 → 22. What is left: Default::
default() as a struct literal's field value, whose type is the field's (8);
==, which the reader does not extract as a call (9); a derived clone on
an Option<Str> field, which the field table types as Str (2, false);
and the 109 sites rustc files on a struct's field lines, where a derive
writes self.field.clone() and no source writes a call — kept in the
denominator, as protocol v2 requires, and published here.
Kotlin's fields, continued (October 5)¶
builder.body.add(x) inside FunSpec is FunSpec.Builder's body.
Eleven classes nest a Builder, and a field table keyed by the bare type
name either kept whichever file came last or, once two files disagreed,
nothing: the Kotlin reader files a nested class's fields under the dotted
path too (FunSpec.Builder), and a path is walked from the class the call
is written in — or, for a property initialiser, whose caller is the file,
from every class the file declares. A field filled by a factory holds what
the factory returns: internal val body = CodeBlock.builder() is recorded
as =CodeBlock.builder, and the walk follows it to CodeBlock.builder's
declared return, Builder, nested beside the factory — so the member asked
for at the end is CodeBlock.Builder's build, not one of the eleven. The
same walk serves C# and Java, whose field tables already carried types.
KotlinPoet, which the cross-hierarchy overload set and the direct-member
rule had moved from 79.2 % to 77.5 %, is measured again in the table.
C++: a bare macro line is blanked (October 5)¶
FMT_BEGIN_NAMESPACE on a line of its own is a macro, and no grammar can
read what it expands to. fmt's core.h recovered from 183 syntax errors,
the first at that line, and the regions under them were not read. A line that
is nothing but a name in capitals, with or without an argument list and
without a ;, is blanked before parsing, byte for byte, the way the
visibility macro between class and a name already was (mask_bare_macro_lines
in core/src/parser/mod.rs). What it buys is small and published in the
table: fmt's errors fall from 183 to 173 in core.h, the rest being
FMT_CONSTEXPR and FMT_API inside declarations, templates the grammar does
not close, and #if branches, which is where fmt's figure stays bounded.
Four corrections the evening measurement asked for (October 5)¶
The fourteen corpora were measured again after the C#, Rust, Kotlin and C++ work, and four rows had lost something. Each loss was traced to a rule and the rule was narrowed; the published table is the run after the corrections, in which no row lost a call or a point of precision.
- The base's normal form, in nominal languages only. TypeScript's arities
read "any number" because JavaScript checks none, so every call there
passing fewer arguments than declared looked like a spread over
params, and five excalidraw calls moved to a base class's method of the same name. The rule now applies to C#, Java and Kotlin. - Private access is judged by class, not by file. C++ defines
Table::Openintable.ccand declares the privateTable(Rep*)intable.h; comparing qualified-name prefixes made the constructor unreachable from its own class and left the deleted copy constructor as the only candidate — one false edge on leveldb. - A sole candidate's arity is not checked in C or C++. A default argument
is written on the declaration in the header and the definition the tree
indexes repeats none, so
TableCache::NewIteratortook four intable_cache.ccand refused the call passing three. Rust keeps the check: relaxing it there let three clap calls through to a method that could not take them. - A field's factory is found at the line that filled it.
bodywas filled byCodeBlock.builder()on line 337, and the call nearest above the use on line 538 wasParameterSpec.builder(…); the "type beside the factory" rule now looks the factory up at the binding's own line. A Kotlin field writtenkdoc: CodeBlock.Builderkeeps its dotted type, so the member asked of it isCodeBlock.Builder's and not one of elevenBuilders'.
A Kotlin lambda's receiver (October 9)¶
buildCodeBlock { add("\n") } reaches CodeBlock.Builder.add. The
lambda passed to buildCodeBlock(builderAction: CodeBlock.Builder.() -> Unit)
runs with a CodeBlock.Builder as this, and a bare call inside it asks
that receiver before the class it is written in. The reader records each
trailing lambda that has a receiver and where the receiver is read from
(collect_lambda_receivers in core/src/parser/kotlin.rs): the function
called, when every declaration of that name writes T.() -> U as its last
parameter; with(x), x.apply, x.run, from the type written for x;
Owner.factory().apply, from what the factory returns; Type().apply.
also, let, takeIf and takeUnless give it and no receiver.
Kotlin's order is kept (a_member_of_the_lambdas_receiver in
core/src/analyzer.rs, first of the bare-name rules): a local function, a
local or a parameter first; then each receiver from the innermost outward;
then the class's own this, which the rules after it answer as they did. A
receiver whose member cannot take the call passes it outward: by count
(add(1, 2, 3) against a one-parameter add), and where the call passes
this to a member whose every declaration types that parameter as
something the receiver is not — KotlinPoet's toString(): String =
buildCodeString { emit(this, null) } is the class's own
emit(CodeWriter, String?), never CodeWriter.emit(String, Boolean). For
that check Kotlin's this is now an argument shape (this_expression), as
Java's was; a labelled this@Outer is not. A receiver whose type is not
known stops the walk, and the call is answered as before. The edge is placed
by written_type, or return_type when the receiver came from a call.
Measured against the oracle the same day, once a JDK and kotlinc were
installed in the container, against the October 5 engine measured the same
way: klaxon 69.1 % -> 75.7 % (301 in-tree calls, precision 1.000),
KotlinPoet 79.8 % -> 80.6 % (857, precision 0.993), no false edge added.
Before the oracle, muundo's own report before and after showed KotlinPoet
52 calls and klaxon 73 moving from undetermined to in-tree, 6 emit("…")
moving from the class's own emit to CodeWriter.emit — which the class's
emit(out: CodeWriter) could not have taken — and 3 more overloads named
(out.lookupName(this)); every sample read against the source was right.
The proof is core/tests/a_lambda_has_a_receiver.rs.
What is left: a receiver filled by a call on a local
(val b = builder(); b.apply { … }) where no type is written; this@label
naming the lambda; and an argument other than this whose type refuses the
receiver's member, which only the count guards today.
A parameter holds what its callers pass, in four more spellings (October 9)¶
The Python rule that follows a called parameter to its callers
(a_parameter_its_callers_fill) read one spelling of argument: a plain name
declared somewhere. PyCG writes four more, and each now names the function
when the file says which, and refuses otherwise:
- A call that hands a function back by name.
func(func2()), wherefunc2returnsfunc3: the reader recordsfunc2()(no arguments) as written, and the table of what each declaration returns by name answers. - A name bound to a function.
b = param_func; c = func; c(b):bis read through the binding the collector recorded. The call sites of a declaration are now filed under the name the source WROTE (c), which is how the argument rows are keyed; under the callee's own name (func) the row was never found. - A method of an instance whose class is constructed where the name was
bound.
i = MyClass(); i.m1(i.m2, i.m3). A method'sself(orcls) is declared and not passed, so positions shift by one; a call spelled through the class,MyClass.m1(i, …), passes it and is refused rather than guessed. - A parameter passed on.
s(t)insidem1passest, itself a parameter ofm1: the rule follows it up tom1's own callers, three levels at most.
Every caller must still name one and the same function; two callers passing
two functions say nothing (lambdas/calls_parameter stays a miss by design —
an edge has one callee). Measured: PyCG 71.6 % -> 76.1 %, 185 of 243
pairs, no wrong pair added, precision 0.995. The proof is the second test in
core/tests/the_shapes_pycg_taught.rs.
Kotlin, continued: delegation, safe calls, properties anywhere in the class (October 9)¶
Measured against kotlinc the same day, each step on both corpora; klaxon 75.7 % -> 85.7 %, KotlinPoet 80.6 % -> 81.0 %, no edge lost anywhere.
with(stateMachine) { put(…) }wherestateMachineis a property typed by what was constructed into it. The lambda-receiver rule read written types only; it now falls back on the field table for a property of the class around the call. The table records a class's property as a local of the class's own scope, so the "a local shadows it" guard is asked of functions and blocks only (bound_in_a_function_or_a_block). klaxon's parser: 16 calls.- A member a class delegates with
byreaches the class.class JsonObject(val map: MutableMap<…>) : MutableMap<…> by map— kotlinc writes a forwardingget,put,containsKeyinto the class, on the class's own line, and no source declares them. A bare call inside the class, or one on a receiver typed as it, reaches the class when the class declares no member of that name and the member is an abstract member of a standard interface (KOTLIN_INTERFACE_MEMBERS: Map, MutableMap, List, Collection, Set, Iterable and their mutable forms, CharSequence, Comparable). An interface the tree declares keeps the hierarchy's answer — its own member, the declaration that runs through the delegate. klaxon declaresfun JsonObject(…)aboveclass JsonObject, so the edge names the class's line (type_line_by_qn), and the overload pass now leaves a line the resolver already decided alone. klaxon 75.7 % -> 84.4 %, its one wrong edge and its one falseoutsidegone. obj?.int(n)andpeeked!!.value()keep their receiver. The resolver readobj?as no receiver at all and kept the bareint, which a bare function of that name could then claim. The Kotlin reader writes the callee with.in their place: a safe call and a non-null assertion change what happens to a null, not the declaration reached.- A property types a receiver anywhere in its class.
lexer.nextToken()in a method aboveval lexer = Lexer(reader): the per-file tables look above the call only, and a Kotlin property is visible in its whole class.through_a_property_of_the_classreads the field table for a receiver that is ONE plain name — not the last segment of a path:funSpec.body.isEmpty()handsbodyas the hint, and the first version of the rule took it for one of the elevenBuilders'body, two false edges on KotlinPoet before it was narrowed. - And a property no longer types every class of the file. A Kotlin
property's binding held from its line to the end of the FILE, the class
body not counted as a region:
val lexer = Lexer(r)in one class typed alexeranother class inherits from a base outside the tree. Its region is its class body now. Neither corpus moved; the test pins it.
What is left, by the audit: a constructor or the factory function
sharing its name, which the overload pass cannot tell apart by count
(JsonArray(…), JsonObject(…), 13 on klaxon); it inside a lambda
typed by the function type it is passed for (mapChildren { it.string(id) });
a local filled by an elvis expression (passedLexer ?: Lexer(reader)); and
a parameter that a property declared ABOVE it still types through the
binding pass (fun f(lexer: Any) = lexer.next() beside val lexer =
Lexer()), which predates this work.
Go: grouped parameters, variadics, named collections (October 9)¶
cobra 90.9 % -> 93.1 % against x/tools callgraph, precision 1.000, the
October 5 engine reproducing 90.9 % in the same container first.
func configEnvVar(name, suffix string)takes two arguments. Go groups names under one type in a singleparameter_declaration, and the arity reader counted the node: every two-argument call to such a function was refused as one argument too many, in the same file as anywhere else. The names are counted now (Go only; C'sparameter_declarationholds one declarator).cmds ...*CommandholdsCommands. The variadic declaration was not a typed name, so arangeover it had nothing to read; it is now, read as its element as[]*Commandalready was.s[i].Name()ontype commandSorterByName []*CommandreachesCommand.Name. The named type is a type of its own — it declaresLessandLen— so it is not read as its element everywhere, only where the receiver is written indexed (collection_elements, after every other rule has failed).
What is left, and why it waits. The 9 false outside are doc/
calling cobra.WriteStringAndCheck through import
"github.com/spf13/cobra" — the module's own root package. Only go.mod
says that path is this tree, and muundo reads no go.mod. Reading it is a
decision of the same kind as reading tsconfig.json: an input the report
depends on, which verify and the snapshot store must then record and
check. The other misses are a loop variable reassigned by a call
(for p := c; p != nil; p = p.Parent()), which the table cannot type
without knowing Parent returns what c is.
C#: a value type before its nullable form (October 9)¶
Newtonsoft.Json 90.5 % -> 91.5 % against Roslyn, the same 97 wrong answers and no other, the October 5 engine reproducing 90.5 % first.
ToverT?, never the reverse.WriteValue(uint),WriteValue(uint?)andWriteValue(object?)all take auint, and C# names the first: an identity conversion is better than a nullable one.at_least_as_specificstripped the?from both sides, souintanduint?were each as specific as the other and the most-specific rule named neither — most of the 402WriteValuecalls leftoverload_unnamed. A reference type's?is an annotation and C# cannot declare both overloads, so the rule only meets value types.- C#'s other integral keywords are value types.
uint,ulong,ushortandsbytejoin the closed types, so auint?argument is refused byWriteValue(uint)as adecimal?already was; their implicit numeric conversions are those of the C# specification, so auintstill widens to alongwhere nouintoverload exists.DateTime,Guidand the other library structs are left out: closing them would refuse an interface parameter they implement, and the specification does not list them. -
9ULis aulong. C#'s integer suffixes were read as Java's, and the first measurement of the two changes above namedWriteValue(long)for it — one false edge, gone once the literal had its type. -
A call writing no type argument does not reach an uninferable generic.
T? DeserializeObject<T>(string value)mentionsTin its return only, and C# infers type arguments from the ARGUMENTS alone — never from what the result is assigned to, as Java does. SoDeserializeObject(json)can only be the non-genericDeserializeObject(string); counting the generic one among the candidates left the overload unnamed. The C# reader records such declarations (collect_uninferable_generics) and the overload pass sets them aside for a call that writes none. 91.5 % -> 91.7 %, +40 named, no wrong answer added. Java keeps its candidates: its inference reads the assignment.
Not encoded: C#'s "better conversion target" tie-break between a signed
and an unsigned type (long over ulong for a uint). Where both fit the
overload stays unnamed, which is the answer the set gives.
Kotlin: what it is (October 9)¶
klaxon 85.7 % -> 87.7 % against kotlinc, precision 1.000; KotlinPoet unchanged.
it.string(id) inside mapChildren { … } reaches JsonObject.string
when mapChildren is declared fun <T : Any> mapChildren(block :
(JsonObject) -> T?): a lambda that declares no parameter receives the one
the function type names, as it. And x.let { it.f() } / x.also { … }
reach what x is, typed the way x.apply { } already was. it belongs to
the INNERMOST lambda that declares no parameter, so the reader records every
lambda (collect_it_lambdas), a lambda declaring parameters is looked
through, and one declaring a parameter called it stops the search.
Which declarations a trailing lambda can reach. The lambda-receiver rule
of the same day asked every declaration of the called name to agree, and
klaxon's tests declare a fun mapChildren() taking nothing — the six calls
stayed unplaced. A trailing lambda can only be passed to a function whose
last parameter is a function type, so only those are counted now
(collect_trailing_lambda_takers), with every declaration of the name in
another language counted too, since nothing here says what those take.
What is left: it narrowed by a smart cast (if (it is JsonArray<*>)
it.mapChildren(block)), which is flow typing and not written on any
declaration.
TypeScript: a member typed typeof X is X (October 9)¶
ky 93.2 % -> 94.4 % against the checker, no wrong answer added; excalidraw and express identical to the call.
KyInstance declares readonly retry: typeof retry, and ky.retry(…) —
39 of ky's calls — reaches the retry that name means in the file declaring
the member, here an import from core/constants.ts. The checker names that
declaration: the member is a property whose type is the function's own type,
and declares no function of its own. It is the reading the parameter rule
already took ("func: typeof resizeFrameOverElement names a declared
function"), applied to a member. The TypeScript reader records such members
(collect_typeof_members); the resolver, where the receiver's type has no
member of that name, resolves X from the declaring file — its own
declaration, else its imports — and claims nothing for a global (typeof
fetch) the tree does not declare.
A type alias of an object type names its members. type KyInstance = {
… } had no owner for its properties in the field index, while the entity
extractor already filed get under KyInstance. It is the owner now, as an
interface would be.
JavaScript: a function expression calls under its binding's name (October 9)¶
express 92.0 % -> 93.0 % against the checker, no wrong answer added; ky and excalidraw identical to the call.
express writes its application as an object of function expressions:
var app = exports = module.exports = {} and then app.set = function
set(setting, val) {…}. The entity has been app::set since "A property
assigned a function is under its object"; the calls in its body were filed
under <module>, because the JavaScript caller walk counted declarations,
methods and arrows but no function expression. And had it counted one, it
would have named it by its own name — set — where the entity is named by
its binding. A function expression now names its caller as its entity is
named: by what holds it, or not at all, in which case the walk goes on to
the function around it, as an anonymous arrow's does.
With the caller right, this.set(…) inside app.enabled reaches app.set
through the rule that already answered this.m() in a method: a sibling in
the caller's own scope. A function assigned to a computed property
(app[method] = function (…) {…}) has no name to give, and its calls stay
with the module.
A language server is ready when it says so (October 9)¶
--lsp waited for the server to fall silent after initialized — a quarter
of the start deadline, four seconds under the test's — and counted silence
as "the project is open". rust-analyzer can stay silent that long while
cargo metadata runs, before it reports any progress; on a loaded machine
the definition was asked too early, the answer was null, and the call was
recorded as unanswered. a_server_places_a_call_and_the_report_names_it
failed that way all day whenever other work ran on the machine, and once
with nothing of this repository's changed.
The client now advertises window.workDoneProgress and rust-analyzer's
experimental.serverStatusNotification, and reads what comes back
(Readiness): a server that reports its status is ready when it says
quiescent with no progress open, and silence never counts while a progress
it began has not ended or its last status said otherwise. A server that
reports neither is judged by silence, as before. The test passes on the
loaded machine in 8 s where it took 20.
CommonJS: a required function carries its properties (October 9)¶
express 93.0 % -> 93.6 %, no wrong answer added; ky and excalidraw identical to the call.
user.js writes module.exports = User, declares function User(…), and
assigns User.count = function (fn) {…} beside it; index.js writes var
User = require('./user') and calls User.count(…). The require bound
User — new User(…) already resolved — but a member of it went through
the CommonJS member rule, which reads the properties of an exported OBJECT
and stamped anything else unresolved. A member of the default export is now
asked first (member_of_the_default_export), under three conditions, each
the cost of a version that measured wrong:
- the receiver is still the import at the call (the CommonJS proof the member rule already asked);
- the default export names a FUNCTION or CLASS declared at module level — not
a value: ky's
const ky = createInstance()is its default export, and the first version sent 192 of ky's calls to theky.create = …written insidecreateInstance; - the member is EXACTLY
<the export>::<name>and written outside the export's own body: members filed under a scope of the same last name (createInstance::ky::create) are not the export's, and a function nested inUser's body is not a property the function object carries.
Python: a star import binds what the module publishes (October 9)¶
PyCG 76.1 % -> 77.8 % (189 of 243 pairs), no wrong pair added.
from from_module import * and then func1(): the Python reader recorded
named imports only, and the * reached nothing. It is recorded now
(star_imports), with each module's __all__ (collect_dunder_all), and a
bare call nothing in the file declares or imports by name reaches the
top-level declaration of that name in the module — public (no leading
underscore), or listed in a literal __all__; an __all__ built any other
way publishes nothing provable, and the call stays unplaced. Of several star
imports binding the name, the last one wins, as the later import overwrites
the earlier.
This is not the refusal under "Import resolver — static-only limitations":
those are star imports from namespace packages and runtime __all__. A
module file in the tree, with its top level written out, is read like any
other.
Kotlin: an extension function's receiver is this (October 9)¶
KotlinPoet 81.0 % -> 81.5 %, the right method 92.3 % -> 94.6 %, wrong answers 5 -> 4; klaxon unchanged.
KotlinPoet's JvmAnnotations.kt is a file of extension functions —
public fun PropertySpec.Builder.transient(): PropertySpec.Builder =
addAnnotation(Transient::class) — and every bare call in them reaches the
extended type, which is this there. The reader records each extension
function's body and its receiver, dotted as written
(collect_extension_bodies), and a bare call inside one asks that type
through the same member lookup the lambda receivers use: after the lambdas
around the call, before the class around the function, as Kotlin orders its
implicit receivers. A local binding the name wins, and a lambda receiver of
unknown type around the call stops the rule, since Kotlin would ask that
receiver first. A member that cannot take the call passes it on — a
two-argument addAnnotation(1, 2) reaches the file's top-level function.
Most of the 20 calls this places land on the right method with the overload
unnamed: addAnnotation(AnnotationSpec), (ClassName), (Class<*>) and
(KClass<*>), told apart only by argument types the overload pass does not
always know.
Kotlin: X::class is a KClass (October 9)¶
KotlinPoet 81.5 % -> 86.5 % on the overload, the same four wrong answers; klaxon, Newtonsoft.Json, gson and ky identical to the call.
KotlinPoet declares addAnnotation(ClassName), addAnnotation(Class<*>)
and addAnnotation(KClass<*>), and the same three for builder,
addImport, addAliasedImport and the asClassName() extensions. What
tells them apart is the argument's type, and three things hid it:
Strictfp::classhad no type. It is aKClass, as Java'sString.classis aClass; tree-sitter-kotlin-ng reads it as a navigation through::, and the argument reader now types either spelling.KClasswas an open type. Two types the tree does not declare never refuse each other, so aKClassfitted aClass<*>parameter. It is a closed type now, besideClass: no conversion makes one the other.KClass<*>?kept its type arguments. They are stripped only from a type that ENDS with>, and the nullable mark came after them, so the parameter matched nothing spelledKClass— klaxon'sfindJsonPaths(kc)lost its overload on the first measurement of the closed type. The arguments are stripped under the mark, and the mark put back.
Rust: a loop or a closure over .iter() receives the element (October 10)¶
clap 80.0 % -> 80.7 % (1 824 -> 1 838 of 2 279), the same 19 wrong answers, no call lost. The oracle counts 2 279 in-tree calls where the October 5 table counted 2 278; the baseline run of the same day, before any change, counts 2 279 too.
for g in &self.groups and self.get_subcommands().filter(|a| …) already
typed g and a by what the collection holds. Three spellings did not:
.iter()on a field or a name.self.groups.iter().any(|grp| …)looked throughiter, which hands its receiver back, found the fieldself.groupsand gave up, because it was not a call. It now ends at the field or the name and reads its written type, as&self.groupsdid. Only for a loop variable or a closure parameter:let it = self.groups.iter()holds the iterator, andit.find(…)must not reach afindthe element declares.- A slice.
list: &[Group]named no element;[T]and[T; N]now nameT. impl IntoIterator<Item = &'a mut Command>. The element is written in a binding, not as a type argument, and was read as the trait alone —outside, twice wrong on clap'sdid_you_mean_flagonce the first rule reached it. The binding's type is the element now.
What is left: a closure destructuring a pair, .filter(|(_, matched)|
matched.check_explicit(…)) over a map's .iter() (about ten of clap's
misses), which needs the binding to say WHICH component it holds.
A language-server test placed items[0].go() on items: &[S] to prove the
server's answer is used; muundo now places that call itself, so the fixture
calls through a function pointer instead.
Rust: == calls the tree's eq, and a field path types its elements (October 10)¶
clap 80.7 % -> 81.3 % (1 838 -> 1 853 of 2 279), wrong answers
19 -> 15, false outside 8 -> 4 (outside precision 0.990 ->
0.995); no new wrong answer.
A comparison is a call. rustc compiles grp.id == *g to PartialEq::eq
and files it on Id's #[derive(PartialEq)], or on the fn eq an impl
writes; muundo extracted no call there. The Rust reader now writes a == b
and a != b as a.eq, with the operator kept in wrote, so the rules that
resolve a receiver — the derived-trait rule among them — resolve it, as the
C++ reader writes *it as it.operator*. The left operand must be a name or
a field path. A comparison nothing in the tree implements is the
language's: an edge still written on its receiver after resolution is
dropped, so k == 3 and s == "x" add nothing.
A loop or a closure over a field reads the field on its struct.
self.groups.iter_mut().find(|grp| …) typed grp through the field's NAME,
and command.rs declares three groups — Command's Vec<ArgGroup>,
Arg's Vec<Id> and a method's parameter — so it was typed by none. The
binding now keeps the whole path, self.groups[], and the field table holds
each collection field's element under a name no source writes (groups[]),
so the path walk that already reads a.settings on Arg reads it. A field
written Vec<Self> holds its own struct. Where the path's head has no known
type, the field's name is read as before.
What the second change gave back. for existing in &self.inner over
Vec<T> typed existing as Vec, and existing.borrow() was answered
outside — right, since borrow is std's, but for the wrong reason. The
element is now the generic T, which nothing types, and the answer is
undetermined. One outside agreement fewer, kept as the honest answer.
What is left: a comparison whose left operand is a call or an index
(w[0] == w[1], arg.get_value_hint() == ValueHint::… on a receiver the
chain does not type), and a closure destructuring a pair over a map.
Rust: filter, rev, skip and their kin pass the elements on (October 10)¶
clap 81.3 % -> 82.0 % (1 853 -> 1 869 of 2 279), the same 15 wrong answers, none lost.
self.groups.iter().filter(…).map(|grp| grp.id.clone()) hands map's
closure a group: filter yields what it was given, fewer of them. The
element rule stopped at the first call it did not know, so only the closure
of the FIRST adapter was typed. For a loop variable or a closure parameter,
the chain is now read past filter, rev, skip, take, skip_while,
take_while, peekable, inspect, step_by, fuse, by_ref and cycle
— back to the field, the name, or the call whose impl Iterator<Item = …>
says what is yielded. Never past map, enumerate or zip, which change
it; and never for a let, which holds the adapter itself, so let mut it =
….filter(…); it.find(…) does not reach a find the element declares.