Skip to content

feat: publish, navigate, and stabilise a budgeted community hierarchy - #332

Merged
forhappy merged 26 commits into
mainfrom
codex/community-hierarchy-stability
Sep 24, 2026
Merged

forhappy merged 26 commits into
mainfrom
codex/community-hierarchy-stability

Conversation

@forhappy

Copy link
Copy Markdown
Contributor

What this does

Community detection publishes a flat partition whose size scales with the repository — 112 communities for pallets/flask, 2,781 for colinhacks/zod — so the overview is a dot field and a screenshot stops matching the build a week later. This branch adds a budgeted, versioned hierarchy: compass.community-hierarchy/1 beside community-quality.json, level navigation in the viewer, and a measured stability gate.

It implements plan 027 (build + export + viewer) and plan 028 (stability + qualification), 25 commits.

The hierarchy

  • Levels are defined over the level below. Level 0 is a bounded root, the last level is the published partition itself (group i pairs with community i; ids can be sparse after an incremental remap), and childIndices index the next finer level, so the artifact never repeats node ids.
  • Two merge rules, each recorded per level. Relationship levels run the seeded Leiden local moving on the level below at halved resolution. Where evidence cannot reduce a level, levels are cut from the directory tree the groups already cite: measurable, because 86 of flask's 112 and 2,725 of zod's 2,781 communities have no cross-community edge at all. A group citing no dominant directory stays its own group, and budgetSatisfied reports the achieved count instead of merging it away.
  • Labels carry provenance: dominant directory (≥60% of members) → module prefix → hub member → generic community id, with the counts that produced the choice in label.evidence. Consumers read rule/generic, never label text.
  • Completeness is proved before publication: children partition the level below exactly once with matching member counts, or the build fails with a typed error.
  • Group ids are evidence: h<level>-<signature16> over the group's member signatures, with a per-level signature and hierarchy-signature/v1 on the artifact.

Measured roots: flask 24 groups, budget satisfied; zod 70 groups (24 directory buckets plus 46 groups that cite no dominant directory) with budgetSatisfied: false — the honest shortfall rather than a merge nothing supports.

Navigation

  • Exports embed compass.viewer.hierarchy/1 beside the model; the toolbar scope reads Level 0 … N | Symbols; double-clicking a group descends into its children inside the level below; the breadcrumb walks back up one group or to the repository; the deepest click opens that community's symbols.
  • All four overview designs (bubbles, matrix, area, tiers) render whichever level is open, and a narrowed scope lists only the groups it draws.
  • compass export hierarchy-json reproduces the artifact byte for byte and rejects an unknown major or a mismatched graph; compass export html|workbench-json --hierarchy-level N opens on a chosen level; the workbench code view carries the levels with coverage.hierarchyLevels.

Stability

  • Reconciliation: every rebuild matches groups against the published hierarchy by member overlap, hands a surviving group the id a reader already learned, and reports split/merged/appeared/disappeared/ambiguous. Splits and merges transfer nothing, parent matches gate a child's inheritance, ambiguity is reported rather than resolved, events are bounded (256) with an exact omittedEvents, and membership, labels, and evidence are never rewritten. community-hierarchy.json.sig persists the identity ledger.
  • compass.community-hierarchy-diff/1: identity-first comparison of two generations, riding on the history workbench view (kind: "history", hierarchyDiff) with one bounded banner line. Absent when either side published no hierarchy — unavailable, not "nothing changed".
  • Level-stable layouts: level projections name nodes after their groups, descending maps child indices to ids, and the overview seed takes its bearing from identity while rank still sets the radius.

Verification

  • cargo fmt --all -- --check; cargo clippy -p compass-graph -p compass-core -p compass-history -p compass-semantic-diff -p compass-output -p compass-cli --lib --bins -- -D warnings
  • cargo test for those six crates (--lib --bins plus the touched integration suites): compass-graph 13 suites, compass-core 86 lib, compass-semantic-diff 5+5, compass-output 65 lib, compass-cli 93 lib + product 9 + viewer-export 13 + help 5
  • ./scripts/qualify_code_graph_v1.sh --hierarchy — 3 fixtures × 2 runs byte-identical, all six acceptance entries true
  • ./scripts/qualify_code_graph_v1.sh --hierarchy-stability — add/move/delete replays byte-identical, all eight entries true; measured ARI/AMI 1.0 between generations (threshold 0.5), untouched ids kept 5/5, 6/6, 5/5, exactly one disappearance at the delete, and a forced even split reported ambiguous with no id inherited
  • npm run typecheck:js; npm test -w @compass/viewer — 300 tests; cd tests/viewer && npx playwright test --project=chromium — 97; node scripts/check_viewer_assets.mjs; sh scripts/check_product_boundary.sh

Two pre-existing failures on this machine, identical on base 20848d46, unrelated to this branch: eight compass-core integration tests in loading_coverage / code_graph_v1_publication_resilience, and one --all-targets clippy lint in crates/compass-output/tests/agent_query.rs (CI runs --lib --bins).

Compatibility

Additive and versioned. graph.json, graph-overview.json, orientation.json, CompassQL, MCP results, and query semantics are untouched; no query path reads the hierarchy, and absence means navigation is unavailable. COMPATIBILITY.md states the group-id contract, that a changed signature means changed membership, that consumers must handle ambiguous, and that any signature-algorithm or threshold change needs a version bump.

Two adaptations, recorded in the plan

  • Reconciliation compares member nodes, but the artifact deliberately stores no node lists, so the published partitions are supplied alongside the two hierarchies.
  • The diff is an additive optional field on the history workbench view rather than a change to the frozen compass.semantic_diff.report/1 schema.

Known gap, deliberately deferred

At relationship levels the dominant-directory rule can name sibling groups identically (pallets/flask publishes 23 groups labelled tests at level 1) because neither group's members concentrate in one subdirectory. Fixing it means appending a qualifier whose coverage is below the label threshold — a label-policy decision recorded in the plan's status note rather than guessed at here.

`cargo fmt --all` reshapes three call sites that #329 added: the node
colour insert in the HTML renderer, the Obsidian colour-group radix
parse, and the palette wrap test.

`cargo fmt --all -- --check` fails on the current `main`, which stops the
`quality` job at its Formatting step and skips Lints and Native tests.
No behaviour changes here; the diff is rustfmt's own output.
Community detection publishes a flat partition whose size scales with the
repository, so a committed reader cannot scan it on a large codebase. Derive
`compass.community-hierarchy/1` from that partition instead: recursive
grouping under an explicit level budget, evidence-derived labels with
provenance, and an exact coverage proof.

Levels are defined over the level below. Level 0 is the coarsest, the last
level is the published partition itself with one group per community, and
`childIndices` index the next finer level, so the artifact never repeats node
ids. Coarsening reuses the seeded Leiden local moving clustering already runs,
halving the resolution until a level fits its target and recording the
achieved count with `budgetSatisfied: false` when the floor is reached first.
Labels fall back in a fixed order: dominant directory, module prefix, hub
member, then a generic community id. Group quality mirrors the published
per-community metrics, so the finest level reports the same cohesion and
conductance the quality artifact already carries.
Publish the budgeted community hierarchy beside community-quality.json from
both clustering entry points, bound to the same graph generation and digest,
so a reader can open the overview at a bounded level instead of a flat
partition sized by modularity alone.

The hierarchy is derived from the partition the build just published and the
same typed topology projection, never re-read from the artifacts it describes.
It joins ROOT_ARTIFACTS, the build-state required-artifact list that records
each artifact's digest and byte size, and the removal branch that drops
community sidecars when a fact-neutral incremental build stops clustering.

An unclustered rebuild that leaves a sidecar behind keeps the existing
graph-bound contract: the artifact fails `validate_for_graph` against the new
graph instead of describing it.
Add `compass export hierarchy-json [--graph PATH] [--output PATH]`, which
reproduces the artifact atomically published with the selected graph
generation and refuses anything else: an unknown schema major, a missing
artifact, or a graph identity mismatch exits non-zero with the typed error
instead of emitting a hierarchy that describes another graph.
The export emits the published bytes exactly, so stdout carries no extra
trailing newline, and a root-projection copy never stands in for the resolved
artifact a consumer reads.
98% of colinhacks/zod's 2,781 communities share no cross-community
relationship at all, so relationship evidence alone can never bound a
root level. Cut the levels above out of the directory tree those groups
already cite: the cut starts at the repository root and repeatedly
expands the largest directory whose children still fit the budget, so
every merged group is a real directory its members share, and a group
that cites no dominant directory stays a group of its own.

Levels now record the rule that produced them with its evidence, the
artifact records the merge policy, and a relationship pass must remove
at least a tenth of a level to be worth publishing so a near-duplicate
level no longer costs megabytes of sidecar.
Describe `compass.community-hierarchy/1` where readers look for it: the
artifact table and a section in the output reference, a "hierarchy levels
and budgets" section in the community-detection concept, the `--no-cluster`
inventory, the CLI help topic for `compass export hierarchy-json`, and the
compatibility note that the artifact is additive and never required to
answer a query. The plan records the measured STOP condition and the merge
policy it was resolved with.
Export the published hierarchy beside the model the standalone page already
carries, as `compass.viewer.hierarchy/1`: one bounded `compass.viewer.graph/1`
projection per level plus the evidence a reader inspects — the rule that
merged the level, each group's label and provenance, cohesion, conductance,
boundary kinds, and the child indices that make descending exact. A level with
more groups than the export's node budget publishes its groups and no model, so
the payload stays bounded on repositories with thousands of communities, and
graphs without a published hierarchy embed nothing.

Both build paths pass their levels through, so `graph.html` from
`compass update` and from `compass init`/recluster describe the same tree.
Parse `compass.viewer.hierarchy/1` with the same strictness as the rest of the
workbench contracts: unknown schema majors, merge rules, and label rules fail,
and a label must say whether it is generic instead of leaving a reader to guess
from its text.

`hierarchyLevels.ts` holds the pure level logic the canvas will use: which
model a scope draws, how a group narrows to its children inside the level below
without redrawing it, the breadcrumb trail, and whether every published level
carries a drawable projection.
Open an export that carries a hierarchy on level 0 instead of a derived
overview, and read the levels as a zoom: `Level 0…N | Symbols` switches the
whole level, double-clicking a group descends into its children inside the
level below, the breadcrumb walks back up one group or all the way to the
repository, and the deepest click opens that community's symbols when the
export embedded them.

The level scope drives the designs too, so the bubbles, matrix, area, and tier
views read whichever level is open. Levels are named from the hierarchy rather
than from `labels.json`, where index 0 is an unrelated community, and the
narrow-stage rule keeps the level numbers visible because hiding them would
leave four identical icons.
Descending kept the whole level's community list, so the panel listed groups
outside the scope with zero members. A narrowed projection now publishes the
nodes, edges, stats, and communities it actually draws.
`compass export workbench-json --code-graph` now loads the hierarchy published
beside the selected graph, validates it against that graph's generation and
digest, and carries its levels on the code view with the bound recorded as
`coverage.hierarchyLevels`. A missing artifact stays valid — older graphs and
unclustered builds have no levels — while a present artifact that describes
another graph fails the export instead of shipping navigation that lies.

The viewer contract and the workbench graph both accept the payload, so the
VS Code surface can navigate levels without waiting for an HTML export.
`compass export html` and `compass export workbench-json` accept
`--hierarchy-level N` and open the page on that published level, defaulting to
level 0. A level the artifact does not have, or a graph whose build published no
hierarchy, fails with a bounded error naming the level and the artifact rather
than rendering a different page.

The embedded view carries `initialLevel`, the viewer opens on it, and a level it
cannot draw falls back to the root.
The workbench renderer carried the hierarchy but not the level the export asked
for, so `--hierarchy-level` only affected standalone pages. It now forwards the
payload's `initialLevel` to the canvas.
`./scripts/qualify_code_graph_v1.sh --hierarchy` runs
`compass.community-hierarchy-qualification/1` twice, byte-compares the reports,
and requires every acceptance entry: the root budget is met wherever a
repository's shape can support it, the tree partitions exactly (checked
independently of the builder), labels carry provenance, at most a quarter of
root labels are generic, digests repeat, and no fixture exceeds its level
budget.

Three fixtures: clustered directories merge on relationship evidence,
communities that share no relationship merge on shared location, and groups
that cite no location keep their own level with `budgetSatisfied: false`. The
report is committed beside the quality qualification it mirrors.
Steps 1–7 of plan 027 are implemented; the status note now points at the
qualification command that proves it and names the deferred label-policy gap
rather than leaving it to be rediscovered.
The implementation notes describe how the artifact is derived, bound, and
proved complete, and how the two merge rules fill the budget; the advisor-plan
index moves 027 to DONE.
Replace position-derived group identity with evidence ids:
`h<level>-<signature16>`, where the signature digests the group's sorted
member-community signatures, plus a per-level signature over its groups and the
`hierarchy-signature/v1` algorithm recorded on the artifact. `index` stays
presentation order.

A group therefore keeps its id when unrelated groups appear or move around it,
and a node-order permutation leaves every id untouched — the prerequisite for
reconciling one build against the next.
Add `reconcile_hierarchy`, which matches groups between two builds by Jaccard
overlap of their member nodes, hands the previous id to a group that survives,
and reports what happened: stable, split, merged, appeared, disappeared, plus
`ambiguous` when two successors are within the policy margin, because the
evidence does not name a single heir.

Reconciliation is identity only: membership, labels, and evidence are never
copied or rewritten, a split or a merge transfers no id, parent matches gate a
child's inheritance so a reorganised subtree reports as new rather than as a
chain of guesses, and the event list is bounded with exact omitted counts. The
artifact is resealed afterwards so its digest still describes it.
Both clustering entry points load the hierarchy published beside the graph they
are rebuilding, reconcile the fresh one against it, and publish the result with
an identity ledger at `community-hierarchy.json.sig`: flattened group ordinal
to `"<id> <signature>"`, written beside the artifact, inventoried with the other
required artifacts, and removed with `--no-cluster` like the rest.

A typed graph keeps its communities as typed metadata rather than a legacy
attribute, so the previous partition is read from the typed document when there
is one; a schema-less graph still falls back to its node attributes.
`./scripts/qualify_code_graph_v1.sh --hierarchy-stability` replays one fixture
through add / move / delete edits, reconciles each generation against the one
before it, runs the whole sequence twice, and byte-compares the reports. Every
acceptance entry is measured, not asserted: root budget and tree completeness,
ARI/AMI between consecutive generations (threshold 0.5, measured 1.0 on the
fixture), untouched groups keeping their ids, change events with no false
alarms, a forced even split reported as ambiguous without inheriting an id, and
repeatable digests.

Stable matches are now a count rather than event entries, so the bounded event
list carries only what changed instead of one row per surviving group.
A level projection now names each node after the group it stands for, so the
payload carries durable ids instead of level positions: descending maps child
indices to group ids, and a rebuild that renumbers a level no longer renames
what the reader is looking at.

The overview seed takes its bearing from that identity rather than the rank, so
an untouched group keeps its direction when a neighbour grows, appears, or
disappears; rank still sets the radius, so the most important group anchors the
centre, and the packing pass still guarantees non-overlapping bubbles and label
room. Reordered-but-equivalent input keeps producing identical positions.
Describe the evidence-derived ids, the reconciliation policy with its
thresholds, the ambiguity outcome, the identity ledger, and the measured
stability acceptance entries, and record the stability thresholds and edit
sequence beside the quality qualification they extend.
Add `compass.community-hierarchy-diff/1`: comparing two published hierarchies
is identity-first, so a group whose id survives is stable, and where an id does
not survive the members decide — one base group reappearing across several
target groups is a split, several folding into one is a merge, and a base group
with two candidates inside the ambiguity margin is reported as ambiguous with
both named rather than resolved. Events carry the overlap and member counts,
are bounded with an exact omitted count, and the payload is digested.

The comparison rides on the history workbench view (`kind: "history"`,
`hierarchyDiff`) and is absent when either realization published no hierarchy,
so absence means unavailable rather than "nothing changed". The viewer's history
banner shows one bounded line about the community structure.
Steps 1–6 are implemented and verified; the status note names the two
adaptations (partitions supplied beside the artifacts, and the diff carried on
the history workbench view rather than inside the frozen semantic-diff report
schema) and the advisor-plan index moves 028 to DONE.
The Django-sized fixture spent 2.65 s of its 3 s startup budget packing 3,400
bubbles, which pushed `performance.spec.ts` over its budget on CI runners even
though it passes on a laptop. Two changes bring that to 0.16 s:

- pair bookkeeping is numeric (dense entry indices) instead of a string key per
  candidate pair on every pass, and the candidate filter compares indices
  rather than ids;
- the pass count is bounded by the work it can do (24k pair-passes) so a
  repository-sized overview spends eight passes instead of forty-eight, while
  the 174-bubble fixture that asserts label clearance keeps its full budget.

Measured on the fixture: seed and pack 2,654 ms -> 160 ms, and the spec's own
wall time 2.9 s -> 1.3 s locally.
@forhappy

Copy link
Copy Markdown
Contributor Author

Addendum: this branch also fixes main's failing JS gate.

The javascript-and-vscode job on main (run for f3a8711e) fails at npm run test:js → tests/viewer/performance.spec.ts "Django-sized community overview opens directly with the static profile": the first paint took 4,310 / 4,917 / 5,315 ms against its 3,000 ms budget. The spec predates this work, but the community overview introduced in #329 made the packing pass part of the load path, and packing 3,400 bubbles cost 2.65 s of the budget on its own.

50275846 bounds that:

before after
seed + pack (3,400 bubbles) 2,654 ms 160 ms
performance.spec.ts wall time, local 2.9 s 1.3 s

Pair bookkeeping is now numeric (dense entry indices) instead of a string key per candidate pair per pass, the candidate filter compares indices rather than ids, and the pass count is bounded by the work it can do (24k pair-passes → eight passes for a repository-sized overview) while the 174-bubble fixture that asserts label clearance keeps its full 48.

Left as-is deliberately: the spec's 3,000 ms budget. The measured cost is now ~10% of it locally, and the fix removes the cause rather than the gate. If CI still reports it over budget on a slower runner, that number is the evidence to raise it with.

@forhappy
forhappy merged commit 11fb778 into main Sep 24, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant