Skip to content

feat(ai): establish layered code-intelligence stack - #144

Open
dporkka wants to merge 75 commits into
mainfrom
codex/code-intelligence-stack-20261003
Open

dporkka wants to merge 75 commits into
mainfrom
codex/code-intelligence-stack-20261003

Conversation

@dporkka

@dporkka dporkka commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • make codebase-memory-mcp the primary graph MCP for Dev Plane while retaining GitNexus as a migration fallback
  • add shared Zoekt indexing for canonical local clones
  • replace vendor-specific agent guidance with a stable packages/repo-intel routing contract
  • add deterministic routing for locate / understand / impact / refactor / security / history / verify
  • add executable GitNexus vs codebase-memory benchmarking against current production paths in Adacavo, Nulang, Nulang Cloud, and Dev Plane
  • add a recall-first scorer and fail-closed codeintel-promote migration gate
  • add exact source/tool/corpus provenance to every benchmark suite
  • add a pinned container benchmark worker and Woodpecker PR + manual live-bakeoff workflows

Architecture

Shared: codebase-memory + Zoekt canonical indexes.
Per worktree: native LSP, ast-grep, rg/fd, compiler/tests.
On demand: Joern. SCIP remains deferred until a measured cross-repository semantic gap justifies another index lifecycle.

packages/repo-intel exposes stable intents so callers do not depend on GitNexus/codebase-memory-specific tool names.

Benchmark ground truth

The corpus is a curated set of required core production paths, not an exhaustive relevance set.

It currently covers:

  • Adacavo proposal lifecycle + durable proposal booking, including the legacy lifecycle implementation and transactional booking path
  • Nulang MIR-to-bytecode codegen + core production callers in CLI, REPL, DAP, AOT, and C API
  • Nulang Cloud's Dev Plane provider, local Wasmtime runtime, and Firecracker host-agent/VM factory boundary
  • Dev Plane scheduler/readiness/budget/executor admission chain and worker wiring

Broad codebase-memory scenarios use deterministic BM25 search_graph(query=...); the Nulang caller scenario uses trace_path. Every codebase-memory call is explicitly scoped to the indexed project.

Promotion rule

The migration gate is deliberately recall-first:

  • every codebase-memory candidate run must succeed
  • required-core-path recall must be 1.00 on every scenario
  • candidate recall may not regress against a successful GitNexus baseline on any scenario
  • precision/F1 are diagnostic only because extra genuinely useful files are not necessarily represented in the non-exhaustive gold set
  • latency/token cost matter only after complete required-core recall

codeintel-promote exits non-zero for an unsafe result. GitNexus is not removed unless the live suite returns promote: true.

Reproducibility

The container benchmark worker pins Node 24, Go 1.26.8, codebase-memory-mcp@0.11.0, and gitnexus@1.6.12, invokes the installed binaries directly, and isolates backend caches.

Final suite.json records:

  • exact Git SHA for every repository
  • backend/runtime versions
  • corpus version + SHA-256
  • exact prompt, intent, required paths, and backend MCP calls for every scenario
  • raw/scored backend observations

The normal portable lane is:

bash scripts/verify-codeintel.sh

A pinned local live run is:

export CODEINTEL_REPOS_ROOT="$HOME/src"
docker compose -f docker-compose.code-intelligence-benchmark.yml build benchmark
docker compose -f docker-compose.code-intelligence-benchmark.yml run --rm benchmark

Woodpecker

.woodpecker/code-intelligence.yaml runs harness/corpus/provenance tests and repo-intel Go tests on the existing linux/amd64 pool.

.woodpecker/code-intelligence-bakeoff.yaml is manual-only. It needs Dev Plane activated in Woodpecker plus repository secret codeintel_github_token with read-only Contents access to the private Adacavo and Nulang Cloud repos. The workflow removes/unsets the temporary Git credential before indexing/querying.

See docs/code-intelligence-bakeoff-runbook.md for the complete operation/migration sequence.

Verification status

TDD covers routing, scorer behavior, corpus contracts, provenance binding, promotion policy, and CLI exit behavior.

The chat execution sandbox cannot run the authoritative final suite: it lacks the required Go 1.25.11+ toolchain/backends/canonical private clones and cannot download them because outbound networking is restricted. The repository-side portable and Woodpecker lanes are therefore the authoritative execution paths for the final branch state.

GitHub-hosted Actions remains independently broken at admission level. Main-branch run 36910222862 already showed the same failure mode before this PR: Lint, Build, and Test were created but executed zero steps. Later PR runs show the same zero-step/no-log behavior. Do not interpret those runs as compile/test failures.

Migration sequence

Do not remove Adacavo's GitNexus-specific safety/freshness assets in this PR.

After a trusted worker produces promote: true:

  1. repeat once from a cold benchmark cache
  2. remove the disabled GitNexus fallback from Dev Plane only
  3. run/bake in Dev Plane with codebase-memory as the sole graph backend
  4. rerun the live corpus
  5. migrate Adacavo's GitNexus safety/freshness integration in a separate rollback-friendly PR

Nulang and Nulang Cloud already point at codebase-memory and do not need a GitNexus-removal step from this migration branch.

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, you can upgrade your account or add credits to your account and enable them for code reviews in your settings.

dporkka added 30 commits October 3, 2026 20:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant