The GitHub App is the primary interface. Architecture decision: ADR 0003. This file is the operational reference.
| Permission | Access | Why |
|---|---|---|
| Repository contents | Read | Read manifests, lockfiles and source (codeload tarball + API) |
| Pull requests | Read | Diffs, to analyse dependency changes in the context of the PR |
| Checks | Write | Create check runs and code annotations — the reporting surface |
| Metadata | Read | Implicit, granted to every app |
Explicitly not requested: issues, actions, contents: write, pull_requests: write, administration, secrets, or anything else. M4 remediation PRs will require a deliberate, separately-communicated permission change.
pull_request(opened, synchronize, reopened)push(configured branches)check_run(rerequested: the re-run button on the GhostDeps check)installation,installation_repositories(setup and initial scan)
No other events are subscribed. The manifest (packages/github-app/app.yml) lists pull_request, push and check_run. GitHub delivers installation and installation_repositories to every app automatically, so they cannot be listed. checks: write already subscribes the app to check_run and check_suite; check_run is listed anyway because re-runs depend on it, and GitHub sends rerequested only to the app that created the run (GitHub docs). check_suite is not used: push is the push trigger, and using both would double up jobs. A test in packages/github-app fails if the manifest and this page drift apart.
- One
ghostdepscheck run per analysed head SHA; duplicate deliveries collapse idempotently. Jobs are keyed by (repository id, head SHA) only: if the same head is opened against, or retargeted to, a different base, the diff-context analysis from the first base stands until the head moves. - Conclusions:
successwhen nothing notable is found (quiet summary: "No significant dependency issues found."),neutralwhen there are findings worth review. Run-level notes (info findings about the whole run, such as theunusedconfidence cap or an incomplete scan) never count as findings: they are listed in a Notes section after the findings and stay out of the title, the count and the confidence groups. A run with notes but no findings isneutral, titled "Analysis incomplete - see notes", never the quietsuccess, because a note can mean an adapter failed or the scan was cut short. Only a run with no findings and no notes getssuccess(#195). The app groups findings only through core'sfindingGroup(#239, #209) and never classifies them itself.verdictfindings are the findings above.incompletenotes (engine caps and incompleteness, such as an adapter failure, a partial scan or an unverified package with no imports) are listed in Notes, under their dependency when they have one, and make a finding-free runneutralas described. A plain adapternote(such as unavailable graph edges) is listed in Notes too, but a run whose only notes are these keepssuccess.awarenessfindings (such as cross-ecosystem overlap) go in a collapsed "awareness notes" section and never change the conclusion, the title or the count. Notes include the app's own status notes when the app itself skipped a step: a source-only PR whose diff was too large or unavailable to read in full gets "Removed-usage check skipped" (core's notes cover what core decided, and the app never repeats them). When no finding is high confidence (for example while theunusedconfidence cap applies), the lower-confidence group is shown expanded. GhostDeps never concludesfailure— it advises, it does not gate. - If the analysis queue is full, the job is dropped and GitHub will not redeliver. GhostDeps then writes a completed
neutralrun titled "GhostDeps was busy" that asks for a new push, at most one per repository per minute. Per-installation scheduling and the durable-queue move are designed in durable-queue-fairness.md (#258). Writing check runs needs the app id (APP_ID). - Annotations attach findings to the manifest or source lines they came from.
- The current App posts no PR comments. Comments are deferred: they need a separate write-permission change and product decision. Findings are reported through Checks. When comment delivery is enabled, a repository opts out with
"commentsOff": truein its.ghostdeps.json; the App skips maintaining the comment and logs the suppression, and the check result stands either way.
Event filtering lives in packages/github-app/src/events/. Re-runs bypass the queue's per-SHA duplicate collapse (each click gets its own job key); the reporter still updates the one check run for that SHA. The rules are default-deny: anything not listed here short-circuits quietly with no job and no API calls beyond the changed-file lookup.
| Event | Accepted when | Changed files from | Result |
|---|---|---|---|
pull_request |
action is opened, synchronize or reopened |
PR files API (first 5 pages / 500 files inside the delivery) | job if a manifest/lockfile or .ghostdeps.json changed; with the source trigger on (default), also if an analysable source file changed |
push |
branch push to the default branch; not a tag, not a deletion, not a brand-new branch | payload commits[] when under 20 commits, otherwise the compare API (capped at 300) |
job if a manifest/lockfile or .ghostdeps.json changed |
check_run |
action is rerequested on the ghostdeps check run (the re-run button) |
n/a | job for that head SHA, always |
installation, installation_repositories |
always (handled separately) | n/a | created or added only: resolve each new repository’s default branch head and queue a full scan; deletion, suspension and removal do not scan |
| anything else | never | n/a | skipped |
The repository scope config .ghostdeps.json (#354) always triggers at the root: an edit re-interprets every scan, so the PR gets a check disclosing the old/new config comparison (or, on a push, a full scan under the new scope). A config change also keeps the PR out of source-only classification.
Manifests and lockfiles are matched by file name at any depth: package.json, package-lock.json, npm-shrinkwrap.json, pnpm-lock.yaml, pnpm-workspace.yaml, yarn.lock, bun.lock, bun.lockb, pyproject.toml, poetry.lock, uv.lock, Pipfile, Pipfile.lock, setup.py, setup.cfg, requirements*.txt/.in, requirements/*.txt, Cargo.toml, Cargo.lock, go.mod, go.sum, go.work.
Source trigger (on by default). Source-only pull requests are analysed too, because removing the last import of a dependency is what makes it unused (#101). Set GHOSTDEPS_SOURCE_PR_TRIGGER=false (or 0, or the app option sourcePrTrigger: false) to go back to manifest/lockfile-only triggers. Analysable source files (pull requests only) are the extensions an adapter scans for imports: .ts, .mts, .cts, .tsx, .js, .mjs, .cjs, .jsx (JS/TS, .d.ts included), .py and .pyw (Python, not .pyi stubs), .go and .rs (#297). Paths under core's excluded directories (node_modules, dist, build, vendor and the rest) and generated bundles (*.min.js, *.map) don't count. The analysis is PR-scoped through pullRequestChanges, so a source-only PR passes an empty change list, never a full-repository verdict, and quiet PRs keep the quiet success check. When the PR's diff was read in full, the worker also passes extractDependencyChanges(...).sourceLineChanges as core's pullRequestSourceChanges. The adapter reports removed usages from those lines, and core's policy turns a removed last usage into the removed-last-usage verdict. If the diff can't be read in full, a PR that changed a manifest or lockfile, or whose file list was capped, is analysed in full with a Notes line saying so. Re-runs rebuild the file list from the PR files API with the same caps (#196), so a re-run scopes exactly as the first run did for the same head SHA. If that lookup fails, the re-run takes the capped-list path (full analysis with the note). A PR whose complete file list was source-only stays PR-scoped with an empty change list and a quiet check, so it never gets repository-wide verdicts, but it also gets no removed-last-usage. When its diff was not read in full, any excluded-fixture-path disclosure on that check is stated as a minimum count, not the PR total (#354). Pushes stay manifest/lockfile-only. One job per head SHA, as before.
Renamed files count under both their old and new names. When a file list was capped or the lookup failed, the change is analysed anyway: missing a relevant change is worse than one extra job. The PR file lookup stops after five pages because it runs before the webhook responds and GitHub times deliveries out after 10 seconds. Duplicate deliveries collapse in the job queue (one job per repository id and head SHA).
check_suite.requested is deliberately not used as a trigger: push already covers it, and using both would double jobs.
packages/github-app/src/worker/ turns each queued job into a check run:
- Mint an installation token scoped to the one repository and to
contents: read+checks: write. The App-side job client holds the token; adapter workers receive neither token nor private key, but share the Node process with the App. This is not process or OS network isolation (see security model and #326). - Claim the SHA through the reporter. A SHA that already has our run is skipped; re-runs (
check_run.rerequested) always get a fresh run. - Ask the API for the tarball and follow the redirect only to
https://codeload.github.com. The body is streamed under a 512 MiB compressed ceiling and a 5 minute timeout. - Extract only through core
extractTarball(no git, no hooks, nothing executed; see the security model). - Run
analyseRepositoryIsolatedwith the JS/TS, Rust, Go and Python adapters, each in a worker thread by default (worker/analyse-job.ts). If the checkout scan was truncated or skipped paths, the worker passes core's scan-completeness notes (scanCompleteness, plusscanIncomplete). Core adds the notes to the result and caps absence findings, so a partial checkout never gets an "unused" verdict and the app never edits the result itself (#136, #161). Full scans (pushes and unlinked re-runs) also honour a committed repo-root.ghostdeps.jsonat the analysed ref: its declared fixture roots are omitted from the scan, and the result carries core'sscanScopedisclosure - aScan scopeblock (including the analysed head SHA), the omitted-file note and the absence-reference cap - while a repo with no config is analysed exactly as before (#354). - Verdicts come from core's default recommendation policy (
createDefaultPolicy(), the same default the CLI scan uses).unusedfindings are capped at medium severity and medium confidence until the corpus check stays green (#173, #178), and never appear on a partial checkout. SetGHOSTDEPS_RECOMMENDATIONS=false(or the app optionrecommendations: false) to report facts only. - Same-SHA re-runs (#174): the check output posted for a finished analysis is kept in an in-process cache, bounded by a 16 MiB byte budget (least recently used entries go first) and 4 entries per repository. It holds the rendered output, not the analysis, because that is all a re-run posts. The key is the repository, head SHA, base SHA, source-only flag, the core/adapter contract versions and the worker's policy, adapter and scan config. A committed
.ghostdeps.jsonadds no key material of its own: it is versioned at the analysed SHA, and the key already pins that SHA, so a config change arrives with a new head and never reuses an old scoped result. A re-run with the same key posts it again and skips the download and analysis, so its output equals the first run's. Only clean analyses are kept. Anything with an adapter error or an app note (skipped step, full-repository fallback) is re-analysed, because a re-run is how users retry those. The cache is empty after a restart. The worker optionresultCache: falseturns it off. - Complete the run. Any failure ends as a
neutralrun titled "GhostDeps could not run" with a plain reason, never a crash or a silent drop. The checkout directory is always removed.
For pull request jobs (and re-runs GitHub links to a same-repo PR), the worker also reads the PR's dependency changes (#115). It fetches the base...head compare diff (the same token, contents: read only), reads each changed package.json at both SHAs as raw text, parses it statically with the JS/TS adapter, and runs core extractDependencyChanges. The result goes to core as AnalyseOptions.pullRequestChanges, so the policy scopes findings to the dependencies the PR touched, and annotations go only on lines the PR adds. If any part of that can't be read (diff too large, malformed or unreadable manifest), the worker analyses the full repository instead: scoping to a partial list could hide a finding. Fork re-runs carry no PR link and get a full analysis. Fixture scope applies to PR analyses too (#354). The head's effective scope governs both snapshots: dependency changes, changed source files and source-line changes under excluded roots are dropped from the analysis, and the check notes disclose the count of changed paths that were excluded (with up to ten examples), so a fixture-only edit never drives a verdict and is never invisible. The filter covers every extracted record - dependency changes, manifests, lockfiles and source lines - and a rename crossing the boundary is excluded on both sides, so the in-scope half of a rename cannot speak for the excluded half. A PR that edits .ghostdeps.json is disclosed with the old and new config digests, the added and removed roots, and a note that the diff interpretation is incomplete; when the base side's config cannot be read, that unknown is disclosed the same way rather than treated as “no config” - even when the head has no config either, because an unreadable base never implies “unchanged”. A committed base config with an empty root list still counts as a config in the comparison, so adding roots in the PR is disclosed as a change. A source-only PR whose changes all sit under excluded roots still gets a check with this changed-path disclosure - scoped event selection never filters it out. Only when the diff itself cannot be read does the run fall back to an UNSCOPED whole-repository analysis with the prominent fallback note - no fixture roots applied, no excluded-path disclosure, no scope-config comparison - because unknown is not an empty change set. The source-only quiet path keeps its skipped-removed-usage note and still adds the excluded-path disclosure and scope comparison when the diff was read. PR-scoped runs stamp the analysed head SHA and carry the Scan scope block like full scans; the same-SHA result cache pins head and base SHAs, so the committed config needs no extra key material, and fallback runs are never cached.
Core can add an approximate install footprint to each direct dependency's impact entry (#59 slice B), from sizes the caller supplies through a PackageMetadataProvider. The app routes through RegistryMetadataService to its only registered fetcher, NpmMetadataService in worker/npm-metadata.ts (#174). It is off by default: set GHOSTDEPS_FOOTPRINT=true (or 1, or the app option footprint: true) to turn it on. Each job gets its own provider, so the fetch budget below applies per run.
- npm only (
javascript-typescript). Sizes are the registry'sdist.unpackedSizefor each exact name@version. PyPI, crates.io and Go get no footprint. - Only packages whose lockfile evidence puts them on the public registry are queried: core's
originmust be one of exactly two origins,https://registry.npmjs.orgorhttps://registry.yarnpkg.com(yarn classic's public mirror). The allowlist lives in core. Everything else is skipped silently: no request, no footprint, no note. A private or mirrored package's name never leaves the installation. - pnpm: expect no footprint for most pnpm projects. Unscoped
pnpm-lock.yamlentries carry only an integrity hash, with no resolved URL, so there is no origin evidence and they are skipped by design (fail-closed, not a bug). Packages under a scope with a.npmrcregistry binding for the public registry do get one. - Names and versions from lockfiles are validated before they reach a URL: npm name syntax and exact semver only. Anything else stays unsized.
- Sizes, public package health facts and known misses are cached per name@version in a process-local LRU (50,000 entries, 24-hour TTL). Complete footprint answers are cached under a hash of the resolved dependency set (one-hour TTL), so an unchanged source-only PR can be answered without the registry. Transient errors and truncated answers are not cached. A restart clears these caches.
- Each run starts with a budget of 300 registry requests. A larger public-registry version set raises the budget up to 1,000; an explicitly configured budget stays fixed. Fetches run 8 at a time (16 for a set over 300), with a 3 s per-request timeout and an 8 s deadline for the provider call (core stops waiting at 10 s). A response over 1 MiB is dropped: the cap is enforced on the streamed body, so a chunked response can't be buffered past it. An answer cut short by the budget, deadline or a failed request is used but not cached as a whole, so the next run fills the gaps. A run whose footprint was cut short also never goes into the same-SHA result cache, so a re-run of that head doesn't repeat an under-reported footprint.
- Proven-public package packuments may also be queried at
GET https://registry.npmjs.org/<name>for exact-version publication time and deprecation facts, under the same validation, no-redirect, response-size and budget rules. No repository identifiers or tokens are sent to npm. Missing data leaves facts absent; the metadata is advisory and does not by itself change a check conclusion. No other ecosystem registry is queried.
The service receives signed GitHub webhook payloads, reads selected repository content through the GitHub API and a bounded codeload tarball, and writes findings to GitHub Checks. Each checkout lives in a temporary directory removed after the job. The queue, recent job dedupe state and clean same-SHA rendered check outputs are in process memory; there is no persistent scan-history database or hosted dashboard in this version. The rendered output cache is bounded to 16 MiB total and four entries per repository and clears on restart. Public npm metadata caches are described above. GitHub retains check output under its own policies; server/operator logs and backups are deployment responsibilities, so do not promise a universal deletion period for those. The App does not run repository scripts. Adapter workers have an advisory offline policy, not an OS network sandbox.
What the app costs per analysis against GitHub's rate limits, what already protects it, and the open gaps: github-api-limits.md (#37).
The worker retries a rate-limited GitHub request only when the wait is 60 seconds or less, and at most twice (#255). Otherwise the run ends neutral as "GhostDeps could not run", saying the rate limit was reached and to wait a few minutes before re-running, instead of holding a worker slot until the limit resets. The webhook never waits: if the changed-files lookup takes longer than 5 seconds, the event takes the same path as a failed lookup and the change is analysed anyway. The worker logs every GitHub response's x-ratelimit-* reading at debug level and warns once (per job client and resource) when the remaining budget drops below 10% of the limit (#256). When a pull_request.synchronize event queues a new head for a PR, a queued (not yet running) job for the head it replaced (the payload's before) is dropped, because its check would already be out of date (#257). Only that exact head is dropped, so a late or redelivered older event never drops the current head's job. A running job finishes. A re-run the user asked for is never dropped and never drops other work. Each drop is logged at info level.
Local development uses smee.io or a tunnel for webhook delivery; credentials come from a development-only GitHub App registration, never the production app. Setup steps will land here with the app skeleton (M0).