Skip to content

feat: add scoped Spec Studio audits and confirmed remediation - #260

Open
Coding-Dev-Tools wants to merge 9 commits into
mainfrom
codex/spec-crawl-readiness-20261010
Open

Coding-Dev-Tools wants to merge 9 commits into
mainfrom
codex/spec-crawl-readiness-20261010

Conversation

@Coding-Dev-Tools

@Coding-Dev-Tools Coding-Dev-Tools commented Oct 10, 2026 •

Copy link
Copy Markdown
Owner

Description

Spec Studio adds a deterministic v2 specification crawler and scoped memory-cluster audit to the dashboard, REST API, Classic MCP, and Smart MCP discovery. It reports ambiguity, injection patterns, claim support, and advisory memory remediation while leaving analysis read-only.

Parameter conflicts require positively matching nonempty subject keys and compatible claim kinds. The index groups records by subject before comparison so unrelated parameter values cannot create or hide conflicts. Supersession suggestions use distinct finite valid_from dates; ingestion order cannot select a historical backfill as the keeper. Missing, invalid, or equal effective dates and unproven divergences produce clarification without keeper/retirement targets. The UI displays those instructions and effective dates safely. Failed evidence lookups remain unchecked, do not count as unsupported claims, and do not penalize the score; successful abstentions remain untraced.

Scope and owner checks restrict memory analysis to approved live records. HTTP accepts direct text or approved procedural-memory inputs; local-operator MCP file reads stay within approved roots. Clarifications remain pending for review. Supersession requires explicit confirmation and validates both records before atomically linking and closing validity. The UI rejects stale scoped responses, renders untrusted text safely, reports actual failures, and supports accessible clarification dialogs.

The full browser gate exposed a deferred-fit camera race in the every-node renderer. Hidden zero-size measurements now preserve the last usable viewport. Worker preview, ready, and final-layout fitting respect manual keyboard, wheel, drag, pinch, and focus navigation; explicit Fit remains available. Controlled real-worker browser tests preserve exact projected camera equality across hiding and replay, including tiny drags and ordinary automatic fitting.

Evaluation enforces section counts and score bounds in full-stack and NumPy-only CI. MCP contracts, skill hashes, guides, and public charts are synchronized with immutable offline evidence. README benchmark links are pinned to the source/evidence commit for PyPI rendering. Historical artifacts and unrelated primary-checkout work remain preserved.

Type

  • Bug fix
  • New feature
  • Documentation
  • CI/CD

Verification

  • 136 focused Python checks passed, followed by the affected 30 multi-node and 73 evidence/documentation/release checks after assertion and URL corrections.
  • All eight Spec Studio Chromium checks passed, including safe rendering and no supersession for ambiguous keeper dates.
  • 290 graph asset checks and the existing 24-test graph lifecycle run passed; six targeted checks passed after the final tiny-drag and explicit-fit corrections.
  • ruff check ., Pyright 1.1.414, commercial/CSP drift checks, MCP contract export, and the Spec Crawl evaluation passed.

Current head: d87bb31d9c09265904b2c4477dda9831c7ecf5db. Fresh full CI is running at https://github.com/Coding-Dev-Tools/engraphis/actions/runs/38070240255; results from the previous head do not qualify this changed head. The three review findings are implemented and their original threads are outdated; formal thread resolution remains subject to the repository approval protocol.

@Coding-Dev-Tools
Coding-Dev-Tools marked this pull request as ready for review October 10, 2026 08:56
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-10T17:07:14.666827Z d87bb31 New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 3a7e1dd391

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread engraphis/core/spec_crawl.py Outdated
Comment on lines +1043 to +1047
if (node_a["subject_key"] and node_b["subject_key"]
and node_a["subject_key"] != node_b["subject_key"]):
continue
if (node_a["claim_kind"] and node_b["claim_kind"]
and node_a["claim_kind"] != node_b["claim_kind"]):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Require a shared subject before reporting parameter conflicts

When either record lacks a subject_key or claim_kind, these guards allow any two memories with the same parameter name and different values to become a high-severity conflict. For example, an unrelated web-server port: 8000 and database port: 5432 are presented as contradictory with a one-click supersession remedy, even though no common subject was established. Require positively matching claim identity or other subject evidence before proposing retirement.

Useful? React with 👍 / 👎.

Comment thread engraphis/core/spec_crawl.py Outdated
Comment on lines +1053 to +1054
newer = node_a if node_a["ingested_at"] >= node_b["ingested_at"] else node_b
older = node_b if node_a["ingested_at"] >= node_b["ingested_at"] else node_a

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Choose conflict keepers by valid time

For backfilled or imported facts, ingestion order does not represent world-time recency: a historical record can be ingested after the currently valid record. Selecting the keeper solely by ingested_at can therefore recommend retaining the historical value and retiring the current one; use valid_from (with a deliberate fallback) or require the user to choose when temporal ordering is ambiguous.

AGENTS.md reference: AGENTS.md:L175-L177

Useful? React with 👍 / 👎.

Comment thread engraphis/service.py Outdated
Comment on lines +12957 to +12959
except Exception:
pass
return None

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve unchecked status when claim lookup fails

If grounded_recall raises because the vector index or store is temporarily unavailable, this wrapper swallows the exception and returns None; crawl_spec interprets that as a completed lookup with no evidence, marks the claim untraced, and penalizes the score. Let the exception reach crawl_spec (which already converts lookup failures to unchecked) or return a distinct failure result so infrastructure errors are not reported as unsupported claims.

Useful? React with 👍 / 👎.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 99189b0990

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread engraphis/service.py
Comment on lines +13129 to +13131
if act == "supersede":
self.retire(node_b_clean, workspace=ws or "default", repo=route["repo"],
reason="spec_crawl_resolve", actor="spec_crawl")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Close the retired fact at the replacement's effective time

For a backfilled conflict where the retained node's valid_from predates the remediation, this calls ordinary retire(), which closes the old node at the current wall-clock time rather than at the replacement's effective time. Although the remedy now selects the correct keeper by valid_from, an as_of read between that effective time and the remediation time still returns both contradictory facts; close validity at the retained node's valid_from instead.

AGENTS.md reference: AGENTS.md:L175-L177

Useful? React with 👍 / 👎.

Comment thread engraphis/service.py
Comment on lines +13114 to +13115
for mid in (node_a_clean, node_b_clean):
self._check_owns(mid, wid, rid)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Permit ancestor nodes in repo-scoped remediation

When a repo-scoped audit includes a workspace/user ancestor, as its SearchFilter(include_ancestors=True) intentionally does, _check_owns(..., rid) rejects that ancestor because its repo_id is None. Consequently every generated supersede or link remediation involving a repo node and an ancestor fails with “does not belong to that workspace/repo”; validate both nodes against the same ancestor-aware audit scope and avoid reapplying exact-repo ownership during the mutation.

AGENTS.md reference: AGENTS.md:L178-L179

Useful? React with 👍 / 👎.

Comment thread engraphis/service.py
Comment on lines +13073 to +13077
for mid in mem_ids_set:
node_links = self.store.get_links(mid)
for nl in node_links:
if nl.get("a") in mem_ids_set and nl.get("b") in mem_ids_set:
links.append(nl)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Deduplicate stored links before cluster analysis

For every stored A–B link, both get_links(A) and get_links(B) return the same row, so this loop appends every edge twice. Besides duplicating the API's links and replay events, these duplicates inflate resolved_links, whose length is used to decide whether semantic-link inference runs; for example, two stored edges among four nodes appear as four links and suppress inference entirely. Deduplicate by endpoints and relation before calling analyze_memory_nodes.

Useful? React with 👍 / 👎.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d87bb31d9c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +320 to +323
elif _XML_CLOSE.match(line):
title = "" # closing tag: consumes the line, starts no section
if title is not None:
heads.append((offset, min(line_end + 1, len(text)), title))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve text after XML closing tags

When a spec uses XML-style sections, adding the closing tag as an empty heading ends the preceding section but creates no section for subsequent unheaded text. For example, content after </role> is never tokenized, so trailing rules, vague wording, or injection payloads are omitted and the report can incorrectly appear clean. Preserve that trailing span as a section instead of silently dropping it.

Useful? React with 👍 / 👎.

Comment on lines +765 to +766
if target not in _ENFORCE_VERBS | _FORBID_VERBS | _STOPWORDS:
directives[target] = action

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep outer negation from being overwritten

For directives such as Never require authentication or Never enforce deployment, the outer never first records a forbid directive, but the scan later processes require/enforce independently and overwrites the same target with enforce. As a result, comparison with Always require authentication reports no conflict, hiding a direct policy contradiction; preserve the governing negation when consuming nested directive verbs.

Useful? React with 👍 / 👎.

Comment thread engraphis/mcp_server.py
Comment on lines +2794 to +2796
"title": "Crawl spec or prompt quality",
"readOnlyHint": True,
"destructiveHint": False,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Register crawl-only MCP tools as viewer reads

In hosted/team MCP usage, minimum_role() does not derive access from this annotation and defaults unlisted tools to member. Because neither engraphis_spec_crawl nor engraphis_spec_crawl_memories was added to _READ_ONLY_TOOLS, viewers are denied these explicitly read-only audit operations even though they can use the existing viewer-level memory reads; register both tools in the read-only role set.

Useful? React with 👍 / 👎.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant