All future implementation work belongs in:
/Users/yjw/agent/Agent Engineering Toolkit
This repository was migrated here on 2026-07-11 with its Git history, phase tags, local fixtures, and dogfood snapshots intact. Treat this path as the only active workspace; the prior Codex-generated directory is archival.
At the start of any new task, read this file, then run git status --short,
git tag --sort=-creatordate | head, and
PYTHONPATH=src uv run --no-editable python -m unittest discover -s tests.
The explicit source path avoids loading a stale non-editable environment copy.
Do not start a
later phase until the current phase's acceptance checks pass and this file is
updated.
aet is a local Evidence Plane for coding Agents. It binds human intent, Git
state, command execution, artifacts, normalized Agent runs, review context,
and limitations into portable, hash-bound evidence. Quick is read-only by
default; only an explicit /aet-proof argv executes project code and writes
the requested compact Receipt. AET reports unknowns and never invents a
holistic trust score or promotes PROPOSED advice into authority.
Primary users are individual developers and small teams using Codex, Claude Code, Cursor, Copilot, or compatible agent Skills across multiple repositories.
The default product surface is four dedicated cross-agent Skills:
aet-check, aet-scope, aet-proof, and aet-fresh; the aet CLI is their
deterministic local runtime. The canonical agent-engineering-toolkit Skill is
the compatibility and explicit AET Lab route. Every shipped capability retains
machine-readable evidence contracts so any Agent Host can invoke it.
aet-plan is a separate fifth Host Skill and not part of automatic Quick
chaining. It compiles existing evidence and current source into a bounded,
read-only PROPOSED Plan for a Host Planner; it grants no implementation or
verification authority.
Out of scope: an agent runtime, a Skill marketplace, automatic prompt rewrites,
and model-dependent prompt regression. The only retained later expansion is
Repo Archaeologist, planned as aet evolve, not a dependency of the static
core.
- Local, deterministic, and read-only by default; no LLM or API key in v0.1.
- Every finding has a stable ID, status, severity, evidence location, remediation, and rule version.
- Use
PASS,FAIL,UNKNOWN, andNOT_APPLICABLE; do not produce a single opaque score. - Keep the core dependency-light. Python 3.11+ standard library is sufficient for v0.1.
- New rules require a fixture that proves both a clean and failing case.
- End every phase by: testing, committing, tagging, updating this file, and bumping/installing the local Skill version.
discovery.py locates context assets → rules.py produces evidence-backed
findings → reporters.py serializes Markdown, JSON, or SARIF → cli.py
controls output and CI exit status.
Supported v0.1 assets: AGENTS.md, CLAUDE.md, CODEX.md,
copilot-instructions.md, .cursorrules, and every SKILL.md below the
target root.
| Stage | Version | Git tag | Result | Rollback |
|---|---|---|---|---|
| Phase 0 | Skill 0.0.1 | phase-0-dogfood |
Workspace, fixtures, dogfood baseline, project memory, and conservative path semantics. | git checkout phase-0-dogfood |
| v0.1 | Skill 0.1.0 / package 0.1.0 | v0.1.0 |
Static context and Skill audit CLI with Markdown, JSON, SARIF, CI example, tests, and wheel verification. | git checkout v0.1.0 |
| v0.2 | Skill 0.2.0 / package 0.2.0 | v0.2.0 |
Intent Gate: human-reviewable contract, changed-path budget, scope checks, and proof-evidence checks. | git checkout v0.2.0 |
| Skill portability | Skill 0.2.1 / package 0.2.0 | skill-v0.2.1 |
Tool-neutral SKILL.md cleanup and cross-agent contract. |
git checkout skill-v0.2.1 |
| v0.3 | Skill/package 0.3.0 | v0.3.0 |
Host-neutral Evidence Pack compiler and opt-in, redacted command Trace. | git checkout v0.3.0 |
v1.19.1 patch release candidate. The deterministic core includes Quick, Intent Gate,
Run Normalization, Portable Investigation and Evidence Bundle, eleven-view
Evidence Atlas, Improvement, Evidence-Guided Planner, graph-first Review Graph,
diagnosis-only Behavioural Risk, Learn, and Repo Archaeologist. Review Graph
defaults to a bounded root slice; Risk forecast remains hard-disabled as
research-only UNKNOWN. GitHub Release is authorized for this task; PyPI
publication is out of scope.
- Initialized this repository and the
agent-engineering-toolkitSkill. - Added clean and broken fixtures plus an initial evidence schema and static rule prototype for dogfooding.
- Audited read-only shallow clones stored outside Git tracking:
stock-analysisate875974992e8a5258df9723ead115390efecf5a1pain-mineratddf3ce20a1bc0cfbd04b79ac76c8e713b6ff0fdacli-creator-skillat418e941607a53be95921edbbb8e2196411b7893d
- Final dogfood reports are in
docs/dogfood/; each discovered one Skill and emitted zero FAIL/UNKNOWN findings under the v0.1 prototype. - Dogfood corrected two false-positive risks before release: generated output names are not local paths, and verification detection recognises Chinese as well as English wording.
- Local Skill installed at
~/.codex/skills/agent-engineering-toolkit, version0.0.1.
- Released the read-only
aet audit [path]command with Markdown, JSON, and SARIF output plus CI exit semantics (FAILalways fails;--strictalso fails onWARN). - Implemented deterministic checks for missing local Markdown/explicit command targets, root instruction bloat, duplicate long directives, required Skill frontmatter, Skill name/directory mismatch, and verification instructions.
- Added four standard-library tests covering clean and failing fixtures, CLI exit codes, and SARIF parsing. Use the exact resume command above.
- Added a GitHub Actions SARIF example and built a source distribution and wheel. Release verification includes running the wheel from a fresh temporary virtual environment.
- Upgraded and reinstalled the local Skill at
~/.codex/skills/agent-engineering-toolkit, version0.1.0. - On this Python 3.13 + uv environment, editable installs create an
underscore-prefixed
.pthfile that CPython skips. Useuv run --no-editablefor local verification; released wheels are unaffected.
- Added
aet review --base <revision>, which reads a human-authoredaet.intent.json, compares the working tree (including untracked files) to the Git base, and emits evidence-backedAET-REV-001throughAET-REV-004findings for contract validity, changed-path budget, scope, and local proof evidence. - Intent Gate stays read-only: it never executes declared proof commands. A PASS for proof evidence means the command and local evidence were declared; human review must run the command separately before claiming it passed.
- Added passing and failing Git-backed review tests, then verified the 0.2.0 wheel in a fresh virtual environment. The local Skill is updated to 0.2.0 with the v0.2 contract reference.
- Removed the generated-template residue from the canonical Skill and made its
workflow tool-neutral: the shared boundary is
aetplus JSON/SARIF output, not any vendor-specific API. - Added
references/cross-agent-use.md. Native Skill hosts install the whole folder; agents without native Skill support can loadSKILL.mdas project instructions and invoke the same CLI. agents/openai.yamlremains optional UI metadata. It must never become a runtime dependency or reduce compatibility for another agent host.- Upgraded and reinstalled the local Skill to
0.2.1after structural validation and the current test suite passed.
- Added
aet trace --output ... -- <command> [args...]: the sole opt-in execution path. It records redacted argv, exit status, timestamps, working directory, Git HEAD/worktree digest, and SHA-256 digests of redacted stdout/stderr artifacts. A non-zero command produces a validFAILTrace; it is not successful proof. - Added
aet evidence pack, which schema-validates independently produced audit, review, and trace JSON, records source SHA-256 values, atomically writes a portable JSON pack, and marks missing optional componentsUNKNOWN. It preserves component summaries and excludes raw command logs. - Added acceptance coverage for successful and failing commands, built-in secret redaction, stable hashes, invalid schemas, missing inputs, atomic replacement, and an audit → review → trace → pack temporary Git fixture.
- Ran 11 unit tests through Trace, built 0.3.0 source/wheel distributions, and installed the wheel into a fresh virtual environment for a Trace smoke test.
- Updated the canonical and local cross-agent Skill to 0.3.0 with the portable
v0.3-contract.mdreference.
Turn the existing audit and review reports plus explicitly requested command execution into one portable, content-addressed Evidence Pack that any agent can attach to a handoff or CI run.
# Explicit execution only; `--` separates the trace options from the command.
aet trace --output .aet/evidence/trace.json -- <command> [args...]
# Compile independently generated audit, review, and trace artifacts.
aet evidence pack \
--audit .aet/evidence/audit.json \
--review .aet/evidence/review.json \
--trace .aet/evidence/trace.json \
--output .aet/evidence/evidence-pack.jsontraceis opt-in and executes only the explicit argv after--; neither audit nor review may start executing commands implicitly.- Store argv, exit code, start/finish timestamps, working directory, Git HEAD and diff digest, plus SHA-256 digests of captured stdout/stderr artifacts. Do not write raw output to the Evidence Pack by default.
- Redact configured secret patterns from command metadata and persisted log
excerpts. If redaction confidence is insufficient, mark the field
UNKNOWNrather than retaining the value. evidence packvalidates the input report schemas, records each input's SHA-256, preservesPASS/FAIL/UNKNOWNwithout collapsing them to a score, and writes atomically.- Inputs may be absent, but the pack must record the missing component as
UNKNOWNand never imply that a test or review happened. - Keep the format host-neutral JSON. It must be consumable by any agent that can read files; no MCP, model API, or vendor trace API is permitted.
- Unit tests cover a successful command, a non-zero command, secret redaction, stable hashing, invalid input schema, missing optional inputs, and atomic output replacement.
- A clean temporary Git fixture produces audit → review → trace → pack with source hashes and no fabricated status.
- A failing command has a recorded non-zero status and a valid Trace artifact; it does not become a successful proof.
- Build and install the 0.3.0 wheel in a fresh virtual environment.
- Upgrade the canonical and local cross-agent Skill to 0.3.0, update this
memory with actual results, commit, and tag
v0.3.0.
The static core will be complete. Repo Archaeologist remains aet evolve and
must not become a dependency of audit, review, trace, or Evidence Pack. No
model-generated judgement should be the sole release gate.
The complete post-v0.3 product plan is recorded in
docs/productization-plan.md. It was produced from the original “设计 Agent
工具包方案” conversation, the active repository at v0.3.0, and a source
review of yaojingang/yao-meta-skill at commit
4eb11f923dc71173736ebf541a7eebfff942d10e.
aet is an Evidence Plane, not a general-purpose Skill OS. Its stable product
surfaces are Context/Skill Hygiene (audit), Intent Change Control (review),
Execution Evidence (trace/evidence pack), and Repository Evolution
(aet evolve). Repo Archaeologist is therefore a first-class usage scenario
of the canonical cross-agent Skill, but remains independent of the offline
deterministic core.
- Reuse an Evidence IR with source hashes and verification levels; preserve
PASS/FAIL/UNKNOWNrather than creating a health score. - Scores may only prioritize reviewer work; they cannot release a change or convert an unknown into a pass.
evolvemust distinguish direct, corroborated, candidate, and unknown links across Git, docs, releases, PRs, and Issues. Model narration is an optional, provenance-bound inference and cannot become source evidence.- Implement
v0.3.1first: fix the stale README v0.3 claim, add auditable discovery excludes/config, and make self-audit usable despite intentional failing fixtures. Then execute v0.4 Evidence IR/proof binding, v0.5 Skill UX/governance, v0.6 offline evolve, and v0.7 GitHub evolve.
aet audit . --strictproduces expected FAIL/UNKNOWN results fromtests/fixtures/broken_project, proving the current discovery layer lacks a configurable test-fixture boundary.README.mdstill contains an obsolete statement that v0.3 Trace and Evidence Pack are planned/not implemented, althoughv0.3.0implements them. This is a documentation defect.- This entry is a design/memory update only; no v0.3 behavior was changed and no new release tag has been created. Resume implementation from the v0.3.1 acceptance criteria in the productization plan.
- v0.1 parses local Markdown and paths; it cannot prove that a remote MCP server is reachable or that a command semantically succeeds.
- Detection is intentionally conservative. A missing local target is a FAIL; a remote, dynamic, or bare filename is left unverified rather than guessed.
- Repo Archaeologist needs GitHub history, Issues, PRs, and releases, so it is deferred behind the stable evidence schema.
- Implemented the Evidence Plane: configurable
audit, intentreview, proof-boundtrace/ Evidence Pack / static viewer, transparent non-gatingtriage, andaet evolveas the Repo Archaeologist surface (plan, local Git/docs, export or explicit GitHub API, graph, report, query). - Added Evidence IR metadata and L0–L5 boundaries while preserving
PASS/FAIL/UNKNOWNas authoritative. Weighted triage factors are visible and cannot alter a gate. - Fixed self-audit with reasoned
aet.tomlexclusion for intentionally broken fixtures. Added stale absolute-path detection after auditing Codex globalAGENTS.md, whose Skill index can drift from installed paths. - Scoped Codex
AGENTS.mddogfood found 52 stale absolute Skill paths and one root-context-bloat warning; the audit did not mutate global instructions. - Added v1 README, contracts, schema, canonical Skill flow, changelog, CI and tag-driven GitHub release workflow.
- Before tagging, require strict self-audit, full unit suite, wheel plus isolated CLI smoke, a proof-bound Evidence Pack, GitHub audit/evolution evidence, and a clean intentional commit.
- Reworked the English README from a command-first reference into a product entrypoint: user problem, four capability surfaces, evidence architecture, verified quality boundary, install paths, workflow guides, Repo Archaeologist, audience fit, and repository map.
- Added
docs/README.zh-CN.mdas the complete Simplified Chinese companion, with an explicit language switch at the top of both README files. - Added
CONTRIBUTING.md, a copyable generic intent example, and GitHub Issue forms so external users can report sanitized evidence-boundary defects or propose concrete workflows without exposing private repository content. - GitHub discovery metadata should describe AET as evidence-first engineering guardrails for coding agents and use focused topics rather than broad AI hype. Keep public claims tied to reproducible release checks; do not claim PyPI publication unless it actually occurs.
- Added
aet context discover,record, andverify. The Context Manifest records local instruction/Skill discovery and hashes; a declared read is explicitly anagent_attestation, never evidence that a model understood or used the file. - Added the local JSON Decision Ledger with
init,add,list,verify, andsupersede. It stores source hashes, evidence state, lifecycle state, and replacement history; it is project decision provenance, not generic Agent memory or RAG. - Regression coverage now includes both a clean and a changed-source/context path, as well as direct supersession. The v1.3 release gate requires 27 unit tests, strict self-audit, reviewed intent, a proof-bound Trace/Evidence Pack, and an isolated wheel smoke test.
- Invest-Vault dogfood confirmed that v1.3 Trace successfully ran its complete pytest process; the gap was report portability, not subprocess execution. Trace held stdout/stderr but not a test framework's generated report.
- Added explicit
trace --artifact <relative-path>capture. It is not a report-file guesser: only a declared regular UTF-8 file under the workspace is captured, redacted, hashed, and embedded into Trace plus Evidence Pack. - A missing, outside-root, non-regular, undecodable, or unredactable declared
artifact is
UNKNOWN. AET returns non-zero after an otherwise successful child command, while preserving the child'sexecution: PASSfact. - A real pytest dogfood trace initially exposed unrelated collection of nested
work/dogfoodrepositories.pyproject.tomlnow limits optional pytest discovery to AET's owntests/directory. - This adopts the useful Harness Engineering idea of durable, inspectable filesystem artifacts and failure traces. It explicitly rejects the article's broader runtime, autonomous optimization, and generic-memory directions.
- The v1.4 release gate requires 30 unit tests, a real pytest JUnit artifact dogfood trace/pack, strict self-audit, reviewed intent, a proof-bound release Evidence Pack, and an isolated wheel smoke test.
- Added three commit-locked, static-only cases: SWE-agent
3ea751c087f32b16e039a2233dd6eefecef325d5, Google ADK67ab27f2547db48f7248b1689aab4c18502aee17, and OpenHands96f902a9ac14bf5edfb2e47d759d75c91e4faf28. - Added the independent
repository-audit-profile/v1contract andaet audit swe-agent|google-adk|openhands --repo <checkout>. Existingaet audit <path>behavior andaudit-profile/v1remain unchanged. - Each case writes two shared machine artifacts and five human-readable
artifacts under both
en/andzh-CN/. Findings are deterministic, evidence-located engineering observations; no holistic score, upstream code execution, upstream test execution, source redistribution, or LLM-authored Finding is permitted. - The measured runtime includes evidence collection, rule analysis, complete report rendering, and staged artifact writes. Clone, dependency installation, LLM network time, and manual review remain outside the 900-second contract.
- OpenHands
enterprise/**andtests/**/enterprise/**are prohibited. The latest upstream has moved its Agent core into separately versioned dependencies, so the local Agent-core claim remainsUNKNOWN. - The maintainer approved all three bilingual report snapshots after
English/Chinese parity review and desktop/mobile visual inspection; their
tracked
review.statusisAPPROVED. New runs still default toPENDING. - Acceptance evidence: 12 focused repository-audit tests passed; the full
regression gate passed 216 tests; the complete business-quality gate passed
75 tests and 237 subtests; the wheel built and returned version
1.12.0in an isolated environment. All three Chinese HTML reports passed a 390 × 844 viewport check with no overflow or broken images. - Remaining release work is operational: bind the final Diff in
release-classification.json, commit, tag, push, wait for exact-commit CI, and publish GitHub Releasev1.12.0. Do not publish to PyPI.
- Reframed the default product surface around four independent, bounded Skills:
/aet-check,/aet-scope,/aet-proof, and/aet-fresh. Each emits one result and stops. Existing 1.x CLI and AET Lab surfaces remain compatible and require explicit opt-in. - Added
aet quick check|scope|proof|fresh. The new layer reuses the existing deterministic audit, Git, Trace, snapshot, redaction, and receipt core rather than changing legacy command semantics. - Added host-neutral investigation contracts and standard-library validation for Intent provenance, competing hypotheses, immutable result references, counter-explanation requirements, Finding strength, tool authority, write and execution permission, command budgets, and stop conditions.
- Preserved
PASS/FAIL/UNKNOWN/NOT_APPLICABLEas authoritative evidence status. Finding origin and semantic support remain in the Investigated Finding contract; Scope disposition and Freshness state remain command-level fields and cannot overwrite source evidence. - Quick Proof writes one compact JSON receipt after an explicit request and records argv, exit status, workspace snapshot, relevant paths, artifacts, the selected executable identity, explicitly named environment-input hashes, Python/platform identity, and dependency lockfile hashes. Quick Fresh distinguishes exact, unrelated-workspace, HEAD-only, relevant-file, artifact, environment, and unknown drift while retaining legacy evidence fallback.
- Added deterministic narrative routing: only a Chinese slash-command request defaults to Simplified Chinese; all other requests use English. Rendering changes no machine state or evidence reference.
- Rebuilt the English and Simplified Chinese README around Quick, added four dedicated portable Skills, six JSON Schemas, synchronized static and animated SVG architecture sources, 1600 × 900 PNG renders, and English and Chinese WebM introductions. The three commit-locked Repository Audit Showcase cases remain unchanged and are documented as AET Lab.
- Completed the full Investigation Contract runtime: contract shape,
finding_type, explicit-user source references, negative-search coverage, material recorded conflicts, semantic disclosure, and allowed stop reasons are now checked by the standard-library Grounding Validator. The Validator always validates the supplied Ledger before trusting its references. - Added the opt-in
eval/quick-investigation/AET Lab harness with the four frozen comparison groups and eight Scope scenarios. A realgpt-5.6-sol/ medium, two-repetition run produced 64 observations. Effective recall / false discovery proportions were 60% / 50% for pure rules, 80% / 38.5% for one-shot LLM, and 90% / 25% for both investigated groups. The Grounded group used the shipped Validator and rejected zero claims in this sample. The tracked result records time, tools, and Tokens; unmeasured manual-review time and user understanding remainUNKNOWN. - Acceptance evidence in the working tree: 260 unit tests passed; strict
self-audit returned zero findings; an isolated wheel contained the Quick,
investigation, narrative, built-in RulePack, and new Schema assets, exposed
all four Quick subcommands, reported version
1.13.0, and passed a legacy Audit smoke test. The runnable stale-proof demo producedEXACT_MATCHand thenRELEVANT_FILES_CHANGED. The tracked 30-sample performance report measured Check P95 0.622 s, Scope P95 0.059 s, and Fresh P95 0.037 s on the recorded local environment. These are local bounded checks, not cross-repository or model-service P95 claims. - Released
v1.13.0from commit2683479cc742775674be75483fdb1606b62b3e60on 2026-07-25. Exact-commit CI Run30113760362and GitHub Release Run30113835956passed; the Release publishes the Wheel, sdist, CI manifest, Diff-bound classification, verified classification, and release-evidence record athttps://github.com/AdvancingTitans/agent-engineering-toolkit/releases/tag/v1.13.0. Real Host Gate isNOT_APPLICABLEbecause this deterministic release adopts no governance asset and makes no release-authorizing Agent behavior claim.
- Replaced the README architecture media with bilingual, motion-validated AET Quick workflow GIFs and bilingual static project panoramas. The panorama distinguishes the Quick request path, shared protocol support, human authority, and the explicit AET Lab entrance; every arrow terminates at a named component or authority boundary.
- Replaced the former WebM introductions with exact 30-second English and Simplified Chinese H.264/AAC MP4 videos. The tracked media manifest binds all published GIF, SVG, PNG, and MP4 assets by SHA-256.
- Reworked both READMEs around daily Agent coding problems, bounded product promises, measurable evaluation trade-offs, the distinct roles of grounding-aware investigation and the in-project Grounding Validator, and Chinese explanations for internal evidence terminology. Chinese documentation and generated media consistently use “契约”.
- Added Codex and Claude Code Run Normalizers, stable source identity, incremental ingestion, tool-call/result linking, generation boundaries, and fail-visible diagnostics. Run Records establish only what the normalized run contains.
- Added strict Observation, Evidence Candidate, Verified Evidence, and Portable Claim boundaries. Every exported Observation declares what it proves and does not prove; model reasoning and recorded tool output cannot become reproduced evidence without deterministic verification.
- Added deterministic Candidate verification for explicitly authorized AET Proof receipts. Command, workspace, budget, path, receipt integrity, and Freshness bindings are enforced; stale results remain historical and cannot become current proof. Added the bounded OptimizationCandidate entry contract with independent-task/high-severity evidence prerequisites and mandatory isolated evaluation.
- Added the read-only Portable Investigator, immutable investigation ledger, JSON Schemas, Bundle Compiler, strict loader/validator, canonical hashing, redaction, Index/Core/Archive layout, content-addressed Blobs, deterministic Markdown, Review Result validation, bounded MCP server, and optional Python and TypeScript SDKs. A reviewer can consume the Bundle without installing AET or either SDK.
- Real prompt-only consumption checks used ten deterministic synthetic Bundles.
Codex CLI 0.144.1 /
gpt-5.6-sol, Hermes Agent 0.17.0 /kimi-k2.6, and Ollama 0.32.3 /qwen3:8beach produced strict, independently rescorable JSON for all ten scenarios: 62PASS, 38NOT_APPLICABLE, zeroFAIL, and zeroUNKNOWN. Runtime, model, elapsed time, command-argv digest, response, report, and publication integrity are tracked. These results are a bounded interoperability check, not a general accuracy claim or trust score. - The English and Simplified Chinese READMEs now describe the complete Quick and portable handoff surfaces. Their static SVG/PNG panoramas, 115-frame motion-validated GIFs, exact 30-second H.264/AAC videos, and media Manifest reflect v1.14.0.
- Acceptance evidence: 418 Python unit tests passed after a non-editable
reinstall; the TypeScript Bundle SDK passed 11 adversarial tests, Node 20
build/compatibility checks, and package dry-runs; the Python wheel and sdist
built successfully and the isolated wheel validated a Bundle plus the
packaged Optimization Schema; all Bundle result and media hashes were
independently recomputed; strict Hermes/Ollama responses were independently
reparsed and rescored; README links, forbidden-name scan, and
git diff --checkpassed. - Released
v1.14.0from commitfee324fe1ee5681035e15b146d3fab8ccaee7f12. Exact-tag CI Run30171294767passed, the Diff-bound classification verified all 58 behavior-sensitive paths, and the public GitHub Release is available athttps://github.com/AdvancingTitans/agent-engineering-toolkit/releases/tag/v1.14.0. Its five assets include the exact CI wheel and sdist, CI Manifest, release classification, and classification verification report. No governance asset was adopted and the Real Host Gate isNOT_APPLICABLE. - Resume point: begin the next change from
v1.14.0; preserve the Portable Evidence Bundle v1 compatibility and do not reinterpret the bounded Hermes/Ollama/Codex fixture results as a general accuracy claim.
- Added Evidence Atlas as a deterministic derived layer over Portable Evidence Bundle v1: a canonical source-backed Graph, eight fixed Perspectives, bounded typed recursive decomposition, Mermaid/Markdown/JSON projections, strict provenance and Schema validation, incremental rebuilds, Atlas Diff, and an offline recursive Viewer. Graph records remain authoritative; Mermaid, documents, and Viewer state create no evidence or authority.
- Added
aet atlas build|validate|view|export|query|explain|diff, including comma-separated--perspectivesselection and no-LLM operation. The Python and TypeScript SDKs expose graph build/load/query/trace/validate/render surfaces, and MCP exposes eight bounded read-onlyaet_graph_*tools. - Freshness timelines use the recorded
checked_atvalue, normalized only to Mermaid-safe punctuation, and showUNKNOWNwhen no timestamp is recorded. Current, stale, conflict, counter-evidence, limitation,does_not_prove, missing Change Group, cycle, deduplication, and maximum-depth semantics have direct regression coverage. Malformed Mermaid declarations fail closed. - The English and Simplified Chinese READMEs now position Evidence Atlas, present bilingual static architecture, a real six-state recursive Viewer GIF, exact 30-second H.264 walkthroughs, the end-to-end flow, and an exact tracked Mermaid Claim Chain generated from AET's source-bound self-review. The example narrowly states what the Portable Evidence v1 Evidence record Schema establishes and retains its unresolved Change Group boundary.
- Acceptance evidence in the release-candidate working tree: 444 Python unit
tests passed after a non-editable reinstall; the TypeScript SDK passed 16
tests plus its built-distribution smoke test; the Viewer runtime check
passed; Mermaid 11.16.0 parsed all 93 recursive diagrams from both the source
and isolated-wheel Atlas; seven media files matched their recorded byte
lengths and SHA-256 values;
git diff --checkpassed. Fresh Wheel and sdist were rebuilt, and the isolated Wheel reported1.15.0, contained the vendored Mermaid runtime and all ten Atlas Schemas, rejectedflowchart BOGUS, built and validated the real self-review Atlas, and kept explicitdoes_not_provedocumentation. - The first exact-tag CI attempt, Run
30202354466, passed the full Python, stale-proof, and real-agent gates, then failed before parsing becausenpm --prefixresolved the relative Atlas argument from the package directory. No Release was created. CI now passes the explicit$GITHUB_WORKSPACE/.aet/evidence/atlas-self-review.atlaspath, and the delivery-gate test freezes that binding. - The final independent compliance audit approved publication with no remaining
P0/P1 or acceptance gap. Released
v1.15.0from commit039cc5a10f4ee2a9c9056060f48af671e060f5c9; exact-tag CI Run30202586689and GitHub Release Run30202642550passed. The public Release publishes the exact CI Wheel, sdist, CI Manifest, Diff-bound classification, classification verification, and release-evidence record athttps://github.com/AdvancingTitans/agent-engineering-toolkit/releases/tag/v1.15.0. Real Host Gate isNOT_APPLICABLE, and no PyPI publication was performed. - Resume point: begin the next change from
v1.15.0; preserve Portable Evidence Bundle v1 compatibility, the deterministic Graph authority boundary, explicit counter-evidence/UNKNOWN/Freshness semantics, and the seven pre-existing untracked files outside version control.
- Added the planned deterministic
src/aet/improvement/subsystem with Issue, Constraint, Candidate, Verification Contract, and Outcome models; Finding normalization/aggregation; bounded rules; human/Agent/PR renderers; Candidate grounding, Scope, reference, strength, and anti-gaming validators; and the Proof-bound verification lifecycle. - Added
aet improvement doctor,aet improve <bundle>, and theprompt,validate,verify, andcompareImprovement actions. Candidate state remainsPROPOSED;verified_improvementrequires a recorded code change, current contract-bound passing Proof, and no comparison regression. - Portable Evidence Bundle v1 has no independent Finding or Improvement
collection. The deterministic adapter consumes validated portable Claims as
Finding-compatible inputs without changing Bundle, Evidence, or Finding
schemas. Missing independent Improvement records remain explicit
UNKNOWN. - Added the non-blocking, no-LLM PR Improvement Summary workflow and the
improvement-chain/regression-lineageAtlas Perspectives. Python Atlas schemas, MCP description, and the TypeScript SDK now agree on ten fixed Perspectives. - Added a reproducible empty-tool-result review case. Its failing regression
evidence grounds
IMP-001, a human report, and a boundedPROPOSEDAgent task. The same Bundle independently drives Claim Chain and Improvement Chain Atlas views; the latter remainsUNKNOWNbecause Bundle v1 has no independent Improvement records. - Updated the English and Simplified Chinese READMEs, static architecture, animated workflow, project panorama, Atlas architecture, and silent 30-second H.264 product/Atlas videos. Added bilingual case SVG/PNG/GIF media and SHA-256 manifests. Geometry, composition, semantic motion, frame, hash, and manual visual checks passed.
- Added six planned Golden Fixture families and
docs/improvement.md. Final acceptance: unittest discovered 478 passing tests; the release-gate pytest run passed 488 tests and 636 subtests; TypeScript SDK 16/16, distribution smoke 1/1, compatibility guard, Viewer/Mermaid, stale-proof, four suites, strict audit, build, isolated Wheel smoke, andgit diff --checkpassed. - Remaining unmeasured product targets are the SC-001 human comprehension
percentage and real Codex/Claude
pass@1/pass^3execution metrics. CI publication behavior has local contract coverage but has not been observed on a live pull request. - Released
v1.16.0from commitbfe062a9f68b805a5b629f32828510a411c9a1f9. Exact-tag CI Run30420607956and GitHub Release Run30420768704passed. Release ID361498297publishes the exact CI Wheel, sdist, manifest, Diff-bound classification, classification verification, and release-evidence record athttps://github.com/AdvancingTitans/agent-engineering-toolkit/releases/tag/v1.16.0. Real Host Gate isNOT_APPLICABLE; no PyPI publication was performed. - Resume point: begin the next change from
v1.16.0; preserve Bundle v1 compatibility, Graph authority, prompt/Atlas sibling projection, explicit counter-evidence/UNKNOWNsemantics, and the seven pre-existing untracked files outside version control. Real Agent and user-comprehension metrics remain unmeasured.
- Implemented the delivery package's Phase 0–6 in order. The new deterministic
src/aet/planning/subsystem normalizes Requests, consumes validated Bundle and Atlas projections, builds bounded Planning Context, validates strict Host-produced Plan Candidates, writes integrity-bound portable Plan packages, exports read-only per-Plan Skills, and maps external unified diffs to pending verification handoffs. - Added six strict Planning Schemas; the
aet plancontext, candidate, package-Helper, Skill-export, and handoff commands; eight read-only MCP tools; Python and TypeScript consumer APIs; theaet-planHost Skill; E2E examples; and a 20-contract frozen localization benchmark. The Planner does not execute commands, edit source, promote evidence, or grant merge/release authority. Plan authority remainsPROPOSED, and verification remainsUNKNOWN/PENDINGuntil separate Proof execution. - The frozen Planner protocol benchmark passes all P1 localization thresholds:
required-path recall
1.0, recommended-path precision0.9706, required-test recall1.0, critical-linkage omission0.0, ungrounded-path rate0.0294, protected-path rate0.0, and full conflict, unknown, and reference preservation. - Added a real Codex
gpt-5.6-solAET self-review, with Gold, raw structured observations, and deterministic scoring kept separate. One run per group compared source-only, v1.16 Bundle + Atlas, and the v1.17 validated Plan. The Plan group reached1.0on all eight separately reported metrics. Against source-only, production decision precision improved by 55.56 percentage points, reference coverage by 100 points, linkage coverage by 66.67 points, andUNKNOWNpreservation by 100 points. Required path and test recall did not improve because source-only was already1.0. Against v1.16 evidence-only, disposition improved by 50 points, test recall by 100 points, andUNKNOWNpreservation by 50 points. - The first unconstrained source-only attempt traversed vendored/minified assets and was rejected as a harness failure. The recorded rerun applies the same source-root and generated/vendor exclusions as the other groups. These observations are a one-case, pass@1-style localization measurement, not a general reliability claim or holistic score.
- Rebuilt the bilingual README product story, Planner workflow GIF/SVG/PNG, real-case screenshots, project panoramas, and exact 30-second silent H.264 product videos. Motion, geometry, composition, frame, SHA-256, and manual visual checks passed.
- Acceptance evidence in the implementation working tree: 536 unittest tests
passed after a non-editable reinstall; Planner line coverage is
90%and validator branch coverage is96%; TypeScript SDK build/test/check passed 19 tests; strict repository Skill audit passed; Host Skill validation passed; benchmark and three E2E scenarios reproduced;git diff --checkpassed. Fresh Wheel and sdist built successfully, the isolated Wheel reported1.17.0, exposed the Planning Python APIs/CLI, and contained all six Planning Schemas. - Released
v1.17.0from commit19efd8426572de8f75164a78c917a51a245976de. Exact-tag CI Run30435701149and GitHub Release Run30435886299passed. Release ID361603943publishes the exact CI Wheel, sdist, manifest, Diff-bound classification, classification verification, and release-evidence record athttps://github.com/AdvancingTitans/agent-engineering-toolkit/releases/tag/v1.17.0. The deterministic classification covers 34 behavior-sensitive paths with exact test exceptions; Real Host adoption Gate isNOT_APPLICABLEbecause no governance asset was adopted. No PyPI publication was performed. - Separate human sign-off of the independently authored frozen Gold set, human
comprehension, implementation success after Plan consumption, and
computational reference-reuse rate remain
UNKNOWN. - Resume point: begin the next change from
v1.17.0. Preserve Bundle v1 compatibility, Graph authority, strictPROPOSEDplanning authority, explicit counter-evidence/UNKNOWNsemantics, and the pre-existing user-owned untracked files.
- Implemented the repository-side Phase 0–4 deliverables from
aet-github-star-growth-delivery-package.mdin order. Phase 5 and Phase 6 remain operational time windows and conditional follow-up work; they are documented asDEFERRED_WITH_REASON, not reported as elapsed results. - Added the installed-wheel
aet demo stale-proofsurface. It copies a packaged fixture into a bounded temporary Git workspace, runs real standard-library tests through the existing Quick Proof implementation, verifiesEXACT_MATCH, applies the manifest-bound source mutation, and verifiesRELEVANT_FILES_CHANGED. The demo adds no network, LLM, telemetry, edit, commit, push, merge, or release authority. - Added strict Demo v1 Schemas and edge-case coverage; focused conversion documentation and a 224-line root README; a static bilingual, no-tracking site; community templates; deterministic hero/social assets; Skills catalog validation; a read-only, argv-array GitHub Action template; privacy-preserving growth snapshots; launch-readiness checks; manual settings/outreach checklists; and channel-specific human publication briefs.
- Live T0 recorded 2 Stars, 0 forks, 2 contributors, 1 open issue, 3 total pull
requests, latest GitHub Release
v1.17.0, and PyPIv1.11.1. Traffic endpoints returned HTTP 403 and remainUNKNOWN; skills.sh indexing remainsUNAVAILABLE. - Version sources and
uv.lockare1.18.0. Builtagent_engineering_toolkit-1.18.0-py3-none-any.whl(1,437,358 bytes, public PyPI SHA-2569b6000ebd8e6cf9d174f1fb6797cf7299cf9246a5adcacd6c16033a63bd24f76) plus the matching CI sdist. The wheel contains the manifest/source/test fixture and passed clean-venv and publicuvxsmoke tests outside the checkout. - Final acceptance: pytest passed 591 tests and 607 subtests; the canonical
unittest gate passed 581 tests; focused compatibility regression passed 92
tests and 219 subtests; README/link/social/Skills/launch/Action checks passed;
Demo JSON validated against its Draft 2020-12 Schema;
git diff --checkpassed; AET self-audit has no FAIL/UNKNOWN and self-review has 7 PASS with no FAIL/UNKNOWN. - Released tag
v1.18.0from commit7ff1e0032827cce92b17777adc47af78977f0af3. Exact-tag CI Run30465380241, GitHub Release Run30466008349, and PyPI Run30466191944passed. Release ID361859093publishes the exact CI wheel, sdist, manifest, diff-bound classification, classification verification, and release-evidence record. - Pages is live at
https://advancingtitans.github.io/agent-engineering-toolkit/; About, Homepage, all 20 Topics, Social Preview, and Discussions are configured. Community Profile reports 100%, the repository remains pinned, and Known limitations are published in Discussion #5. - Resume point: preserve the released
v1.18.0authority boundaries and the pre-existing user-owned files. Phase 5/6 time-window measurements, skills.sh indexing, MCP Registry eligibility, and human-owned outreach remain follow-up work. The separate public Action repository is published from commit55f3c8f0e505fab5567694da2d1b5d834a227922with tagv1, Release ID361873799, and the publicAET EvidenceMarketplace listing. The user confirmed and accepted GitHub Marketplace Developer Agreement v2.4 at action time. Never convert unavailable or deferred state into a success claim. - Manual outreach progress on 2026-07-29/30: X, Reddit, Zhihu, Juejin, and
LinkedIn are public;
sdras/awesome-actions#874is open. Product Hunt is fully configured and scheduled for 2026-07-30 Pacific Time. Show HN rejected the authenticated submission under its temporary Show HN restriction for newer/unfamiliar accounts; no bypass was attempted, and comments remain human-only under the platform's current AI-comment rule. The owner cancelled V2EX publication. Exact evidence and URLs are recorded inops/growth/launch/manual-actions.md. - The reusable Codex Skill
promote-github-repois installed at/Users/yjw/.codex/skills/promote-github-repo. It generalizes the complete baseline, activation, README/conversion, Release/ecosystem, multi-channel publication, measurement, and reporting workflow while explicitly excluding V2EX. Its network-free audit script, promotion-pack templates, YAML, UI metadata, quick validation, credential-redaction smoke, real-repository smoke, and isolated forward tests pass.
- Implemented Phase 1–5 from
aet-behavioural-risk-diagnosis-delivery-package.mdas an AET Lab surface. The newsrc/aet/risk/core is deterministic, local, standard-library only, and consumes existing normalized Run, explicit Intent v2, Risk Policy, and optional validated Bundle evidence. It preservesPASS/FAIL/UNKNOWN/NOT_APPLICABLE, emits no aggregate risk rating or internal motive claim, and keeps every intervention atPROPOSEDauthority. - Added strict Policy, Diagnosis, and Forecast v1 schemas;
aet risk diagnose; gated experimentalaet risk forecast; same-context pathways; fixed intervention mapping; read-only MCP; an optional eleventh Evidence Atlasbehavioural-riskPerspective; TypeScript protocol types and local validators; offline evaluation fixtures; and bilingual Lab documentation. - Frozen evaluation report
.aet/risk-eval-report.jsonpasses with precision1.0, recall1.0, FPR0.0, 21 labels, and 3 explicitly preservedUNKNOWNlabels. Codex/Claude Code semantic parity is covered by E2E tests. - Controlled shadow code-review experiment
.aet/risk-business-experiment.jsoncovers 24 episodes / 72 factor labels across AET, ControlArena, and AgentRx source snapshots and two hosts. AET Risk reached exact-label accuracy1.0, positive recall1.0, FPR0.0, UNKNOWN preservation1.0, and cited-failure coverage1.0, exceeding the declared AET v1.18 and peer-inspired fixture baselines. This is reviewer decision-support evidence, not upstream peer execution or a production model-safety result. - Phase 4 prediction eligibility remains
FAIL: 24 controlled episodes are below the 200 independent-episode / 30 positive-outcome gate and do not provide independent production outcomes. Forecast therefore remains experimentalUNKNOWN, and the conditionalskills/aet-risksurface was not created. - Acceptance evidence: 50 Risk tests and the Risk E2E pass; full non-editable
Python regression passes 632 tests; TypeScript package build/check and 20
runtime tests pass; wheel/sdist build, isolated wheel install, all three
installed Risk schemas, and installed CLI smoke pass;
git diff --checkpasses. No commit, tag, remote release, or package publication was performed. - Independent two-human annotation/kappa and the five-minute human tabletop
usability threshold remain
UNKNOWN; do not infer them from automated fixtures. Resume by collecting independent adjudicated labels and at least the preregistered calibration volume before reconsidering forecast or Skill promotion. Preserve the user-owned dirty worktree and all existing Evidence First authority boundaries.
- This section supersedes the human-labelling resume point above. The user chose “release diagnosis first, keep prediction in research” and explicitly removed new manual annotation/usability work from the release scope.
- The synthetic 7-case suite is now explicitly contract regression only
(
release_gate=false,label_authority=synthetic_contract_fixture). The diagnosis release gate iseval/behavioural-risk/public-corpus.json: nine minimal action/outcome summaries from the official AgentDojo repository at commit089ed468cf3ed0322acc66b0211f26d9d90dbf60, with upstream programmatic utility/security labels and source/argument/effect hashes. No prompts, content, or benchmark solutions are redistributed. .aet/risk-public-benchmark-report.jsonpasses 9 cases / 27 factor labels with exact-case accuracy, precision, recall, and cited-failure coverage 1.0 and FPR 0.0. It declareshuman_validation_claimed=false,forecast_eligible=false, and scopediagnosis_only. AgentDojo has no monitoring-evasion label, so that factor is deliberatelyNOT_APPLICABLE.- Forecast has a code-level
FORECAST_RELEASE_STATE = "research_only"lock. Even otherwise valid calibration data returns gate FAIL and forecast UNKNOWN withforecast_research_only; public diagnosis cases must not be reused to bypass this lock. - Current-source acceptance: Risk 53 tests pass; full regression 635 tests
passes with
PYTHONPATH=src; protocol-types and evidence-bundle build/check, 20 runtime tests, and 2 dist smoke tests pass. Wheel/sdist build, clean-venv wheel install, installed Risk CLI diagnosis, and all 3 installed Risk schemas pass. 执行报告.mdcontains the final scope delta, evidence, and remaining risks. Resume by awaiting user acceptance. Do not commit, tag, push, create a GitHub release, or publish packages until the user explicitly authorizes release.
- Implemented Phase 0–5 of the graph-first review package without changing the
Evidence First authority boundary.
src/aet/review_graph/now builds a Git-bound Python Code Graph, composes it with Bundle Claim/Evidence and Improvement control nodes, validates strict graph/slice/manifest contracts, and returns bounded root, one-hop expansion, or staleUNKNOWNstop slices. - Added four Review Graph v1 Schemas, a non-overwriting hash-bound Review
Package,
aet review-graph build/validate/open/expand/export-compat, and two read-only MCP tools (aet_review_open,aet_review_expand). The default Agent input is canonicalreview/root.slice.json; full JSON/JSONL supports expansion, Mermaid remains a human projection, and legacyagent-context.json/agent-task.mdrequires explicit compatibility export. - Fail-closed checks cover missing references, incomplete controls, changed or protected scope, dynamic/ambiguous Python relations, budgets, extra/missing or tampered package files, symlinks, output overwrite, and snapshot drift. No aggregate score, model call, command execution, source edit, or Repo Archaeologist dependency was added.
- The AET empty-tool-result business case produces a 12-node/13-edge root at 6,505 stored bytes versus 6,522 bytes for the legacy Agent Task plus two Mermaid files, and 8,468 bytes for the legacy minimum raw materials. A single-run, isolated read-only Codex comparison scored the final root 8/8, legacy input 7/8, and code-only graph 1/8 on the preregistered intervention fields. This is a one-case preliminary observation, not a general model or peer-superiority claim.
- Acceptance: Review Graph
21 tests / OK; MCP, portable CLI, and original Improvement integration22 tests / OK; README/growth gate4 tests / OK; full current-source regression656 tests / OKin 72.790 seconds.执行报告.mdanddocs/review-graph.mdcontain the detailed results and limitations. - Resume by awaiting user acceptance. Preserve the user-owned dirty worktree. Do not commit, tag, push, publish, or create a GitHub Release until the user explicitly authorizes it.
-
The user explicitly authorized a direct
mainpublication and GitHub Release. PyPI publication remains out of scope; public PyPI is v1.18.0 and no PyPI action was requested. -
Combined the previously unreleased Behavioural Risk Diagnosis and Review Graph implementations without weakening four-state authority. Forecast stays hard-locked to research-only
UNKNOWN; Review Graph snapshot drift and package tampering remain fail-closed stop conditions. -
Rebuilt the 213-line English README and 201-line Chinese README around AET's current definition as a local Evidence Plane. Corrected public PyPI status, exact GitHub wheel installation, current CLI/MCP surfaces, static and dynamic flows, case library, context-byte measurement, and explicit comparison limits. Added a commit-pinned factual comparison with
code-review-graphand a production-shaped refresh-token race case that separates the human Mermaid view from the bounded Agent slice and explicit stop rules. -
Rebuilt and visually reviewed the bilingual v1.19 architecture SVG/PNG/GIF, 1600×1000 panoramas, and exact 30-second silent 1600×900 H.264 introductions. Fireworks geometry/composition and 115-frame semantic motion reports pass; the v3 media manifest binds 16 distinct artifacts by SHA-256.
-
Version sources and lock are 1.19.0. The 657-test current-source regression, strict self-audit, focused README/media/Atlas checks, all six README Mermaid parses, the 9-case AgentDojo diagnosis gate, TypeScript gates, package build, package-content checks, and isolated-wheel Demo/CLI smoke pass. Diff-bound release classification, exact-commit CI, and GitHub Release evidence also passed in the runs recorded below.
-
Released immutable tag
v1.19.0from commitd9beff78af1bf2079af110f4d009686959d664ff. Main CI Run31265258405, tag CI Run31265258444, and GitHub Release Run31265675517passed. Release ID367240504publishes the exact CI wheel, sdist, manifest, diff-bound classification, classification verification, and release evidence. -
Post-release verification caught one stale live-distribution fact: public PyPI is v1.18.0, not v1.11.1. Do not mutate the released v1.19.0 tag or assets; v1.19.1 is the documentation/package-metadata patch candidate with no Evidence or authority semantic change.
- Verified the official PyPI JSON reports v1.18.0 and corrected the English, Chinese, PyPI, quick-start, launch-gate, Action-template, and exact GitHub install documentation accordingly.
- Version metadata and lock are 1.19.1. This patch changes no Evidence, authority, Review Graph, Behavioural Risk, media, or protocol semantics and performs no PyPI publication.