Skip to content

fix(bench): kill a TTY exchange on inactivity instead of total time - #3997

Merged
Alan-TheGentleman merged 2 commits into
Gentleman-Programming:mainfrom
Denver2828:fix/3971-pty-axis-deadline
Sep 5, 2026
Merged

Alan-TheGentleman merged 2 commits into
Gentleman-Programming:mainfrom
Denver2828:fix/3971-pty-axis-deadline

Conversation

@Denver2828

@Denver2828 Denver2828 commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

🔗 Linked Issue

Closes #3971


🏷️ PR Type

What kind of change does this PR introduce?

  • type:bug — Bug fix (non-breaking change that fixes an issue)
  • type:feature — New feature (non-breaking change that adds functionality)
  • type:docs — Documentation only
  • type:refactor — Code refactoring (no functional changes)
  • type:chore — Build, CI, or tooling changes
  • type:breaking-change — Breaking change (fix or feature that changes existing behavior)

📝 Summary

On Linux /dev/ptmx under transition-axis load (#3971), waitForReviewModeTTY kills an exchange on a fixed total deadline even while the TUI banner is still actively trickling in, producing the CI failure read TUI before "RDD is currently ENABLED globally." ... context deadline exceeded.

This PR replaces the total-time budget with an inactivity watchdog: the exchange dies only after ttyTimeout (10s) without receiving a single PTY byte, and every received byte resets the timer. A new ttyOverallTimeout (2min) caps the whole exchange so a forever-chattering terminal still dies. Both kill causes unwrap to context.DeadlineExceeded, so caller classification is unchanged.

A blind total-budget bump was considered and rejected (rationale recorded in the commit body): any fixed total still loses under enough load, and it slows detection of every genuine hang.


📂 Changes

File / Area What Changed
bench/runner.go Replace the total-time PTY read deadline with an inactivity watchdog (ttyTimeout, reset on every byte) plus an overall exchange cap (ttyOverallTimeout); both unwrap to context.DeadlineExceeded
bench/runner_test.go Regression test replaying a slow trickling first frame through the real waitForReviewModeTTY (RED on merge base, GREEN on head), plus two tests pinning the overall cap and the inactivity budget

🤖 AI Assistance

Select exactly one option. Do not check both options.

  • None — No material AI assistance was used.
  • Material assistance used — Complete all applicable declaration fields below.

Tool/model (if known): Claude Code (Fable)

Material scope: Implementation drafting, test authoring, and verification orchestration, under the contributor's direction and review.

Verification performed: Regression test observed failing on the merge base (1863432) with the CI failure shape and passing on the head (including -count=5); gofmt -l clean; go run ./internal/gofmtcheck OK; go vet ./... (bench module) OK; go build ./... (root module) OK; git diff --check clean; bench full-module failing set compared byte-identical before/after the change.

Trivial formatting, spelling, minor autocomplete, search/navigation, and trivial, non-substantive mechanical transformations do not need to be itemized. See AI_POLICY.md for the canonical policy.


🧪 Test Plan

Unit Tests

go test ./...

Run in the bench module. New tests:

  • TestRunTTYToleratesASlowFirstFrameUnderLoad — scripted terminal replays the j121 banner in trickling frames (silences of 400/400/700ms under a 1s stand-in budget, 1.5s total, over it), read through the real waitForReviewModeTTY. FAILS on merge base 1863432 with the CI failure shape; PASSES on head, also at -count=5.
  • TestRunTTYOverallCapKillsAForeverChatteringExchange and TestRunTTYInactivityBudgetStillKillsASilentExchange — pin the overall cap and the inactivity budget.

The bench full-module failing set is byte-identical before and after this change: 12 pre-existing failures (the #3934 ConPTY pair plus ten POSIX-fixture failures matching #2603). This PR fixes none of them and adds none. #3934 (native Windows ConPTY) is a different defect from #3971 (Linux /dev/ptmx under load) in the same runner; it is used here only as a cross-lane non-regression bar.

Go Format

go run ./internal/gofmtcheck

OK (gofmt -l also clean).

E2E Tests (Docker required)

cd e2e && ./docker-test.sh

NOT RUN locally: Docker is unavailable in the local environment; deferred to CI. The change is confined to the bench module.

Benchmark Validation

This changes the benchmark runner, so benchmark validation applies. Unit-level validation is covered by the three tests above. The full driven-journey corpus run is NOT RUN locally (it needs the product binary and a long transition-axis run) and is deferred to CI.

Also run locally: go vet ./... (bench module) OK, go build ./... (root module) OK, git diff --check clean.

  • Unit tests pass (go test ./...) — new/changed tests pass; the 12 pre-existing bench failures are unchanged byte-identical before/after (see above)
  • Go format passes (go run ./internal/gofmtcheck)
  • E2E tests pass (cd e2e && ./docker-test.sh) — deferred to CI (Docker unavailable locally)
  • Manually tested locally

🤖 Automated Checks

The following checks run automatically on this PR:

Check Status Description
Check PR Cognitive Load ⏳ PR should stay within 400 changed lines (additions + deletions) or use size:exception
Check Issue Reference ⏳ PR body must contain Closes/Fixes/Resolves #N
Check Issue Has status:approved ⏳ Linked issue must have been approved before work began
Check PR Has type:* Label ⏳ Exactly one type:* label must be applied
Unit Tests ⏳ go test ./... must pass
Go Format ⏳ go run ./internal/gofmtcheck must pass
E2E Tests ⏳ cd e2e && ./docker-test.sh must pass

✅ Contributor Checklist

  • PR is linked to an issue with status:approved
  • PR stays within 400 changed lines, or I have requested/obtained maintainer-applied size:exception with rationale documented
  • I have added the appropriate type:* label to this PR
  • Unit tests pass (go test ./...)
  • Go format passes (go run ./internal/gofmtcheck)
  • E2E tests pass (cd e2e && ./docker-test.sh) — deferred to CI (see Test Plan)
  • Benchmark validation completed, or this change is not applicable to the benchmark (explain why in the Test Plan) — unit-level done; full corpus run deferred to CI (see Test Plan)
  • I have updated documentation if necessary
  • My commits follow Conventional Commits format
  • I understand, reviewed, and take responsibility for the complete submission
  • I selected exactly one AI-assistance option and, if material assistance was used, completed all applicable declaration fields
  • My commits do not include Co-Authored-By trailers

💬 Notes for Reviewers

For production Go changes in internal/cli, internal/reviewtransaction, or internal/sddstatus:

Not applicable — this PR changes only the bench module; no qualifying guard population is touched.

  • Identify any qualifying security, integrity, admission, repair, or governance guard and challenge its legitimate input population against real-world evidence.
  • Confirm its guard:population direction and claim are adjacent and accurate, and that .guard-population-baseline.txt changed only when the guard contract intentionally changed.
  • Do not treat a passing declaration/registry check as proof that no qualifying guard was omitted or that the population claim is semantically complete.

Summary by CodeRabbit

  • Bug Fixes

    • Improved terminal interaction timeouts to tolerate slow output while still stopping exchanges that become inactive or run indefinitely.
    • Preserved successful handling of delayed terminal responses, continuous output, terminal closure, and quit actions.
    • Prevented an update notification from interrupting the related terminal workflow.
  • Tests

    • Added coverage for slow, silent, continuously active, and automatically terminated terminal exchanges.

j121-rdd-tui-controls-global-mode fails intermittently in the CI
transition-axis run while passing the core run minutes earlier (#3971):
under that run's CPU contention the TUI keeps painting, but the whole
four-screen exchange outlives the fixed 10s deadline that
context.WithTimeout applied to the entire runTTY call, so the child is
killed mid-read ("read TUI before ...: input/output error; context
deadline exceeded; signal: killed").

Of the two shapes the issue admits — a load-reflecting deadline or a
read that tolerates the slower first frame — this takes the second: a
ttyWatchdog now expires the exchange only after ttyTimeout (still 10s)
WITHOUT receiving a single PTY byte, and every received byte resets that
timer, so a slow, trickling frame under load survives while a hung TUI
still dies after exactly the old budget. A new generous
ttyOverallTimeout (2min) caps the whole exchange so a TUI that paints
forever without reaching the expected screen still terminates
deterministically. Both causes unwrap to context.DeadlineExceeded, so
callers classify them exactly as before. A blind bump of the total
budget was rejected because it would slow every genuinely hung exchange
by the same factor without removing the load sensitivity.

TestRunTTYToleratesASlowFirstFrameUnderLoad is the CI failure scaled
down: a scripted terminal keeps making progress (every silence shorter
than the budget) but completes the j121 banner only after more than the
budget in total, read through j121's own waitForReviewModeTTY helper.
On the previous commit it fails with the CI shape ("read TUI before
\"RDD is currently ENABLED globally.\" ... context deadline exceeded");
with this change it passes. TestRunTTYOverallCapKillsAForeverChattering-
Exchange and TestRunTTYInactivityBudgetStillKillsASilentExchange pin
the two new budgets without real 10s waits.

Native-Windows bar: the bench module's failing-test set is byte-for-byte
identical before and after this change (the #3934 pair plus the ten
pre-existing fixture failures matching #2603); this commit fixes none of
them and introduces no new failure. The driven journey corpus itself is
deferred to CI.
@coderabbitai

coderabbitai Bot commented Aug 31, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: ASSERTIVE

Plan: Team

Run ID: 79907b9f-38b8-4c77-8f93-cf92e12b4c3b

📥 Commits

Reviewing files that changed from the base of the PR and between 7a5b914 and 662aaf8.

📒 Files selected for processing (1)
  • bench/journeys_issue3766.go

Included review availability: Your plan provides up to 8 included reviews per hour; 7 remain after this review.


📝 Walkthrough

Walkthrough

Changes

The TTY runner now enforces separate inactivity and overall exchange deadlines. PTY byte reads refresh the inactivity deadline. Tests cover slow, continuous, and silent output. The issue #3766 journey disables the update modal with a sandbox fixture.

TTY exchange watchdog

Layer / File(s) Summary
Progress-aware TTY watchdog
bench/runner.go
runTTYWithDeadlines tracks PTY progress with ttyProgressReader. ttyWatchdog enforces inactivity and overall budgets. Timeout handling uses context.Cause(ctx).
Deterministic deadline tests
bench/runner_test.go
Fake PTY infrastructure and tests validate delayed progress, overall timeout, inactivity timeout, transcript capture, quit handling, and cleanup.
Issue #3766 journey setup
bench/journeys_issue3766.go
The journey seeds a current last_update_check timestamp in sandbox state before the repository and TUI steps.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Merge Risk: ⚪ Minimal · up to 662aa

The PR changes benchmark TTY timeout handling to tolerate active but slow output while retaining inactivity and overall limits; no actionable merge-blocking risk remains after normal checks and review.

Sequence Diagram(s)

sequenceDiagram
  participant runTTYWithDeadlines
  participant ttyProgressReader
  participant ttyWatchdog
  runTTYWithDeadlines->>ttyWatchdog: Create inactivity and overall budgets
  ttyProgressReader->>ttyWatchdog: Record each PTY byte
  ttyWatchdog->>runTTYWithDeadlines: Cancel with timeout cause
  runTTYWithDeadlines->>ttyWatchdog: Read context cancellation cause
Loading

Suggested reviewers: decode2

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 58.33% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 12 functions across 3 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly describes the primary change: TTY exchanges now terminate on inactivity instead of a fixed total-time deadline.
Linked Issues check ✅ Passed The changes satisfy issue [#3971]. The TTY watchdog refreshes the 10-second inactivity budget on received bytes, enforces a 2-minute overall cap, preserves timeout classification, and adds regression …
Out of Scope Changes check ✅ Passed All changes are related to the linked issue [#3971]. The runner changes, regression tests, and update-check fixture directly support reliable TTY benchmark journeys and do not introduce unrelated scop…
Full details: Linked Issues check

Explanation

The changes satisfy issue [#3971]. The TTY watchdog refreshes the 10-second inactivity budget on received bytes, enforces a 2-minute overall cap, preserves timeout classification, and adds regression tests for slow, continuous, and silent exchanges. The update-check fixture also prevents the modal from obscuring the affected TUI journey.

Full details: Out of Scope Changes check

Explanation

All changes are related to the linked issue [#3971]. The runner changes, regression tests, and update-check fixture directly support reliable TTY benchmark journeys and do not introduce unrelated scope.

  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

ardelperal added a commit to ardelperal/gentle-ai that referenced this pull request Aug 31, 2026
…side

Empty commit; branch content unchanged at fbe6c03 base + cherry-picks.
Targeted packages verified locally: internal/components/sdd (74.6s),
internal/assets (50.5s), internal/reviewtransaction (212.1s) all green.

The previous Windows Runtime failure on TestRepairClassifiedAuthority
ConcurrentExecutionCommitsAndReplays is currently red on Alan's own
in-flight PRs (Gentleman-Programming#3997, Gentleman-Programming#3957) which do not modify reviewtransaction,
so this push is to rule out a diff-local regression on a healthy runner.
…Available modal cannot cover the exchange

The watchdog diagnostic from the previous commit exposed a second
manifestation of #3971, distinct from the load-slowness that commit
addresses: the bench sandbox gives each journey a fresh HOME with no
state.json, so CheckAllWithCooldown (internal/update/cooldown.go) finds
no LastUpdateCheck and runs the launch update check. When upstream main
is newer than the CI-built binary, the TUI shows the Update Available
modal (internal/tui/screens/update_prompt.go), which covers the menu
j121's exchange waits for ("Start installation" never arrives), and
after 10s of true silence the watchdog kills the run.

Seed j121's sandbox HOME with a state.json carrying a current
last_update_check before the TTY exchange, mirroring the precedent in
bench/journeys_issue_3561.go (green in CI for the same reason). No
state.json exists at that point and the journey runs against a
not-installed state, so only the cooldown field is seeded.
@Denver2828

Copy link
Copy Markdown
Contributor Author

The Unit Tests red on commit 1 (7a5b914) is expected and diagnostic, not a regression.

#3971 turns out to have two distinct manifestations in the same fragile journey. The original occurrences are load slowness mid-exchange, which commit 1's inactivity watchdog addresses. Running under that watchdog, CI then exposed the second: the bench sandbox gives every journey a fresh HOME with no update-check cooldown, so the launch update check runs, and when upstream main is ahead of the CI-built binary the "Update Available" modal covers the menu j121 waits on — its first expected text, "Start installation", never arrives. Where this death previously surfaced only as an opaque read /dev/ptmx: input/output error, the watchdog now appends no PTY output for 10s: context deadline exceeded and the captured screen showing the covering modal, which is what made the cause legible.

Commit 2 (662aaf8, +15/−0, j121's journey file only) seeds a current last_update_check into the sandbox state before the exchange, mirroring the #3561 journey fixture that is green in CI for the same reason. It deliberately does not seed installed_agents: j121 exercises the not-installed state. j120 and j122 share the same latent exposure; that is outside this issue's scope and may warrant a follow-up. The green proof landed: this PR's CI run on 662aaf8 completed with j121 green — 61 journeys, 0 failed.

@Alan-TheGentleman Alan-TheGentleman added the type:bug Bug fix label Sep 5, 2026
@Alan-TheGentleman
Alan-TheGentleman merged commit 20d7574 into Gentleman-Programming:main Sep 5, 2026
24 of 26 checks passed
@decode2 decode2 mentioned this pull request Sep 5, 2026
18 of 24 tasks
decode2 pushed a commit that referenced this pull request Sep 10, 2026
#3023)

* fix(opencode): keep SDD phase commands in primary orchestrator session

OpenCode treats `subtask: true` in command YAML frontmatter as a forced
sub-agent invocation, overriding the `gentle-orchestrator` `mode: primary`
setting in `internal/assets/opencode/sdd-overlay-single.json`. The five
SDD phase commands (`sdd-init`, `sdd-explore`, `sdd-apply`, `sdd-verify`,
`sdd-archive`) all carried that flag, producing a redundant nested hop.

The orchestrator's `permission.task` allowlist already lists every phase
worker, so removing `subtask: true` only changes session topology — not
delegation capability — and lets the orchestrator stay primary.

Closes #2939

Source scope (5 files, 1 line removed each):
- internal/assets/opencode/commands/sdd-init.md
- internal/assets/opencode/commands/sdd-explore.md
- internal/assets/opencode/commands/sdd-apply.md
- internal/assets/opencode/commands/sdd-verify.md
- internal/assets/opencode/commands/sdd-archive.md

Golden regeneration (2 files, 1 line removed each):
- testdata/golden/sdd-opencode-cmd-sdd-init.golden
- testdata/golden/sdd-opencode-cmd-sdd-apply.golden

New regression test:
- internal/components/sdd_opencode_subtask_test.go -- byte-level assertion
  that none of the 5 source files reintroduce `subtask: true`

Out of scope (intentional, per user decision 2026-08-11):
- internal/assets/opencode/commands/sdd-onboard.md -- also carries
  `subtask: true` but the issue enumerates it as a negative control.
  Follow-up issue to be filed separately.
- skill-creator.md / skill-registry.md -- different domain.
- All non-OpenCode adapters -- none use per-command `subtask`.

* fix(opencode): include sdd-onboard in the subtask-removal regression per decode2

- Remove the `subtask: true` line from
  internal/assets/opencode/commands/sdd-onboard.md. The other five SDD
  phase commands were already corrected in the parent commit; sdd-onboard
  was missed because it is the only one whose first paragraph mentions
  the hidden "sdd-onboard" sub-agent (rather than the more obvious
  sdd-apply/sdd-archive pair) and was easy to overlook in the original
  sweep. The yaml frontmatter now matches the other six phase commands.
- Extend TestNoSubtaskOnSDDOpenCodeCommands to enumerate all six
  frontmatter files (sdd-init, sdd-explore, sdd-apply, sdd-verify,
  sdd-archive, sdd-onboard). The static byte-level check now catches a
  regression on any of them, and the docstring and comments reflect the
  six-command reality.
- Installation/injection coverage: inject.go's compatibilitySDDSkillIDs
  and claudeModelAssignmentRowOrder already include sdd-onboard; no
  change needed there.

* test(opencode): prove installed SDD commands stay parent-owned

Completes the coverage decode2 requested on PR #3023 for #2939: the
existing guard only scanned the source assets, so an injection path that
re-introduced `subtask: true` would have passed unnoticed.

- TestInjectOpenCodeSDDCommandsRemainParentOwned runs Inject for the
  OpenCode adapter and asserts each installed delegating SDD command
  keeps `agent: gentle-orchestrator` and carries no `subtask: true` in
  its installed frontmatter.
- TestNoSubtaskOnSDDOpenCodeCommands now scopes its assertion to the
  YAML frontmatter block instead of the whole file, so prose documenting
  the field cannot fail the source ratchet.
- Both guards also pin that removing the field never detaches the
  command from gentle-orchestrator routing.

RED/GREEN: re-adding `subtask: true` to any listed asset fails both
tests; removing it passes. go test ./internal/components/... green,
gofmtcheck green, e2e 3/3 platforms.

* fix(opencode): keep sdd-research parent-owned

sdd-research joined the OpenCode command set after the #2939 approval
(74cc41b) with the exact shape the issue fixed: `agent:
gentle-orchestrator` plus `subtask: true`, which forces the primary
orchestrator into a sub-agent invocation before the hidden sdd-research
phase worker can launch.

Same root correction as the approved six: remove only the field, keep
the routing. Both ratchets (source assets and installed commands) now
enumerate it, and its golden is regenerated.

Kept as a separate work unit so maintainers can drop the scope
expansion with a single revert if they prefer the strictly approved
six-command set.

* test(opencode): decode frontmatter YAML in the parent-owned ratchets

Addresses the CodeRabbit finding on #3864: substring-matching the raw
frontmatter block could false-pass or false-fail when a description or
comment mentions the literal `agent:` or `subtask:` values. Both
ratchets now unmarshal the frontmatter into a typed struct and assert
the exact agent value and the effective subtask flag.

Verified:
- RED: re-adding `subtask: true` to an asset fails both ratchets.
- False-pass guard: a YAML comment (`# subtask: true`) and a quoted
  description containing the literal both pass with the field absent,
  and invalid frontmatter fails loudly at decode.
- go test ./internal/components/ ./internal/components/sdd/ green.

* style(opencode): sort yaml import in the injection ratchet

gofmt requires the gopkg.in/yaml.v3 import at the end of the sorted
block; the previous commit inserted it mid-block and every CI job gates
on Go Format passing.

* test(opencode): cover the upgrade and repeat-install paths in the ratchet

Addresses the CodeRabbit finding on #3864: the ratchet ran Inject once,
proving only the fresh-install path. It now also covers:

- Upgrade: a pre-fix install whose sdd-init.md still declares
  `subtask: true` must be overwritten by a re-sync, never preserved,
  so users upgrading from an affected version reach the fixed state.
- Repeat install: a second Inject must report no changes and leave
  the parent-owned metadata in place.

Assertions extracted to assertSDDCommandsParentOwned so both install
passes share the exact same metadata checks.

* ci(retouch): retrigger CI to confirm Windows Runtime flake is runner-side

Empty commit; branch content unchanged at fbe6c03 base + cherry-picks.
Targeted packages verified locally: internal/components/sdd (74.6s),
internal/assets (50.5s), internal/reviewtransaction (212.1s) all green.

The previous Windows Runtime failure on TestRepairClassifiedAuthority
ConcurrentExecutionCommitsAndReplays is currently red on Alan's own
in-flight PRs (#3997, #3957) which do not modify reviewtransaction,
so this push is to rule out a diff-local regression on a healthy runner.

---------

Co-authored-by: danielgap <soydanielgap@gmail.com>
Co-authored-by: ardelperal <ardelperal@users.noreply.github.com>
Alan-TheGentleman added a commit that referenced this pull request Sep 24, 2026
fix(bench): kill a TTY exchange on inactivity instead of total time
Alan-TheGentleman pushed a commit that referenced this pull request Sep 24, 2026
#3023)

* fix(opencode): keep SDD phase commands in primary orchestrator session

OpenCode treats `subtask: true` in command YAML frontmatter as a forced
sub-agent invocation, overriding the `gentle-orchestrator` `mode: primary`
setting in `internal/assets/opencode/sdd-overlay-single.json`. The five
SDD phase commands (`sdd-init`, `sdd-explore`, `sdd-apply`, `sdd-verify`,
`sdd-archive`) all carried that flag, producing a redundant nested hop.

The orchestrator's `permission.task` allowlist already lists every phase
worker, so removing `subtask: true` only changes session topology — not
delegation capability — and lets the orchestrator stay primary.

Closes #2939

Source scope (5 files, 1 line removed each):
- internal/assets/opencode/commands/sdd-init.md
- internal/assets/opencode/commands/sdd-explore.md
- internal/assets/opencode/commands/sdd-apply.md
- internal/assets/opencode/commands/sdd-verify.md
- internal/assets/opencode/commands/sdd-archive.md

Golden regeneration (2 files, 1 line removed each):
- testdata/golden/sdd-opencode-cmd-sdd-init.golden
- testdata/golden/sdd-opencode-cmd-sdd-apply.golden

New regression test:
- internal/components/sdd_opencode_subtask_test.go -- byte-level assertion
  that none of the 5 source files reintroduce `subtask: true`

Out of scope (intentional, per user decision 2026-08-11):
- internal/assets/opencode/commands/sdd-onboard.md -- also carries
  `subtask: true` but the issue enumerates it as a negative control.
  Follow-up issue to be filed separately.
- skill-creator.md / skill-registry.md -- different domain.
- All non-OpenCode adapters -- none use per-command `subtask`.

* fix(opencode): include sdd-onboard in the subtask-removal regression per decode2

- Remove the `subtask: true` line from
  internal/assets/opencode/commands/sdd-onboard.md. The other five SDD
  phase commands were already corrected in the parent commit; sdd-onboard
  was missed because it is the only one whose first paragraph mentions
  the hidden "sdd-onboard" sub-agent (rather than the more obvious
  sdd-apply/sdd-archive pair) and was easy to overlook in the original
  sweep. The yaml frontmatter now matches the other six phase commands.
- Extend TestNoSubtaskOnSDDOpenCodeCommands to enumerate all six
  frontmatter files (sdd-init, sdd-explore, sdd-apply, sdd-verify,
  sdd-archive, sdd-onboard). The static byte-level check now catches a
  regression on any of them, and the docstring and comments reflect the
  six-command reality.
- Installation/injection coverage: inject.go's compatibilitySDDSkillIDs
  and claudeModelAssignmentRowOrder already include sdd-onboard; no
  change needed there.

* test(opencode): prove installed SDD commands stay parent-owned

Completes the coverage decode2 requested on PR #3023 for #2939: the
existing guard only scanned the source assets, so an injection path that
re-introduced `subtask: true` would have passed unnoticed.

- TestInjectOpenCodeSDDCommandsRemainParentOwned runs Inject for the
  OpenCode adapter and asserts each installed delegating SDD command
  keeps `agent: gentle-orchestrator` and carries no `subtask: true` in
  its installed frontmatter.
- TestNoSubtaskOnSDDOpenCodeCommands now scopes its assertion to the
  YAML frontmatter block instead of the whole file, so prose documenting
  the field cannot fail the source ratchet.
- Both guards also pin that removing the field never detaches the
  command from gentle-orchestrator routing.

RED/GREEN: re-adding `subtask: true` to any listed asset fails both
tests; removing it passes. go test ./internal/components/... green,
gofmtcheck green, e2e 3/3 platforms.

* fix(opencode): keep sdd-research parent-owned

sdd-research joined the OpenCode command set after the #2939 approval
(74cc41b) with the exact shape the issue fixed: `agent:
gentle-orchestrator` plus `subtask: true`, which forces the primary
orchestrator into a sub-agent invocation before the hidden sdd-research
phase worker can launch.

Same root correction as the approved six: remove only the field, keep
the routing. Both ratchets (source assets and installed commands) now
enumerate it, and its golden is regenerated.

Kept as a separate work unit so maintainers can drop the scope
expansion with a single revert if they prefer the strictly approved
six-command set.

* test(opencode): decode frontmatter YAML in the parent-owned ratchets

Addresses the CodeRabbit finding on #3864: substring-matching the raw
frontmatter block could false-pass or false-fail when a description or
comment mentions the literal `agent:` or `subtask:` values. Both
ratchets now unmarshal the frontmatter into a typed struct and assert
the exact agent value and the effective subtask flag.

Verified:
- RED: re-adding `subtask: true` to an asset fails both ratchets.
- False-pass guard: a YAML comment (`# subtask: true`) and a quoted
  description containing the literal both pass with the field absent,
  and invalid frontmatter fails loudly at decode.
- go test ./internal/components/ ./internal/components/sdd/ green.

* style(opencode): sort yaml import in the injection ratchet

gofmt requires the gopkg.in/yaml.v3 import at the end of the sorted
block; the previous commit inserted it mid-block and every CI job gates
on Go Format passing.

* test(opencode): cover the upgrade and repeat-install paths in the ratchet

Addresses the CodeRabbit finding on #3864: the ratchet ran Inject once,
proving only the fresh-install path. It now also covers:

- Upgrade: a pre-fix install whose sdd-init.md still declares
  `subtask: true` must be overwritten by a re-sync, never preserved,
  so users upgrading from an affected version reach the fixed state.
- Repeat install: a second Inject must report no changes and leave
  the parent-owned metadata in place.

Assertions extracted to assertSDDCommandsParentOwned so both install
passes share the exact same metadata checks.

* ci(retouch): retrigger CI to confirm Windows Runtime flake is runner-side

Empty commit; branch content unchanged at fbe6c03 base + cherry-picks.
Targeted packages verified locally: internal/components/sdd (74.6s),
internal/assets (50.5s), internal/reviewtransaction (212.1s) all green.

The previous Windows Runtime failure on TestRepairClassifiedAuthority
ConcurrentExecutionCommitsAndReplays is currently red on Alan's own
in-flight PRs (#3997, #3957) which do not modify reviewtransaction,
so this push is to rule out a diff-local regression on a healthy runner.

---------

Co-authored-by: danielgap <soydanielgap@gmail.com>
Co-authored-by: ardelperal <ardelperal@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:bug Bug fix

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bug(bench): j121-rdd-tui-controls-global-mode flakes with a pty deadline in the CI transition-axis run

2 participants