📝 Bug Description
Two grammars in this repository define what counts as a "runnable continuation" for a refusal message, and they disagree today:
- The static ratchet (
internal/cli/refusal_resolution_ratchet_test.go, refusalRatchetNamedContinuationRegexp, ~line 151) matches only gentle-ai [a-z][a-z-]* — product name plus one verb, nothing after. A refusal printing gentle-ai review status --cwd <repo> passes this check.
- The benchmark classifier (
bench/classify.go, HasRunnableCommand with placeholderRun = <[^>\s]*>) requires at least one real argument and rejects unfilled <placeholder> tokens. The same message is classified not runnable there.
So a message the benchmark classifies out-of-band satisfies the ratchet that CI runs as a required merge check. The ratchet's own header states the limitation plainly: "it cannot catch a WRONG named continuation — one that parses, runs, and does not help."
This seam is the shared root of a recurring class of reports where printed guidance cannot be executed as printed:
And per #2834's measurement, internal/sddstatus/runtime_ledger.go alone declares 36 refusals with no runnable exit.
🔄 Steps to Reproduce
- Observe
refusalRatchetNamedContinuationRegexp = regexp.MustCompile(gentle-ai [a-z][a-z-]*) in internal/cli/refusal_resolution_ratchet_test.go.
- Evaluate it against a refusal text such as
run gentle-ai review status --cwd `` — it matches (gentle-ai review), so the site passes the ratchet.
- Run the same text through
bench/classify.go HasRunnableCommand: placeholderRun finds <repo> and the classifier returns false.
- Result: the two enforcement layers disagree about the same emitted sentence, and CI's required check is the weaker of the two.
✅ Expected Behavior
The static ratchet enforces the same printed-command grammar the benchmark already enforces:
- a named continuation must carry real arguments, not stop at the verb;
- an unfilled
<placeholder> tail is accepted only when the site explicitly declares operator knowledge via the existing refusal:by-design operator-knowledge: marker (or holds a baseline row);
- the baseline regenerates once at the switch and may only shrink afterwards;
- existing CI (already a required merge check) then enforces the stronger grammar going forward — no new CI lane needed.
Both command splitters needed already exist and are deliberately mirrored (internal/cli/review_printed_command.go SplitPrintedCommandWords and its bench twin); no third splitter should be written.
❌ Actual Behavior
The ratchet accepts verb-only matches regardless of what follows, including tails the benchmark classes as not runnable. Wrong or dead-end continuations therefore pass the required merge check until someone files a per-instance report.
🌍 Environment
- Current
main (verified against source on 2026-08-22)
- Evidence gathered on Windows; the grammar seam itself is platform-independent
💡 Proposed Scope
Static unification slice only: swap the ratchet's continuation grammar to the shared printed-command grammar, add the operator-knowledge marker path, regenerate the baseline once. Estimated ~150–250 changed lines including tests; one revertible PR from current main. The dynamic half ("follow the printed recipe until the block clears") already has a canonical pattern (internal/cli/review_abandon_message_test.go) and grows case-by-case afterwards.
Non-goals: no execution sandbox in the unit lane, no new wire vocabulary, no parallel command splitter, no change to bench classification semantics.
📝 Bug Description
Two grammars in this repository define what counts as a "runnable continuation" for a refusal message, and they disagree today:
internal/cli/refusal_resolution_ratchet_test.go,refusalRatchetNamedContinuationRegexp, ~line 151) matches onlygentle-ai [a-z][a-z-]*— product name plus one verb, nothing after. A refusal printinggentle-ai review status --cwd <repo>passes this check.bench/classify.go,HasRunnableCommandwithplaceholderRun = <[^>\s]*>) requires at least one real argument and rejects unfilled<placeholder>tokens. The same message is classified not runnable there.So a message the benchmark classifies out-of-band satisfies the ratchet that CI runs as a required merge check. The ratchet's own header states the limitation plainly: "it cannot catch a WRONG named continuation — one that parses, runs, and does not help."
This seam is the shared root of a recurring class of reports where printed guidance cannot be executed as printed:
And per #2834's measurement,
internal/sddstatus/runtime_ledger.goalone declares 36 refusals with no runnable exit.🔄 Steps to Reproduce
refusalRatchetNamedContinuationRegexp = regexp.MustCompile(gentle-ai [a-z][a-z-]*)ininternal/cli/refusal_resolution_ratchet_test.go.rungentle-ai review status --cwd `` — it matches (gentle-ai review), so the site passes the ratchet.bench/classify.goHasRunnableCommand:placeholderRunfinds<repo>and the classifier returns false.✅ Expected Behavior
The static ratchet enforces the same printed-command grammar the benchmark already enforces:
<placeholder>tail is accepted only when the site explicitly declares operator knowledge via the existingrefusal:by-design operator-knowledge:marker (or holds a baseline row);Both command splitters needed already exist and are deliberately mirrored (
internal/cli/review_printed_command.goSplitPrintedCommandWordsand its bench twin); no third splitter should be written.❌ Actual Behavior
The ratchet accepts verb-only matches regardless of what follows, including tails the benchmark classes as not runnable. Wrong or dead-end continuations therefore pass the required merge check until someone files a per-instance report.
🌍 Environment
main(verified against source on 2026-08-22)💡 Proposed Scope
Static unification slice only: swap the ratchet's continuation grammar to the shared printed-command grammar, add the operator-knowledge marker path, regenerate the baseline once. Estimated ~150–250 changed lines including tests; one revertible PR from current
main. The dynamic half ("follow the printed recipe until the block clears") already has a canonical pattern (internal/cli/review_abandon_message_test.go) and grows case-by-case afterwards.Non-goals: no execution sandbox in the unit lane, no new wire vocabulary, no parallel command splitter, no change to bench classification semantics.