Skip to content

fix(cli): make the refusal-resolution ratchet enforce the runnable-command grammar the benchmark already classifies #3595

Description

@biggs-100

📝 Bug Description

Two grammars in this repository define what counts as a "runnable continuation" for a refusal message, and they disagree today:

  1. The static ratchet (internal/cli/refusal_resolution_ratchet_test.go, refusalRatchetNamedContinuationRegexp, ~line 151) matches only gentle-ai [a-z][a-z-]* — product name plus one verb, nothing after. A refusal printing gentle-ai review status --cwd <repo> passes this check.
  2. The benchmark classifier (bench/classify.go, HasRunnableCommand with placeholderRun = <[^>\s]*>) requires at least one real argument and rejects unfilled <placeholder> tokens. The same message is classified not runnable there.

So a message the benchmark classifies out-of-band satisfies the ratchet that CI runs as a required merge check. The ratchet's own header states the limitation plainly: "it cannot catch a WRONG named continuation — one that parses, runs, and does not help."

This seam is the shared root of a recurring class of reports where printed guidance cannot be executed as printed:

And per #2834's measurement, internal/sddstatus/runtime_ledger.go alone declares 36 refusals with no runnable exit.

🔄 Steps to Reproduce

  1. Observe refusalRatchetNamedContinuationRegexp = regexp.MustCompile(gentle-ai [a-z][a-z-]*) in internal/cli/refusal_resolution_ratchet_test.go.
  2. Evaluate it against a refusal text such as run gentle-ai review status --cwd `` — it matches (gentle-ai review), so the site passes the ratchet.
  3. Run the same text through bench/classify.go HasRunnableCommand: placeholderRun finds <repo> and the classifier returns false.
  4. Result: the two enforcement layers disagree about the same emitted sentence, and CI's required check is the weaker of the two.

✅ Expected Behavior

The static ratchet enforces the same printed-command grammar the benchmark already enforces:

  • a named continuation must carry real arguments, not stop at the verb;
  • an unfilled <placeholder> tail is accepted only when the site explicitly declares operator knowledge via the existing refusal:by-design operator-knowledge: marker (or holds a baseline row);
  • the baseline regenerates once at the switch and may only shrink afterwards;
  • existing CI (already a required merge check) then enforces the stronger grammar going forward — no new CI lane needed.

Both command splitters needed already exist and are deliberately mirrored (internal/cli/review_printed_command.go SplitPrintedCommandWords and its bench twin); no third splitter should be written.

❌ Actual Behavior

The ratchet accepts verb-only matches regardless of what follows, including tails the benchmark classes as not runnable. Wrong or dead-end continuations therefore pass the required merge check until someone files a per-instance report.

🌍 Environment

  • Current main (verified against source on 2026-08-22)
  • Evidence gathered on Windows; the grammar seam itself is platform-independent

💡 Proposed Scope

Static unification slice only: swap the ratchet's continuation grammar to the shared printed-command grammar, add the operator-knowledge marker path, regenerate the baseline once. Estimated ~150–250 changed lines including tests; one revertible PR from current main. The dynamic half ("follow the printed recipe until the block clears") already has a canonical pattern (internal/cli/review_abandon_message_test.go) and grows case-by-case afterwards.

Non-goals: no execution sandbox in the unit lane, no new wire vocabulary, no parallel command splitter, no change to bench classification semantics.

Activity

  1. biggs-100 commented on Aug 23, 2026

    @biggs-100
    Author

    Measured the seam locally to put real numbers behind this issue. A temporary probe test (on a local branch, not proposed for merge) walks the production refusal sites with the standard analyzer, then re-classifies every site the legacy grammar passes using bench-grade runnable-command semantics built on SplitPrintedCommandWords from this same package.

    Results against current main (51a0d60b):

    • 102 sites pass the legacy verb-only grammar as satisfiedNamed
    • 22 of those (21.6%) fail the bench-grade check: no argument beyond the verb, or an unfilled <placeholder> token in command position
    • 1,564 sites remain baseline-covered violations; 0 problems introduced by the probe

    The 22 fall into four categories, each with a different remedy:

    1. ~13 carry placeholders such as --cwd <repo> or --contract <same-contract>. Several of these are fixable by interpolation rather than markers: at least one message prints <same-contract> while the runtime holds the actual contract value in scope a few lines away.
    2. 3 are usage strings (usage: gentle-ai codegraph init --cwd <project-root>) where the placeholder is genuinely operator knowledge; these fit the existing refusal:by-design operator-knowledge: marker path.
    3. 4 build the command through %s at runtime; statically unprovable today, so baseline rows until the analyzer grows that reach.
    4. One is a false positive that refines the proposal: rerun gentle-ai sync is verb-only but genuinely runs as printed. Zero-argument verbs are runnable, so the shared invariant should be "parses against a real verb with valid flag shapes", not "carries at least one argument". Bench's HasRunnableCommand currently has the same bias.

    Two takeaways for whoever picks this up:

    • The cost estimate holds. Roughly 16 marker-or-interpolation fixes, about 4 baseline rows, one grammar refinement for argumentless verbs, plus removing the probe lands inside the 150-250 line forecast in the issue body.
    • The category split suggests sequencing: interpolation fixes first (they shrink the count honestly), then markers for genuine operator-knowledge, then the grammar switch so regressions stay dead in CI.

    Happy to share the branch diff if useful.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions