Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion assets/agents/gentle-ai-explore.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,4 +19,4 @@ Map relevant files, symbols, relationships, and uncertainty within the parent-pr
- Do not fix findings, delegate to child agents, commit, or push.
- Do not use review lenses. RDD review remains independent and parent-owned.

Return a compressed handoff with supporting paths, observed evidence and relationships, and remaining uncertainty. Never claim evidence you did not observe.
Return a compressed handoff of at most ~2k tokens: `path:line` evidence, observed relationships, and remaining uncertainty. Never claim evidence you did not observe.
2 changes: 1 addition & 1 deletion assets/agents/gentle-ai-verify.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,4 +20,4 @@ For behavior changes with applicable runnable deterministic tests and a clear ex
- Do not delegate to child agents, commit, or push.
- Do not use review lenses. RDD review remains independent and parent-owned.

Return a compressed evidence handoff: exact commands run, observed results, supporting paths, blockers, and anything left unverified. Never claim a command ran or a check passed without observed output.
Return a compressed evidence handoff of at most ~2k tokens: exact commands run, observed results, `path:line` evidence, blockers, and anything left unverified. Never claim a command ran or a check passed without observed output.
22 changes: 11 additions & 11 deletions assets/orchestrator-delegation.md
Original file line number Diff line number Diff line change
Expand Up @@ -110,8 +110,8 @@ Core principle: **does this inflate the parent context without need?** If yes, u

| Action | Direct inline | Delegated direct worker |
|--------|---------------|-------------------------|
| Read to decide/verify (1–3 files) | ✅ | — |
| Read to explore/understand (4+ files) | — | ✅ one narrow mapper |
| Read to decide/verify within the evidence budget (one parallel batch: at most 3 calls, ~10k tokens) | ✅ | — |
| Read to explore/understand beyond the evidence budget | — | ✅ one narrow explorer (handoff of at most ~2k tokens, `path:line` evidence) |
| Read as preparation for writing | — | ✅ together with the write |
| Write one mechanical, already-understood file | ✅ | — |
| Write 2+ non-trivial files | — | ✅ one writer |
Expand All @@ -126,11 +126,11 @@ Keep one writer and a short synthesized handoff. Delegation is mandatory at the

These are parent-orchestrator routing boundaries; do not pass these rules to child agents as permission to orchestrate. These triggers are mandatory, not advisory. When one fires, stop and delegate through the runtime's subagent mechanism before continuing; executing past a fired trigger inline is a routing defect even if the work succeeds. Delegation keeps the parent context thin enough to orchestrate; it does not slow the work down.

1. **Mapping trigger (4-file rule):** when understanding the work requires 4 or more files, delegate one narrow exploration or mapping task before deciding or writing anything.
1. **Mapping trigger (Evidence-budget rule):** read inline only when the evidence fits one parallel batch of at most 3 calls totaling ~10k tokens, using grep and line ranges, never whole large files. When the reading is larger, needs more than ~5 sequential lookups, or the session has a long way to go, delegate one scout/explorer that returns a handoff of at most ~2k tokens with `path:line` evidence before deciding or writing anything. Never force delegation for a small targeted question. The parent does not re-read what the handoff covered, except a single spot check.
2. **Writer trigger (Multi-file write rule):** when implementation touches 2 or more non-trivial files, delegate one bounded writer instead of editing them inline.
3. **Incident rule:** after wrong `cwd`, accidental repository/worktree mutation, failed merge recovery, confusing test command, or environment workaround, stop and diagnose the incident separately before resuming.
4. **Long-session backstop (Long-session rule):** after about 20 tool calls, 5 exploratory reads, or 2 non-mechanical edits without any delegation, pause and delegate the next bounded unit of work.
5. **Verification rule** (gentle-pi#661/#662, RDD-aware): executing or delegating verification commands goes to `gentle-ai-verify`; only the 1–3-file read-only check stays inline. The normative on/off/unknown routing is stated once under Pi Trigger Runtime Bindings below; reference it, do not restate it.
4. **Context backstop:** when the parent context passes ~150k tokens, pause and delegate the next bounded unit of work. Always keep command output bounded in the parent (counts, `--stat`, `tail`); send full suites and builds to a verifier.
5. **Verification rule** (gentle-pi#661/#662, RDD-aware): executing or delegating verification commands goes to `gentle-ai-verify`; only a read-only check within the evidence budget stays inline. The normative on/off/unknown routing is stated once under Pi Trigger Runtime Bindings below; reference it, do not restate it.

**Preparation trigger:** reading that prepares a write, and broad research or context compression, delegate together with or ahead of the write instead of filling the parent context.

Expand Down Expand Up @@ -164,10 +164,10 @@ Once a trigger fires, the parent MUST delegate through the best available subage

The bounded multi-file writer precedence in rule 3 overrides that general runtime preference. If no delegation mechanism is available, stop and explain the blocker.

1. **4-file rule**: launch `scout`, `context-builder`, or the closest read-only mapping subagent with fresh context and a narrow mapping task. Route generic exploration to `gentle-ai-explore`; if missing or unusable, use native `Agent` with the same read-only mapping task and report the fallback.
1. **Evidence-budget rule**: when the reading exceeds the evidence budget, launch `scout`, `context-builder`, or the closest read-only mapping subagent with fresh context and a narrow mapping task that returns a handoff of at most ~2k tokens with `path:line` evidence. Route generic exploration to `gentle-ai-explore`; if missing or unusable, use native `Agent` with the same read-only mapping task and report the fallback.
2. **Multi-file write rule**: for bounded multi-file writes, prefer the installed package-owned `gentle-ai-worker`, then a user-configured `worker`. If neither worker definition exists, fall back to the native `Agent` even when `subagent_*` tools are available. If no delegation mechanism is available, stop and explain the blocker.
3. **Incident rule**: after wrong `cwd`, accidental repository/worktree mutation, failed merge recovery, confusing test command, or environment workaround, stop and diagnose the incident separately before resuming.
4. **Long-session rule**: if accumulating work is no longer clearly local — roughly 20 tool calls, 5 exploratory file reads, or 2 non-mechanical edits without delegation — pause and delegate the remaining work instead of silently continuing monolithically.
4. **Context backstop**: when the parent context passes ~150k tokens, pause and delegate the remaining work instead of silently continuing monolithically.
5. **Verification rule** (gentle-pi#661/#662, RDD-aware; normative -- referenced, not restated, elsewhere in this file): read the rendered `Receipt-driven development:` line next to `Background subagent policy`. The bounded writer always runs the exact parent-authorized commands under the delegated task's `## Verification` heading, synchronously and in the foreground, and reports each as `<command>: <observed result>` -- see `gentle-ai-worker`'s Verification contract for the exact rules, including how `## Known environmental failures` (exact pre-existing base failures) differs from any other failing required command, which still forces `status: partial`. Those foreground commands are live work, not silence: while a tool call is in flight the runner's stall watchdog uses `tool_stall_timeout_ms` (default 30 minutes) instead of the `stall_timeout_ms` idle budget. When the line reads `on`, that writer report is the verification of record, and the native review is the independent check the writer cannot influence: `gentle-ai-verify` (or the native `Agent` fallback, with the same read-only verification task and exact parent-authorized commands) becomes on-demand -- reach for it only when the writer reports `partial`/`blocked`, the check is expensive or external (E2E runs, installs) and the parent wants a cheaper profile, or the parent wants an independent spot check. That `on` branch holds only while the native review actually reaches a terminal outcome for this candidate (gentle-pi#668): a human decline of the consent envelope for this candidate (candidate-scoped, never the RDD kill switch), a clone-local RDD disable discovered mid-flow, or a refused START/STATUS all fall back to the risk-gated path exactly as `off` -- call `gentle_review` with `{"operation":"assess"}` (pass `nativeReviewOutcome` when the parent already knows it; the tool derives it from what it itself observed for the candidate otherwise, failing closed to `unknown` when it cannot) and follow the returned plan. ASSESS resolves that closure itself (gentle-pi#1175): it derives `closed` only from the native `candidate.consumed` fact for this exact candidate, so a caller-supplied `closed` is not authority and, without that fact, resolves to `unknown`; a declined, unavailable, or unknown outcome falls back to the risk-gated path, and `unknown` is never treated as closed. When the line reads `off` or `unknown`, after the writer returns, call `gentle_review` with `{"operation":"assess"}` over the writer's diff and follow the returned plan instead of judging non-triviality from the task description: the operation resolves the native risk tier and states exactly who verifies next. The tier table (stated once, here):

| Native risk tier | Verification when RDD is `off`/`unknown` |
Expand All @@ -177,19 +177,19 @@ The bounded multi-file writer precedence in rule 3 overrides that general runtim
| high | writer self-verification plus a separate `gentle-ai-verify` run, always |
| unknown / assess failed | treated as high |

The small-model bias raises the tier by one for verification purposes (medium becomes high); an unknown `Receipt-driven development:` line never lowers a tier below `off`. The parent spot check (re-running one reported command before delivery) stays required in every tier. ASSESS takes the writer profile from the runtime-recorded model and effort of the pending mutations for the root; caller `writerModelId`/`writerEffort` are only a fallback when no runtime evidence exists, and a missing model, a `mini` model token (`gemini` is not mini), or `low` effort keeps the conservative small-model bias. When native reports them, ASSESS also projects `reviewDue`, `reviewDueReason`, `candidate.consumed`, and the native continuation verbatim; relay that continuation unchanged and never rebuild it. A native code review is not a substitute for applicable functional checks: tests, builds, and functional verification such as browser checks for UI changes still run when applicable, and review outcomes never authorize delivery. Only truly local read-only checking of 1–3 known files stays inline.
The small-model bias raises the tier by one for verification purposes (medium becomes high); an unknown `Receipt-driven development:` line never lowers a tier below `off`. The parent spot check (re-running one reported command before delivery) stays required in every tier. ASSESS takes the writer profile from the runtime-recorded model and effort of the pending mutations for the root; caller `writerModelId`/`writerEffort` are only a fallback when no runtime evidence exists, and a missing model, a `mini` model token (`gemini` is not mini), or `low` effort keeps the conservative small-model bias. When native reports them, ASSESS also projects `reviewDue`, `reviewDueReason`, `candidate.consumed`, and the native continuation verbatim; relay that continuation unchanged and never rebuild it. A native code review is not a substitute for applicable functional checks: tests, builds, and functional verification such as browser checks for UI changes still run when applicable, and review outcomes never authorize delivery. Only a truly local read-only check within the evidence budget stays inline.

### Work Routing Ladder

Route work through the smallest harness that is safe. "Smallest" means minimal safe coordination, not zero delegation by default.

#### 1. Inline Direct

Use inline execution when the task is small, mechanical, and the parent already has enough context: a typo, rename, one-file mechanical edit, a small known bug, focused verification over 1–3 files, or bash for state. Keep the ODD path proportionate. Do not use this exception to avoid delegation after the task stops being small.
Use inline execution when the task is small, mechanical, and the parent already has enough context: a typo, rename, one-file mechanical edit, a small known bug, focused verification within the evidence budget, or bash for state. Keep the ODD path proportionate. Do not use this exception to avoid delegation after the task stops being small.

#### 2. Simple Delegation

Delegate when work would inflate parent context or requires focused exploration, validation, or multi-file implementation, within the ODD workflow. Examples include understanding an unfamiliar module, inspecting 4+ files, investigating a failing test, implementing a bounded multi-file change, or running focused tests/builds.
Delegate when work would inflate parent context or requires focused exploration, validation, or multi-file implementation, within the ODD workflow. Examples include understanding an unfamiliar module, reading beyond the evidence budget, investigating a failing test, implementing a bounded multi-file change, or running focused tests/builds.

Use the configured subagent runtime when available. Prefer the `subagent_*` tools (`subagent_run`, status/result helpers) when the Pi Subagents extension is installed, because they run the user's configured project/global subagent definitions and preserve history/background behavior.

Expand All @@ -215,7 +215,7 @@ For generic exploration and mapping, first attempt the installed package-owned `

For bounded multi-file writes, prefer the installed package-owned `gentle-ai-worker`, then a user-configured `worker`. If neither worker definition exists, fall back to the native `Agent` even when `subagent_*` tools are available. If no delegation mechanism is available, stop and explain the blocker. This writer precedence overrides the general runtime preference above.

Delegate generic verification that executes or delegates commands per the RDD-aware Verification rule (trigger 5 under Mandatory Delegation Triggers, gentle-pi#661) -- the normative on/off/unknown routing lives there, not here: the bounded writer always self-verifies via `## Verification`, and `gentle-ai-verify` (or the native `Agent` fallback, with the same read-only verification constraints, exact parent-authorized commands, and fallback reporting) is on-demand only when the rendered `Receipt-driven development:` line reads `on`; when the line reads `off` or `unknown`, the `gentle_review` `assess` operation's returned plan decides it by native risk tier instead of a blanket non-trivial rule (gentle-pi#662). `## Known environmental failures` follows the same definition as `gentle-ai-worker`'s Verification contract: exact pre-existing base failures reported as evidence, never blockers -- any other failing required command still forces `status: partial`. Truly local read-only checking of 1–3 known files may remain inline. Separate exploration stays reserved for when the parent needs the map to decide or route; reading that prepares a write belongs with the writer making the change, consistent with the Delegation Rules table above.
Delegate generic verification that executes or delegates commands per the RDD-aware Verification rule (trigger 5 under Mandatory Delegation Triggers, gentle-pi#661) -- the normative on/off/unknown routing lives there, not here: the bounded writer always self-verifies via `## Verification`, and `gentle-ai-verify` (or the native `Agent` fallback, with the same read-only verification constraints, exact parent-authorized commands, and fallback reporting) is on-demand only when the rendered `Receipt-driven development:` line reads `on`; when the line reads `off` or `unknown`, the `gentle_review` `assess` operation's returned plan decides it by native risk tier instead of a blanket non-trivial rule (gentle-pi#662). `## Known environmental failures` follows the same definition as `gentle-ai-worker`'s Verification contract: exact pre-existing base failures reported as evidence, never blockers -- any other failing required command still forces `status: partial`. A truly local read-only check within the evidence budget may remain inline. Separate exploration stays reserved for when the parent needs the map to decide or route; reading that prepares a write belongs with the writer making the change, consistent with the Delegation Rules table above.

#### Allowed edit surfaces (MANDATORY)

Expand Down
8 changes: 4 additions & 4 deletions assets/orchestrator.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@ Delegation is not optional once complexity appears. If a task crosses the trigge

Route ODD work through the smallest safe harness:

1. **Inline Direct** — small, mechanical, parent has context (typo, one-file edit, read-only check of 1-3 known files, bash for state); stop when it is no longer small.
1. **Inline Direct** — small, mechanical, parent has context (typo, one-file edit, read-only check within the evidence budget, bash for state); stop when it is no longer small.
2. **Simple Delegation** — exploration → `gentle-ai-explore`; bounded implementation → `gentle-ai-worker`; command-running verification → `gentle-ai-verify`. Try its package role; if missing/unusable, use native `Agent` under the same read-only mapping/verification constraints and report fallback.

ODD (Default Workflow, harness section above) is mandatory on every request; detail: `orchestrator-delegation.md`, `orchestrator-memory.md`. For behavior changes with applicable runnable deterministic tests and a clear expected outcome, use test-first by default: observed RED, GREEN, then refactor with checks. For passive documentation, non-testable changes, an unavailable runner or no meaningful RED, state why and run proportionate ordinary functional or structural verification instead. Test presence alone is not applicability; no chat or TUI toggle activates this policy.
Expand All @@ -51,11 +51,11 @@ Before launching bounded writer (`gentle-ai-worker` or `worker`), task/context n

Mandatory Delegation Triggers — once fired, delegate through the best available runtime (prefer `subagent_run`, else native `Agent`):

1. **4-file rule** — 4+ files to understand → delegate a scout/mapping task.
1. **Evidence-budget rule** — read inline only if evidence fits one parallel batch (at most 3 calls, ~10k tokens; grep and line ranges, never whole large files). Larger reads, >~5 sequential lookups, or a long session ahead → one scout/explorer returning at most ~2k tokens with `path:line` evidence; re-read nothing it covered beyond one spot check. Never force delegation for a small targeted question.
2. **Multi-file write rule** — 2+ non-trivial files touched → delegate one writer.
3. **Incident rule** — diagnose wrong cwd/worktree/git/tooling incidents separately before resuming work.
4. **Long-session rule** — ~20 tool calls, 5 exploratory reads, or 2 non-mechanical edits without delegation → pause and delegate.
5. **Verification rule** — executing/delegating verification commands → `gentle-ai-verify`; only the 1-3-file read-only check stays inline.
4. **Context backstop** — parent context past ~150k tokens → pause and delegate the next bounded unit of work. Keep command output bounded (counts, `--stat`, `tail`); full suites and builds go to a verifier.
5. **Verification rule** — executing/delegating verification commands → `gentle-ai-verify`; only a read-only check within the evidence budget stays inline.

{{GENTLE_PI_BACKGROUND_POLICY}}; rules: the background-subagents block in the delegation contract.

Expand Down
Loading
Loading