From 302d0950d49ee1611933f60a0f186544130cb3f9 Mon Sep 17 00:00:00 2001 From: Alan Buscaglia Date: Wed, 30 Sep 2026 16:57:15 +0200 Subject: [PATCH 1/5] feat(agents): keep orchestrator-only context blocks out of subagents Delegated children now receive a minimal child-context extension through the existing extensionPaths mechanism. It filters the gentle-ai managed orchestrator, sdd-orchestrator, sdd-model-assignments, and agent-routing blocks out of the child's context files, keeps nested non-orchestrator blocks and all project text, and leaves a file unchanged when its markers are malformed. Measured live: a delegated explorer starts at ~57k instead of ~87k tokens. Refs #1587 --- docs/gentle-shell.md | 2 + extensions/child-context.ts | 26 +++ extensions/gentle-agents.ts | 18 +++ lib/child-context-files.ts | 166 +++++++++++++++++++ tests/child-context-files.test.ts | 255 ++++++++++++++++++++++++++++++ tests/gentle-agents.test.ts | 31 +++- 6 files changed, 496 insertions(+), 2 deletions(-) create mode 100644 extensions/child-context.ts create mode 100644 lib/child-context-files.ts create mode 100644 tests/child-context-files.test.ts diff --git a/docs/gentle-shell.md b/docs/gentle-shell.md index 34c491d13..839a07fe2 100644 --- a/docs/gentle-shell.md +++ b/docs/gentle-shell.md @@ -189,6 +189,8 @@ The card is above the editor in every mode, including fullscreen — it is not o Every subagent is its own `pi --mode rpc` child process, so the terminal never runs subagent work: the host reads JSON lines, applies each one as a small delta to a bounded per-task thread, and notifies only the listeners of that task. A task-mode child's question (`ctx.ui.select`, `confirm`, `input`, `editor`) reaches you as an ordinary pi dialog; a background child's question is dismissed. Subagents have no automatic total execution timeout: a long-running child remains live while it continues emitting RPC events. A silent child still times out through the configurable `stall_timeout_ms` watchdog (default four minutes). An announced tool call that is still running is live work, not silence, so it is bounded by `tool_stall_timeout_ms` instead (default 30 minutes, never below `stall_timeout_ms`). Closing pi stops the children that are still running. +Delegated children (`GENTLE_PI_AGENTS_CHILD=1`) load the same context files as their parent, minus the gentle-ai managed blocks that bind themselves to the orchestrator, such as `orchestrator` and `agent-routing` (the full list lives in [`lib/child-context-files.ts`](../lib/child-context-files.ts)). Every other managed block (for example `codegraph-guidance`, `engram-protocol`, or `remote-authorization` nested inside `agent-routing`) and all unmanaged project text reach the child unchanged. A marker counts only when it is alone on its line outside fenced code; if a file's markers are unbalanced, mismatched, or ambiguous, that file is passed through unfiltered. Gentle Agents launches every child with `--extension` pointing at the packaged `extensions/child-context.ts`, because a child does not load the gentle-pi package on its own in the isolated Gentle Shell home; if that file is missing, the child starts without it and keeps today's unfiltered context. The extension registers only a `before_agent_start` hook and does nothing outside a child session, so primary sessions always receive their context files as-is. + - `subagent_list_agents`, `subagent_run` (`agent`, `task`, `label?`, `context?`, `workspace_root?` or `repository_root?`, `mode?` task or background), `subagent_status`, `subagent_result`, `subagent_list_tasks`, `subagent_reply` (one current-session reply to a live child query), `subagent_cancel`, `subagent_send_message` (steer a running child), `subagent_continue` (resume a finished task in its own session). - `orchestrator_session_id`, `orchestrator_list`, and `orchestrator_send_message` provide local-profile session notifications. List results advertise IDs only and reachability remains unknown. Sending selects the sole advertised peer or asks the user to choose; outbound messages require explicit interactive human consent before dispatch (Allow once, Allow for this session, or Deny), and fail closed if interactive UI or its selector is unavailable, even after a session grant. Each send requires a caller-supplied reason with at least 8 characters after trimming and at most 512 UTF-8 bytes. The consent preview visibly escapes controls and Unicode line separators; the delivered message remains unchanged. A successful ACK means the peer accepted the notification for delivery, not that it read or completed work. This is notification-and-ACK transport only: it has no cross-session queries, offline queue, retries, broadcasts, or read/completion guarantees. On Unix, presence records remain in the profile's private transport directory while socket endpoints use a private, profile-hashed directory below the canonical system temporary directory, keeping endpoint length independent of the profile path and at most 100 encoded bytes. The shared system temporary parent is only validated (current-user-owned without group/other write, or root/current-user-owned, world-writable, and sticky); it is never claimed, permissioned, or cleaned up by gentle-pi. On POSIX, the transport uses private Unix-domain sockets; on Windows, it uses private named pipes scoped by the current account SID, served by a package-local PowerShell helper (`runtime/windows-session-transport.ps1`): the transport selects that fixed helper, and availability and delivery depend on the helper's bounded startup and pipe checks. Notification and ACK limits remain bounded across platforms. - `subagent_run.workspace_root` selects the parent's main worktree or an existing linked worktree in the parent's Git clone only. Validation happens before queueing; the child runs at that canonical root. Successful OS spawn registers it in the originating parent's same-clone registry, including delayed queued launches; failed spawns do not register. diff --git a/extensions/child-context.ts b/extensions/child-context.ts new file mode 100644 index 000000000..1c2f50136 --- /dev/null +++ b/extensions/child-context.ts @@ -0,0 +1,26 @@ +import type { ExtensionAPI } from "@earendil-works/pi-coding-agent"; +import { filterChildSessionContextFiles, type ContextFileOptions } from "../lib/child-context-files.ts"; + +// gentle-shell#1587: delegated children drop the orchestrator-only gentle-ai +// managed blocks from their context files. Gentle Agents passes this file to +// every child with --extension because children do not load the gentle-pi +// package in the isolated Gentle Shell home. It registers one hook and has no +// other side effects; outside a child session it is a no-op. +export function createChildContextExtension(env: NodeJS.ProcessEnv = process.env): (pi: ExtensionAPI) => void { + return (pi) => { + pi.on("before_agent_start", (event) => { + if (env.GENTLE_PI_AGENTS_CHILD !== "1") return undefined; + // The filtered copies replace contextFiles on the same options object + // (pi-claude-bridge rebuilds its prompt from it). This never throws, + // keeps the original files on any error or malformed markers, and is + // idempotent: already-filtered content has nothing left to remove. + const options = (event as { systemPromptOptions?: ContextFileOptions | null } | undefined)?.systemPromptOptions; + filterChildSessionContextFiles(options); + return undefined; + }); + }; +} + +export default function childContextExtension(pi: ExtensionAPI): void { + createChildContextExtension()(pi); +} diff --git a/extensions/gentle-agents.ts b/extensions/gentle-agents.ts index 57de7761d..4528cd2fb 100644 --- a/extensions/gentle-agents.ts +++ b/extensions/gentle-agents.ts @@ -11,6 +11,7 @@ import { randomUUID } from "node:crypto"; import { mkdir, readFile, writeFile } from "node:fs/promises"; import os from "node:os"; import { join, resolve, isAbsolute } from "node:path"; +import { fileURLToPath } from "node:url"; import { keyHint, type ExtensionAPI, type ExtensionContext } from "@earendil-works/pi-coding-agent"; import { Text, type TUI } from "@earendil-works/pi-tui"; import { invalidateSidebar } from "../lib/shell-sidebar-layout.ts"; @@ -96,6 +97,21 @@ export interface AgentsDeps extends RunnerDeps { runtimeMetricsPolicy?: RuntimeMetricsPolicyDeps; metricsNow?: () => number; metricsSchedule?: RunnerDeps["schedule"]; + // Extensions every child loads with --extension (gentle-shell#1587). + childExtensionPaths?: string[]; +} + +// gentle-shell#1587: children do not load the gentle-pi package in the +// isolated Gentle Shell home, so the child-context extension (which drops the +// orchestrator-only managed blocks from their context files) is passed to +// every child explicitly. A missing file fails safe to no extension. +export function childContextExtensionPaths(exists: (path: string) => boolean = existsSync): string[] { + try { + const path = fileURLToPath(new URL("./child-context.ts", import.meta.url)); + return exists(path) ? [path] : []; + } catch { + return []; + } } export function agentRuntimePaths(home: string, agentHome = join(home, ".pi", "agent")): { sessions: string; transcripts: string } { @@ -122,6 +138,7 @@ const defaultDeps = (env: NodeJS.ProcessEnv): AgentsDeps => ({ home: os.homedir(), resolveWorktree: resolveSessionWorktree, env, + childExtensionPaths: childContextExtensionPaths(), }); export function agentsEnabled(env: NodeJS.ProcessEnv = process.env): boolean { @@ -1106,6 +1123,7 @@ export default function gentleAgents(pi: ExtensionAPI, env: NodeJS.ProcessEnv = thinking: profile.thinking, sessionDir, resumeSessionPath: resume, + ...(deps.childExtensionPaths && deps.childExtensionPaths.length > 0 ? { extensionPaths: [...deps.childExtensionPaths] } : {}), env: childEnv, ...(foreign || parentRepositoryIdentity === undefined ? {} : { authorizeParentStandingReviewPermission: (repositoryIdentity: string) => { diff --git a/lib/child-context-files.ts b/lib/child-context-files.ts new file mode 100644 index 000000000..9935c8871 --- /dev/null +++ b/lib/child-context-files.ts @@ -0,0 +1,166 @@ +// gentle-shell#1587: delegated children (GENTLE_PI_AGENTS_CHILD=1) load the +// same context files as the orchestrator, including gentle-ai managed blocks +// that bind themselves to the orchestrator only. Those blocks cost ~37k prefix +// tokens per child and give a worker orchestration rules it must not follow. +// This module removes exactly those blocks and keeps everything else. + +export const ORCHESTRATOR_ONLY_MANAGED_BLOCKS = [ + "orchestrator", + "sdd-orchestrator", + "sdd-model-assignments", + "agent-routing", +] as const; + +export type OrchestratorOnlyManagedBlock = (typeof ORCHESTRATOR_ONLY_MANAGED_BLOCKS)[number]; + +export interface ContextFile { + path: string; + content: string; +} + +export interface ManagedBlockFilterResult { + content: string; + removedBlocks: number; + // True when markers were unbalanced, mismatched, or ambiguous and the + // content was returned unchanged. + failSafe: boolean; +} + +export interface ChildContextFilesResult { + files: T[]; + removedBlocks: number; + removedBytes: number; + failSafePaths: string[]; +} + +export interface ContextFileOptions { + contextFiles?: ContextFile[]; +} + +const REMOVED_NAMES: ReadonlySet = new Set(ORCHESTRATOR_ONLY_MANAGED_BLOCKS); + +// A managed marker counts only when it is alone on its line. Inline mentions +// (prose, inline code) and markers inside fenced code are plain text. +const MARKER_LINE = /^[ \t]*[ \t]*$/; +const FENCE_LINE = /^ {0,3}(`{3,}|~{3,})(.*)$/; +const BLANK_LINE = /^[ \t]*\r?\n?$/; + +interface Line { + text: string; + // Name of the innermost managed block owning this line (marker lines are + // owned by the block they delimit), or null for unmanaged text. + owner: string | null; +} + +function splitLines(content: string): string[] { + return content.match(/[^\n]*\n|[^\n]+$/g) ?? []; +} + +function unchanged(content: string): ManagedBlockFilterResult { + return { content, removedBlocks: 0, failSafe: true }; +} + +export function filterOrchestratorOnlyBlocks(content: string): ManagedBlockFilterResult { + const lines: Line[] = []; + const stack: string[] = []; + let fence: { char: string; length: number } | null = null; + let removedBlocks = 0; + let sawMarker = false; + + for (const text of splitLines(content)) { + const bare = text.replace(/\r?\n$/, ""); + const owner = stack.length > 0 ? stack[stack.length - 1] : null; + const fenceMatch = FENCE_LINE.exec(bare); + if (fence !== null) { + if (fenceMatch !== null && fenceMatch[1][0] === fence.char && fenceMatch[1].length >= fence.length && fenceMatch[2].trim() === "") { + fence = null; + } + lines.push({ text, owner }); + continue; + } + if (fenceMatch !== null && !(fenceMatch[1][0] === "`" && fenceMatch[2].includes("`"))) { + fence = { char: fenceMatch[1][0], length: fenceMatch[1].length }; + lines.push({ text, owner }); + continue; + } + const marker = MARKER_LINE.exec(bare); + if (marker === null) { + lines.push({ text, owner }); + continue; + } + sawMarker = true; + const [, closing, name] = marker; + if (closing === "") { + // A block reopened inside itself is ambiguous: fail safe. + if (stack.includes(name)) return unchanged(content); + stack.push(name); + if (REMOVED_NAMES.has(name)) removedBlocks += 1; + lines.push({ text, owner: name }); + continue; + } + if (stack.length === 0 || stack[stack.length - 1] !== name) return unchanged(content); + stack.pop(); + lines.push({ text, owner: name }); + } + if (stack.length > 0) return unchanged(content); + if (!sawMarker || removedBlocks === 0) return { content, removedBlocks: 0, failSafe: false }; + + // Drop lines owned by a removed block. Blank lines that directly follow a + // removal collapse so at most one empty line remains at that seam; a + // removal at the start or end of the file leaves no blank edge behind. + const output: string[] = []; + let trailingBlanks = 0; + let afterRemoval = false; + for (const line of lines) { + if (line.owner !== null && REMOVED_NAMES.has(line.owner)) { + afterRemoval = true; + continue; + } + const blank = BLANK_LINE.test(line.text); + if (afterRemoval && blank && (output.length === 0 || trailingBlanks >= 1)) continue; + if (!blank) afterRemoval = false; + output.push(line.text); + trailingBlanks = blank ? trailingBlanks + 1 : 0; + } + if (afterRemoval) { + while (output.length > 0 && BLANK_LINE.test(output[output.length - 1])) output.pop(); + } + return { content: output.join(""), removedBlocks, failSafe: false }; +} + +export function stripOrchestratorOnlyBlocks(content: string): string { + return filterOrchestratorOnlyBlocks(content).content; +} + +const encoder = new TextEncoder(); + +// Returns filtered copies; the input array and its entries are never mutated. +export function filterChildContextFiles(files: readonly T[]): ChildContextFilesResult { + const result: ChildContextFilesResult = { files: [], removedBlocks: 0, removedBytes: 0, failSafePaths: [] }; + for (const file of files) { + const filtered = filterOrchestratorOnlyBlocks(file.content); + if (filtered.failSafe) result.failSafePaths.push(file.path); + if (filtered.removedBlocks === 0) { + result.files.push({ ...file }); + continue; + } + result.removedBlocks += filtered.removedBlocks; + result.removedBytes += encoder.encode(file.content).length - encoder.encode(filtered.content).length; + result.files.push({ ...file, content: filtered.content }); + } + return result; +} + +// Replaces contextFiles on the SAME options object: pi-claude-bridge keeps a +// reference to it and rebuilds its prompt from its contextFiles. Never +// throws; on any error the original context files stay in place. +export function filterChildSessionContextFiles(options: ContextFileOptions | null | undefined): ChildContextFilesResult | null { + try { + if (!options || !Array.isArray(options.contextFiles)) return null; + const result = filterChildContextFiles(options.contextFiles); + if (result.removedBlocks > 0) options.contextFiles = result.files; + return result; + } catch { + return null; + } +} diff --git a/tests/child-context-files.test.ts b/tests/child-context-files.test.ts new file mode 100644 index 000000000..1fa5b76b3 --- /dev/null +++ b/tests/child-context-files.test.ts @@ -0,0 +1,255 @@ +import assert from "node:assert/strict"; +import test from "node:test"; +import { + filterChildContextFiles, + filterChildSessionContextFiles, + ORCHESTRATOR_ONLY_MANAGED_BLOCKS, + stripOrchestratorOnlyBlocks, +} from "../lib/child-context-files.ts"; +import type { ExtensionAPI } from "@earendil-works/pi-coding-agent"; +import childContextExtension, { createChildContextExtension } from "../extensions/child-context.ts"; + +function block(name: string, body: string): string { + return `\n${body}\n`; +} + +// Synthetic file shaped like a gentle-ai managed AGENTS.md: unmanaged project +// text, kept blocks, orchestrator-only blocks, a nested kept block inside +// agent-routing, and a nested orchestrator-only block inside sdd-orchestrator. +const REALISTIC_AGENTS_MD = [ + "# Project conventions", + "", + "Run `pnpm test` before pushing.", + "", + block("codegraph-guidance", "## CodeGraph\n\nUse CodeGraph first."), + "", + block("sdd-orchestrator", [ + "", + "# Orchestrator instructions", + "You are a COORDINATOR.", + "", + "", + block("sdd-model-assignments", "| Phase | Model |\n|---|---|\n| sdd-apply | sonnet |"), + "", + "### Sub-agent launch pattern", + ].join("\n")), + "", + block("engram-protocol", "## Engram\n\nSave decisions."), + "", + block("agent-routing", [ + "## Implementation Routing", + "", + "Delegate on the 4-file rule.", + "", + block("remote-authorization", "## Remote operation authorization\n\nAsk before remote work."), + ].join("\n")), + "", + "## Local notes", + "Keep lines short.", + "", +].join("\n"); + +const EXPECTED_REALISTIC = [ + "# Project conventions", + "", + "Run `pnpm test` before pushing.", + "", + block("codegraph-guidance", "## CodeGraph\n\nUse CodeGraph first."), + "", + block("engram-protocol", "## Engram\n\nSave decisions."), + "", + block("remote-authorization", "## Remote operation authorization\n\nAsk before remote work."), + "", + "## Local notes", + "Keep lines short.", + "", +].join("\n"); + +test("the orchestrator-only list names exactly the blocks bound to the orchestrator", () => { + assert.deepEqual([...ORCHESTRATOR_ONLY_MANAGED_BLOCKS].sort(), ["agent-routing", "orchestrator", "sdd-model-assignments", "sdd-orchestrator"]); +}); + +for (const name of ORCHESTRATOR_ONLY_MANAGED_BLOCKS) { + test(`removes the ${name} block and keeps the surrounding text`, () => { + const input = `Before\n\n${block(name, "Orchestrator only\n\nMore text")}\n\nAfter\n`; + assert.equal(stripOrchestratorOnlyBlocks(input), "Before\n\nAfter\n"); + }); +} + +test("filters a realistic managed AGENTS.md: nested kept blocks survive, nested listed blocks go with their parent", () => { + const output = stripOrchestratorOnlyBlocks(REALISTIC_AGENTS_MD); + assert.equal(output, EXPECTED_REALISTIC); + assert.doesNotMatch(output, /COORDINATOR|sdd-model-assignments|Implementation Routing|4-file rule/); + assert.match(output, /\n## Remote operation authorization/); +}); + +test("keeps codegraph-guidance, engram-protocol, and unknown managed blocks byte-for-byte", () => { + const input = [ + block("codegraph-guidance", "cg"), + block("engram-protocol", "engram\n\n\n\nspacing kept"), + block("some-future-block", "future"), + "", + ].join("\n"); + assert.equal(stripOrchestratorOnlyBlocks(input), input); +}); + +test("unmanaged text before, between, and after removed blocks is preserved byte-for-byte", () => { + const input = [ + " leading indent\t", + "", + "", + "double blank kept above", + block("orchestrator", "gone"), + "between text ", + block("agent-routing", "gone"), + "after\ttext", + ].join("\n"); + assert.equal(stripOrchestratorOnlyBlocks(input), " leading indent\t\n\n\ndouble blank kept above\nbetween text \nafter\ttext"); +}); + +test("collapses only blank lines left by a removal to at most one", () => { + const input = `Intro\n\n\n${block("orchestrator", "x")}\n\n\n\nOutro\n\n\n\nTail\n`; + assert.equal(stripOrchestratorOnlyBlocks(input), "Intro\n\n\nOutro\n\n\n\nTail\n"); + const leading = `${block("orchestrator", "x")}\n\n\nFirst line\n`; + assert.equal(stripOrchestratorOnlyBlocks(leading), "First line\n"); + const trailing = `Last line\n\n${block("agent-routing", "x")}\n\n`; + assert.equal(stripOrchestratorOnlyBlocks(trailing), "Last line\n"); +}); + +test("non gentle-ai HTML comments are plain text", () => { + const input = "\nkeep\n\n\n"; + assert.equal(stripOrchestratorOnlyBlocks(input), input); +}); + +test("markers mentioned inline or inside fenced code are plain text", () => { + const input = [ + "Managed blocks start with `` in prose.", + "```markdown", + "", + "example", + "", + "```", + "", + ].join("\n"); + assert.equal(stripOrchestratorOnlyBlocks(input), input); +}); + +test("unbalanced, mismatched, or duplicate-open markers leave the file unchanged", () => { + const cases = [ + `${block("codegraph-guidance", "cg")}\n\nno close\n`, + `\n${block("agent-routing", "x")}\n`, + `\nx\n\n`, + `\n\nx\n\n\n`, + `\n\nx\n\n`, + `\n\nx\n\n\n`, + ]; + for (const input of cases) assert.equal(stripOrchestratorOnlyBlocks(input), input); +}); + +test("handles CRLF input and keeps CRLF line endings", () => { + const input = REALISTIC_AGENTS_MD.replace(/\n/g, "\r\n"); + assert.equal(stripOrchestratorOnlyBlocks(input), EXPECTED_REALISTIC.replace(/\n/g, "\r\n")); +}); + +test("a file without managed markers is returned identical", () => { + const input = "# Plain\n\nNo managed blocks here.\n\n"; + assert.equal(stripOrchestratorOnlyBlocks(input), input); + assert.equal(stripOrchestratorOnlyBlocks(""), ""); +}); + +test("filterChildContextFiles returns filtered copies, never mutates inputs, and reports removals", () => { + const files = Object.freeze([ + Object.freeze({ path: "/home/AGENTS.md", content: REALISTIC_AGENTS_MD }), + Object.freeze({ path: "/repo/AGENTS.md", content: "# Plain\n" }), + Object.freeze({ path: "/repo/sub/AGENTS.md", content: "\nunbalanced\n" }), + ]); + const result = filterChildContextFiles(files); + assert.notEqual(result.files, files); + assert.equal(result.files.length, 3); + assert.equal(result.files[0].path, "/home/AGENTS.md"); + assert.equal(result.files[0].content, EXPECTED_REALISTIC); + assert.notEqual(result.files[0], files[0]); + assert.equal(result.files[1].content, "# Plain\n"); + assert.equal(result.files[2].content, files[2].content); + assert.equal(files[0].content, REALISTIC_AGENTS_MD); + assert.equal(result.removedBlocks, 3, "sdd-orchestrator, its nested sdd-model-assignments, and agent-routing"); + assert.equal(result.removedBytes, Buffer.byteLength(REALISTIC_AGENTS_MD) - Buffer.byteLength(EXPECTED_REALISTIC)); + assert.deepEqual(result.failSafePaths, ["/repo/sub/AGENTS.md"]); +}); + +test("filterChildSessionContextFiles replaces contextFiles on the same options object and tolerates bad input", () => { + const original = [{ path: "/home/AGENTS.md", content: REALISTIC_AGENTS_MD }]; + const options: { contextFiles?: Array<{ path: string; content: string }> } = { contextFiles: original }; + filterChildSessionContextFiles(options); + assert.notEqual(options.contextFiles, original); + assert.equal(options.contextFiles?.[0].content, EXPECTED_REALISTIC); + assert.equal(original[0].content, REALISTIC_AGENTS_MD, "input array entries are never mutated"); + + const empty: { contextFiles?: Array<{ path: string; content: string }> } = {}; + filterChildSessionContextFiles(empty); + assert.equal("contextFiles" in empty, false); + assert.doesNotThrow(() => filterChildSessionContextFiles(undefined)); + assert.doesNotThrow(() => filterChildSessionContextFiles(null)); + + const malformed = { contextFiles: [{ path: "/x", content: 42 }] as unknown as Array<{ path: string; content: string }> }; + const malformedFiles = malformed.contextFiles; + assert.doesNotThrow(() => filterChildSessionContextFiles(malformed)); + assert.equal(malformed.contextFiles, malformedFiles, "on any error the original context files stay in place"); +}); + +type BeforeAgentStart = (event: unknown, ctx: unknown) => unknown; + +function childContextHandlers(env: NodeJS.ProcessEnv): { handlers: Map; registered: string[] } { + const handlers = new Map(); + const registered: string[] = []; + const pi = { + on(name: string, handler: BeforeAgentStart) { + registered.push(name); + handlers.set(name, handler); + }, + } as unknown as ExtensionAPI; + createChildContextExtension(env)(pi); + return { handlers, registered }; +} + +test("the child-context extension registers only before_agent_start", () => { + assert.deepEqual(childContextHandlers({}).registered, ["before_agent_start"]); + assert.equal(typeof childContextExtension, "function"); +}); + +for (const scenario of ["child", "primary"] as const) { + test(`the child-context extension ${scenario === "child" ? "filters" : "leaves"} context files on the shared options object for a ${scenario} session`, async () => { + const { handlers } = childContextHandlers(scenario === "child" ? { GENTLE_PI_AGENTS_CHILD: "1" } : { GENTLE_PI_AGENTS_CHILD: "0" }); + const contextFiles = [ + { path: "/home/AGENTS.md", content: REALISTIC_AGENTS_MD }, + { path: "/repo/CLAUDE.md", content: "# Plain\n" }, + ]; + const systemPromptOptions = { appendSystemPrompt: "", contextFiles }; + const event = { systemPrompt: "base", systemPromptOptions }; + const result = await handlers.get("before_agent_start")!(event, {}); + assert.equal(result, undefined, "the extension never returns a replacement system prompt"); + assert.equal(event.systemPromptOptions, systemPromptOptions, "the same options object stays in place"); + assert.equal(contextFiles[0].content, REALISTIC_AGENTS_MD, "the original entries are never mutated"); + if (scenario === "primary") { + assert.equal(systemPromptOptions.contextFiles, contextFiles); + return; + } + assert.notEqual(systemPromptOptions.contextFiles, contextFiles); + assert.deepEqual(systemPromptOptions.contextFiles.map((file) => file.content), [EXPECTED_REALISTIC, "# Plain\n"]); + // Idempotent: a second run over already-filtered content removes nothing. + const filtered = systemPromptOptions.contextFiles; + await handlers.get("before_agent_start")!(event, {}); + assert.equal(systemPromptOptions.contextFiles, filtered); + }); +} + +test("the child-context extension never throws on missing or malformed options", async () => { + const { handlers } = childContextHandlers({ GENTLE_PI_AGENTS_CHILD: "1" }); + const handler = handlers.get("before_agent_start")!; + await assert.doesNotReject(async () => handler({}, {})); + await assert.doesNotReject(async () => handler({ systemPromptOptions: null }, {})); + const malformed = { systemPromptOptions: { contextFiles: [{ path: "/x", content: 42 }] } }; + const original = malformed.systemPromptOptions.contextFiles; + await assert.doesNotReject(async () => handler(malformed, {})); + assert.equal(malformed.systemPromptOptions.contextFiles, original); +}); diff --git a/tests/gentle-agents.test.ts b/tests/gentle-agents.test.ts index f502fe3bd..6ae83effe 100644 --- a/tests/gentle-agents.test.ts +++ b/tests/gentle-agents.test.ts @@ -13,7 +13,7 @@ import type { TestContext } from "node:test"; import { generateUnifiedPatch, type ExtensionAPI, type ExtensionContext } from "@earendil-works/pi-coding-agent"; import { visibleWidth, type TUI, type TuiMouseEvent } from "@earendil-works/pi-tui"; import { sidebarState } from "../lib/shell-sidebar.ts"; -import gentleAgents, { agentRuntimePaths, agentsCollapseKey, agentsEnabled, agentsStopKey, agentsViewKey, answerThroughUi, completionText, createDefaultSessionTransport, legacySubagentsInstalled, type AgentsDeps, type SessionTransportFactory } from "../extensions/gentle-agents.ts"; +import gentleAgents, { agentRuntimePaths, agentsCollapseKey, agentsEnabled, agentsStopKey, agentsViewKey, answerThroughUi, childContextExtensionPaths, completionText, createDefaultSessionTransport, legacySubagentsInstalled, type AgentsDeps, type SessionTransportFactory } from "../extensions/gentle-agents.ts"; import { ActiveSessionClient, ActiveSessionListener, SessionPresenceRegistry } from "../lib/agents-session-transport.ts"; import { WindowsActiveSessionClient, WindowsActiveSessionListener } from "../lib/windows-session-transport.ts"; import { historyDir, loadHistory, saveTask } from "../lib/agents-history.ts"; @@ -2041,7 +2041,7 @@ test("default Node spawn adapter distinguishes IPC-only and permission-capable c children[2]!.emit({ type: "agent_settled" }); await permission.result; - const args = ["--host-flag", "--mode", "rpc", "--session-dir", join(home, ".pi", "agent", "gentle-agents", "sessions"), "--model", "openai-codex/gpt-5.6-terra:low", "--tools", "read,grep,subagent_parent_message", "--append-system-prompt", "You map things."]; + const args = ["--host-flag", "--mode", "rpc", "--session-dir", join(home, ".pi", "agent", "gentle-agents", "sessions"), ...childContextExtensionPaths().flatMap((path) => ["--extension", path]), "--model", "openai-codex/gpt-5.6-terra:low", "--tools", "read,grep,subagent_parent_message", "--append-system-prompt", "You map things."]; assert.equal(captured.length, 3, "the extension reaches Node's spawn boundary for IPC-only and permission-channel launches"); const permissionChannelStdio = process.platform === "win32" ? "overlapped" : "pipe"; for (const [index, fixture] of ["task", "background", "permission"].entries()) { @@ -4303,3 +4303,30 @@ test("issue #1162: task-mode subagent_run includes question directly in waiting await fire("session_shutdown", ctx); }); + +// gentle-shell#1587: children do not load the gentle-pi package in the +// isolated Gentle Shell home, so every child receives the child-context +// extension explicitly through --extension. +test("children receive the child-context extension, and a missing file is omitted", async () => { + const expected = join(dirname(new URL(import.meta.url).pathname), "..", "extensions", "child-context.ts"); + assert.deepEqual(childContextExtensionPaths(), [resolve(expected)]); + assert.deepEqual(childContextExtensionPaths(() => false), [], "a missing extension file fails safe to no --extension"); + const extensionArguments = (args: string[]) => args.filter((_, index) => args[index - 1] === "--extension"); + for (const scenario of ["present", "missing"] as const) { + const h = fakePi(); + const runtime = deps(); + if (scenario === "missing") runtime.deps.childExtensionPaths = []; + gentleAgents(h.pi, {}, runtime.deps); + const { ctx } = fakeContext(); + await h.fire("session_start", ctx); + try { + await h.tools.get("subagent_run")!.execute(`child-context-${scenario}`, { agent: "explore", task: "Map", mode: "background" }, undefined, undefined, ctx); + await tick(); + assert.equal(runtime.spawned.length, 1); + assert.deepEqual(extensionArguments(runtime.spawned[0]!), scenario === "present" ? [resolve(expected)] : []); + } finally { + await h.fire("session_shutdown", ctx); + await tick(); + } + } +}); From 14bd23568e7a39af1cf6ebbf158f2b91cc0eed71 Mon Sep 17 00:00:00 2001 From: Alan Buscaglia Date: Wed, 30 Sep 2026 17:04:46 +0200 Subject: [PATCH 2/5] feat(orchestrator): replace file-count delegation triggers with an evidence budget Replace the 4-file rule and the ~20 tool calls / 5 reads backstop with the measured evidence-budget rule: read inline only one parallel batch of at most 3 calls and ~10k tokens, delegate larger or sequential mapping to one explorer with a ~2k-token path:line handoff, never force delegation for small targeted questions, and back stop on parent context size with bounded command output. Explorer and verifier handoffs are capped at ~2k tokens. Gentle Shell leads the upstream canon here; the canon fixture is unchanged and the divergence is tracked by gentle-ai#5139. Refs #1587 --- assets/agents/gentle-ai-explore.md | 2 +- assets/agents/gentle-ai-verify.md | 2 +- assets/orchestrator-delegation.md | 22 ++++++------- assets/orchestrator.md | 8 ++--- docs/readme-reference.md | 4 +-- skills/gentle-ai/SKILL.md | 6 ++-- tests/odd-routing-canonical-ratchet.test.ts | 13 +++++--- tests/odd-routing-contract.test.ts | 36 ++++++++++++++++++--- tests/orchestrator-budget.test.ts | 4 +-- tests/package-manifest.test.ts | 2 +- 10 files changed, 65 insertions(+), 34 deletions(-) diff --git a/assets/agents/gentle-ai-explore.md b/assets/agents/gentle-ai-explore.md index 02ac815cf..4e2622778 100644 --- a/assets/agents/gentle-ai-explore.md +++ b/assets/agents/gentle-ai-explore.md @@ -19,4 +19,4 @@ Map relevant files, symbols, relationships, and uncertainty within the parent-pr - Do not fix findings, delegate to child agents, commit, or push. - Do not use review lenses. RDD review remains independent and parent-owned. -Return a compressed handoff with supporting paths, observed evidence and relationships, and remaining uncertainty. Never claim evidence you did not observe. +Return a compressed handoff of at most ~2k tokens: `path:line` evidence, observed relationships, and remaining uncertainty. Never claim evidence you did not observe. diff --git a/assets/agents/gentle-ai-verify.md b/assets/agents/gentle-ai-verify.md index 924a89e94..405ccb8aa 100644 --- a/assets/agents/gentle-ai-verify.md +++ b/assets/agents/gentle-ai-verify.md @@ -20,4 +20,4 @@ For behavior changes with applicable runnable deterministic tests and a clear ex - Do not delegate to child agents, commit, or push. - Do not use review lenses. RDD review remains independent and parent-owned. -Return a compressed evidence handoff: exact commands run, observed results, supporting paths, blockers, and anything left unverified. Never claim a command ran or a check passed without observed output. +Return a compressed evidence handoff of at most ~2k tokens: exact commands run, observed results, `path:line` evidence, blockers, and anything left unverified. Never claim a command ran or a check passed without observed output. diff --git a/assets/orchestrator-delegation.md b/assets/orchestrator-delegation.md index 10c24ee6b..eceb5a07f 100644 --- a/assets/orchestrator-delegation.md +++ b/assets/orchestrator-delegation.md @@ -110,8 +110,8 @@ Core principle: **does this inflate the parent context without need?** If yes, u | Action | Direct inline | Delegated direct worker | |--------|---------------|-------------------------| -| Read to decide/verify (1–3 files) | ✅ | — | -| Read to explore/understand (4+ files) | — | ✅ one narrow mapper | +| Read to decide/verify within the evidence budget (one parallel batch: at most 3 calls, ~10k tokens) | ✅ | — | +| Read to explore/understand beyond the evidence budget | — | ✅ one narrow explorer (handoff of at most ~2k tokens, `path:line` evidence) | | Read as preparation for writing | — | ✅ together with the write | | Write one mechanical, already-understood file | ✅ | — | | Write 2+ non-trivial files | — | ✅ one writer | @@ -126,11 +126,11 @@ Keep one writer and a short synthesized handoff. Delegation is mandatory at the These are parent-orchestrator routing boundaries; do not pass these rules to child agents as permission to orchestrate. These triggers are mandatory, not advisory. When one fires, stop and delegate through the runtime's subagent mechanism before continuing; executing past a fired trigger inline is a routing defect even if the work succeeds. Delegation keeps the parent context thin enough to orchestrate; it does not slow the work down. -1. **Mapping trigger (4-file rule):** when understanding the work requires 4 or more files, delegate one narrow exploration or mapping task before deciding or writing anything. +1. **Mapping trigger (Evidence-budget rule):** read inline only when the evidence fits one parallel batch of at most 3 calls totaling ~10k tokens, using grep and line ranges, never whole large files. When the reading is larger, needs more than ~5 sequential lookups, or the session has a long way to go, delegate one scout/explorer that returns a handoff of at most ~2k tokens with `path:line` evidence before deciding or writing anything. Never force delegation for a small targeted question. The parent does not re-read what the handoff covered, except a single spot check. 2. **Writer trigger (Multi-file write rule):** when implementation touches 2 or more non-trivial files, delegate one bounded writer instead of editing them inline. 3. **Incident rule:** after wrong `cwd`, accidental repository/worktree mutation, failed merge recovery, confusing test command, or environment workaround, stop and diagnose the incident separately before resuming. -4. **Long-session backstop (Long-session rule):** after about 20 tool calls, 5 exploratory reads, or 2 non-mechanical edits without any delegation, pause and delegate the next bounded unit of work. -5. **Verification rule** (gentle-pi#661/#662, RDD-aware): executing or delegating verification commands goes to `gentle-ai-verify`; only the 1–3-file read-only check stays inline. The normative on/off/unknown routing is stated once under Pi Trigger Runtime Bindings below; reference it, do not restate it. +4. **Context backstop:** when the parent context passes ~150k tokens, pause and delegate the next bounded unit of work. Always keep command output bounded in the parent (counts, `--stat`, `tail`); send full suites and builds to a verifier. +5. **Verification rule** (gentle-pi#661/#662, RDD-aware): executing or delegating verification commands goes to `gentle-ai-verify`; only a read-only check within the evidence budget stays inline. The normative on/off/unknown routing is stated once under Pi Trigger Runtime Bindings below; reference it, do not restate it. **Preparation trigger:** reading that prepares a write, and broad research or context compression, delegate together with or ahead of the write instead of filling the parent context. @@ -164,10 +164,10 @@ Once a trigger fires, the parent MUST delegate through the best available subage The bounded multi-file writer precedence in rule 3 overrides that general runtime preference. If no delegation mechanism is available, stop and explain the blocker. -1. **4-file rule**: launch `scout`, `context-builder`, or the closest read-only mapping subagent with fresh context and a narrow mapping task. Route generic exploration to `gentle-ai-explore`; if missing or unusable, use native `Agent` with the same read-only mapping task and report the fallback. +1. **Evidence-budget rule**: when the reading exceeds the evidence budget, launch `scout`, `context-builder`, or the closest read-only mapping subagent with fresh context and a narrow mapping task that returns a handoff of at most ~2k tokens with `path:line` evidence. Route generic exploration to `gentle-ai-explore`; if missing or unusable, use native `Agent` with the same read-only mapping task and report the fallback. 2. **Multi-file write rule**: for bounded multi-file writes, prefer the installed package-owned `gentle-ai-worker`, then a user-configured `worker`. If neither worker definition exists, fall back to the native `Agent` even when `subagent_*` tools are available. If no delegation mechanism is available, stop and explain the blocker. 3. **Incident rule**: after wrong `cwd`, accidental repository/worktree mutation, failed merge recovery, confusing test command, or environment workaround, stop and diagnose the incident separately before resuming. -4. **Long-session rule**: if accumulating work is no longer clearly local — roughly 20 tool calls, 5 exploratory file reads, or 2 non-mechanical edits without delegation — pause and delegate the remaining work instead of silently continuing monolithically. +4. **Context backstop**: when the parent context passes ~150k tokens, pause and delegate the remaining work instead of silently continuing monolithically. 5. **Verification rule** (gentle-pi#661/#662, RDD-aware; normative -- referenced, not restated, elsewhere in this file): read the rendered `Receipt-driven development:` line next to `Background subagent policy`. The bounded writer always runs the exact parent-authorized commands under the delegated task's `## Verification` heading, synchronously and in the foreground, and reports each as `: ` -- see `gentle-ai-worker`'s Verification contract for the exact rules, including how `## Known environmental failures` (exact pre-existing base failures) differs from any other failing required command, which still forces `status: partial`. Those foreground commands are live work, not silence: while a tool call is in flight the runner's stall watchdog uses `tool_stall_timeout_ms` (default 30 minutes) instead of the `stall_timeout_ms` idle budget. When the line reads `on`, that writer report is the verification of record, and the native review is the independent check the writer cannot influence: `gentle-ai-verify` (or the native `Agent` fallback, with the same read-only verification task and exact parent-authorized commands) becomes on-demand -- reach for it only when the writer reports `partial`/`blocked`, the check is expensive or external (E2E runs, installs) and the parent wants a cheaper profile, or the parent wants an independent spot check. That `on` branch holds only while the native review actually reaches a terminal outcome for this candidate (gentle-pi#668): a human decline of the consent envelope for this candidate (candidate-scoped, never the RDD kill switch), a clone-local RDD disable discovered mid-flow, or a refused START/STATUS all fall back to the risk-gated path exactly as `off` -- call `gentle_review` with `{"operation":"assess"}` (pass `nativeReviewOutcome` when the parent already knows it; the tool derives it from what it itself observed for the candidate otherwise, failing closed to `unknown` when it cannot) and follow the returned plan. ASSESS resolves that closure itself (gentle-pi#1175): it derives `closed` only from the native `candidate.consumed` fact for this exact candidate, so a caller-supplied `closed` is not authority and, without that fact, resolves to `unknown`; a declined, unavailable, or unknown outcome falls back to the risk-gated path, and `unknown` is never treated as closed. When the line reads `off` or `unknown`, after the writer returns, call `gentle_review` with `{"operation":"assess"}` over the writer's diff and follow the returned plan instead of judging non-triviality from the task description: the operation resolves the native risk tier and states exactly who verifies next. The tier table (stated once, here): | Native risk tier | Verification when RDD is `off`/`unknown` | @@ -177,7 +177,7 @@ The bounded multi-file writer precedence in rule 3 overrides that general runtim | high | writer self-verification plus a separate `gentle-ai-verify` run, always | | unknown / assess failed | treated as high | -The small-model bias raises the tier by one for verification purposes (medium becomes high); an unknown `Receipt-driven development:` line never lowers a tier below `off`. The parent spot check (re-running one reported command before delivery) stays required in every tier. ASSESS takes the writer profile from the runtime-recorded model and effort of the pending mutations for the root; caller `writerModelId`/`writerEffort` are only a fallback when no runtime evidence exists, and a missing model, a `mini` model token (`gemini` is not mini), or `low` effort keeps the conservative small-model bias. When native reports them, ASSESS also projects `reviewDue`, `reviewDueReason`, `candidate.consumed`, and the native continuation verbatim; relay that continuation unchanged and never rebuild it. A native code review is not a substitute for applicable functional checks: tests, builds, and functional verification such as browser checks for UI changes still run when applicable, and review outcomes never authorize delivery. Only truly local read-only checking of 1–3 known files stays inline. +The small-model bias raises the tier by one for verification purposes (medium becomes high); an unknown `Receipt-driven development:` line never lowers a tier below `off`. The parent spot check (re-running one reported command before delivery) stays required in every tier. ASSESS takes the writer profile from the runtime-recorded model and effort of the pending mutations for the root; caller `writerModelId`/`writerEffort` are only a fallback when no runtime evidence exists, and a missing model, a `mini` model token (`gemini` is not mini), or `low` effort keeps the conservative small-model bias. When native reports them, ASSESS also projects `reviewDue`, `reviewDueReason`, `candidate.consumed`, and the native continuation verbatim; relay that continuation unchanged and never rebuild it. A native code review is not a substitute for applicable functional checks: tests, builds, and functional verification such as browser checks for UI changes still run when applicable, and review outcomes never authorize delivery. Only a truly local read-only check within the evidence budget stays inline. ### Work Routing Ladder @@ -185,11 +185,11 @@ Route work through the smallest harness that is safe. "Smallest" means minimal s #### 1. Inline Direct -Use inline execution when the task is small, mechanical, and the parent already has enough context: a typo, rename, one-file mechanical edit, a small known bug, focused verification over 1–3 files, or bash for state. Keep the ODD path proportionate. Do not use this exception to avoid delegation after the task stops being small. +Use inline execution when the task is small, mechanical, and the parent already has enough context: a typo, rename, one-file mechanical edit, a small known bug, focused verification within the evidence budget, or bash for state. Keep the ODD path proportionate. Do not use this exception to avoid delegation after the task stops being small. #### 2. Simple Delegation -Delegate when work would inflate parent context or requires focused exploration, validation, or multi-file implementation, within the ODD workflow. Examples include understanding an unfamiliar module, inspecting 4+ files, investigating a failing test, implementing a bounded multi-file change, or running focused tests/builds. +Delegate when work would inflate parent context or requires focused exploration, validation, or multi-file implementation, within the ODD workflow. Examples include understanding an unfamiliar module, reading beyond the evidence budget, investigating a failing test, implementing a bounded multi-file change, or running focused tests/builds. Use the configured subagent runtime when available. Prefer the `subagent_*` tools (`subagent_run`, status/result helpers) when the Pi Subagents extension is installed, because they run the user's configured project/global subagent definitions and preserve history/background behavior. @@ -215,7 +215,7 @@ For generic exploration and mapping, first attempt the installed package-owned ` For bounded multi-file writes, prefer the installed package-owned `gentle-ai-worker`, then a user-configured `worker`. If neither worker definition exists, fall back to the native `Agent` even when `subagent_*` tools are available. If no delegation mechanism is available, stop and explain the blocker. This writer precedence overrides the general runtime preference above. -Delegate generic verification that executes or delegates commands per the RDD-aware Verification rule (trigger 5 under Mandatory Delegation Triggers, gentle-pi#661) -- the normative on/off/unknown routing lives there, not here: the bounded writer always self-verifies via `## Verification`, and `gentle-ai-verify` (or the native `Agent` fallback, with the same read-only verification constraints, exact parent-authorized commands, and fallback reporting) is on-demand only when the rendered `Receipt-driven development:` line reads `on`; when the line reads `off` or `unknown`, the `gentle_review` `assess` operation's returned plan decides it by native risk tier instead of a blanket non-trivial rule (gentle-pi#662). `## Known environmental failures` follows the same definition as `gentle-ai-worker`'s Verification contract: exact pre-existing base failures reported as evidence, never blockers -- any other failing required command still forces `status: partial`. Truly local read-only checking of 1–3 known files may remain inline. Separate exploration stays reserved for when the parent needs the map to decide or route; reading that prepares a write belongs with the writer making the change, consistent with the Delegation Rules table above. +Delegate generic verification that executes or delegates commands per the RDD-aware Verification rule (trigger 5 under Mandatory Delegation Triggers, gentle-pi#661) -- the normative on/off/unknown routing lives there, not here: the bounded writer always self-verifies via `## Verification`, and `gentle-ai-verify` (or the native `Agent` fallback, with the same read-only verification constraints, exact parent-authorized commands, and fallback reporting) is on-demand only when the rendered `Receipt-driven development:` line reads `on`; when the line reads `off` or `unknown`, the `gentle_review` `assess` operation's returned plan decides it by native risk tier instead of a blanket non-trivial rule (gentle-pi#662). `## Known environmental failures` follows the same definition as `gentle-ai-worker`'s Verification contract: exact pre-existing base failures reported as evidence, never blockers -- any other failing required command still forces `status: partial`. A truly local read-only check within the evidence budget may remain inline. Separate exploration stays reserved for when the parent needs the map to decide or route; reading that prepares a write belongs with the writer making the change, consistent with the Delegation Rules table above. #### Allowed edit surfaces (MANDATORY) diff --git a/assets/orchestrator.md b/assets/orchestrator.md index 4f4211b29..4b0f9005a 100644 --- a/assets/orchestrator.md +++ b/assets/orchestrator.md @@ -38,7 +38,7 @@ Delegation is not optional once complexity appears. If a task crosses the trigge Route ODD work through the smallest safe harness: -1. **Inline Direct** — small, mechanical, parent has context (typo, one-file edit, read-only check of 1-3 known files, bash for state); stop when it is no longer small. +1. **Inline Direct** — small, mechanical, parent has context (typo, one-file edit, read-only check within the evidence budget, bash for state); stop when it is no longer small. 2. **Simple Delegation** — exploration → `gentle-ai-explore`; bounded implementation → `gentle-ai-worker`; command-running verification → `gentle-ai-verify`. Try its package role; if missing/unusable, use native `Agent` under the same read-only mapping/verification constraints and report fallback. ODD (Default Workflow, harness section above) is mandatory on every request; detail: `orchestrator-delegation.md`, `orchestrator-memory.md`. For behavior changes with applicable runnable deterministic tests and a clear expected outcome, use test-first by default: observed RED, GREEN, then refactor with checks. For passive documentation, non-testable changes, an unavailable runner or no meaningful RED, state why and run proportionate ordinary functional or structural verification instead. Test presence alone is not applicability; no chat or TUI toggle activates this policy. @@ -51,11 +51,11 @@ Before launching bounded writer (`gentle-ai-worker` or `worker`), task/context n Mandatory Delegation Triggers — once fired, delegate through the best available runtime (prefer `subagent_run`, else native `Agent`): -1. **4-file rule** — 4+ files to understand → delegate a scout/mapping task. +1. **Evidence-budget rule** — read inline only if evidence fits one parallel batch (at most 3 calls, ~10k tokens; grep and line ranges, never whole large files). Larger reads, >~5 sequential lookups, or a long session ahead → one scout/explorer returning at most ~2k tokens with `path:line` evidence; re-read nothing it covered beyond one spot check. Never force delegation for a small targeted question. 2. **Multi-file write rule** — 2+ non-trivial files touched → delegate one writer. 3. **Incident rule** — diagnose wrong cwd/worktree/git/tooling incidents separately before resuming work. -4. **Long-session rule** — ~20 tool calls, 5 exploratory reads, or 2 non-mechanical edits without delegation → pause and delegate. -5. **Verification rule** — executing/delegating verification commands → `gentle-ai-verify`; only the 1-3-file read-only check stays inline. +4. **Context backstop** — parent context past ~150k tokens → pause and delegate the next bounded unit of work. Keep command output bounded (counts, `--stat`, `tail`); full suites and builds go to a verifier. +5. **Verification rule** — executing/delegating verification commands → `gentle-ai-verify`; only a read-only check within the evidence budget stays inline. {{GENTLE_PI_BACKGROUND_POLICY}}; rules: the background-subagents block in the delegation contract. diff --git a/docs/readme-reference.md b/docs/readme-reference.md index 57ce449da..a7b587f95 100644 --- a/docs/readme-reference.md +++ b/docs/readme-reference.md @@ -420,11 +420,11 @@ Size and uncertainty can call for scoped exploration or delegation within ODD. T | Trigger | Required behavior | | --------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------- | -| Reading 4+ files to understand a flow | Launch `scout`, `context-builder`, or the closest read-only mapping subagent. | +| Reading beyond the evidence budget (one parallel batch of at most 3 calls, ~10k tokens), more than ~5 sequential lookups, or a long session ahead | Launch `scout`, `context-builder`, or the closest read-only mapping subagent; it returns a handoff of at most ~2k tokens with `path:line` evidence. Never force delegation for a small targeted question. | | Touching 2+ non-trivial code files | Delegate one writer; do not continue inline unless delegation is unavailable. | | Commit, push, or PR after code changes | Follow the loaded native instruction, or ordinary repository policy when none is supplied. | | Wrong cwd, worktree/git accident, merge recovery, confusing test/env issue | Stop, preserve the affected scope, and investigate separately before resuming. | -| Long monolithic session with accumulating complexity, roughly 20 tool calls, 5 exploratory reads, or 2 non-mechanical edits | Pause and delegate the remaining work, or stop and explain the exact blocker. | +| Parent context past ~150k tokens | Pause and delegate the next bounded unit of work, or stop and explain the exact blocker. Keep command output bounded; send full suites and builds to a verifier. | The intended balanced loop for a bounded bugfix is: diff --git a/skills/gentle-ai/SKILL.md b/skills/gentle-ai/SKILL.md index 7ceb2f789..c6aa635f7 100644 --- a/skills/gentle-ai/SKILL.md +++ b/skills/gentle-ai/SKILL.md @@ -17,7 +17,7 @@ When asked who or what you are, answer as el Gentleman: a Pi-specific coding-age - For substantial authorized ODD work, track the feature in its task document and mirror. - For behavior changes with applicable runnable deterministic tests and a clear expected outcome, use test-first by default: observe RED, GREEN, relevant alternate cases, then REFACTOR and record evidence. Test presence alone does not establish applicability. For passive documentation, non-testable changes, an unavailable runner, or no meaningful RED, state why and run proportionate ordinary functional or structural verification. Never invent RED/GREEN, skip checks, or require a chat/TUI toggle. - Keep one parent session responsible for orchestration; child subagents should receive concrete phase work and must not spawn more subagents. -- Parent-only delegation triggers apply after complexity appears: 4+ files for understanding, 2+ non-trivial files to write, tooling/worktree incidents, or long sessions with accumulating complexity. +- Parent-only delegation triggers apply after complexity appears: reading beyond the evidence budget, 2+ non-trivial files to write, tooling/worktree incidents, or a parent context past the context backstop. - Keep writes single-threaded unless the user explicitly approves isolated parallel worktrees. - Forecast review workload before large changes; ask before producing oversized or multi-area diffs. - Keep dangerous-command safety independent and authoritative. @@ -43,10 +43,10 @@ clarify → scout/context-builder when context-heavy → one worker → verify Hard delegation triggers: -- **4-file rule**: reading 4+ files to understand means delegate exploration. +- **Evidence-budget rule**: read inline only when the evidence fits one parallel batch of at most 3 calls, ~10k tokens (grep and line ranges, never whole large files). Larger reading, more than ~5 sequential lookups, or a long session ahead means delegate one explorer that returns a handoff of at most ~2k tokens with `path:line` evidence. Never force delegation for a small targeted question; do not re-read what the handoff covered beyond one spot check. - **Multi-file write rule**: touching 2+ non-trivial files means use one worker. - **Incident rule**: after wrong cwd, accidental worktree/repo mutation, merge recovery, confusing test command, or environment workaround, diagnose separately. -- **Long-session rule**: after roughly 20 tool calls, 5 exploratory reads, or 2 non-mechanical edits with no delegation and accumulating complexity, pause and choose a non-review subagent or justify not doing so. +- **Context backstop**: when the parent context passes ~150k tokens, pause and delegate the next bounded unit of work to a non-review subagent. Keep command output bounded (counts, `--stat`, `tail`) and send full suites and builds to a verifier. ## Review Lens Selection diff --git a/tests/odd-routing-canonical-ratchet.test.ts b/tests/odd-routing-canonical-ratchet.test.ts index 64b3a4ff7..e0a4a843f 100644 --- a/tests/odd-routing-canonical-ratchet.test.ts +++ b/tests/odd-routing-canonical-ratchet.test.ts @@ -86,9 +86,12 @@ const ANCHORS: readonly RoutingAnchor[] = [ { label: "mapping trigger at 4 or more files", canonical: "**Mapping trigger:** when understanding the work requires 4 or more files", + // Gentle Shell intentionally leads the canon here: the local mirrors carry + // the measured evidence-budget rule instead of the 4-file count (tracked by + // gentle-ai#5139). The canonical anchor stays until gentle-ai follows. mirrors: [ - { surface: DELEGATION, includes: "**Mapping trigger (4-file rule):** when understanding the work requires 4 or more files" }, - { surface: CORE, includes: "**4-file rule** — 4+ files to understand" }, + { surface: DELEGATION, includes: "**Mapping trigger (Evidence-budget rule):** read inline only when the evidence fits one parallel batch of at most 3 calls" }, + { surface: CORE, includes: "**Evidence-budget rule** — read inline only if evidence fits one parallel batch (at most 3 calls, ~10k tokens" }, ], }, { @@ -107,9 +110,11 @@ const ANCHORS: readonly RoutingAnchor[] = [ { label: "long-session backstop", canonical: "**Long-session backstop:**", + // Gentle Shell intentionally leads the canon here: a parent-context token + // backstop replaces the tool-call count (tracked by gentle-ai#5139). mirrors: [ - { surface: DELEGATION, includes: "**Long-session backstop (Long-session rule):**" }, - { surface: CORE, includes: "**Long-session rule** — ~20 tool calls, 5 exploratory reads, or 2 non-mechanical edits without delegation" }, + { surface: DELEGATION, includes: "**Context backstop:** when the parent context passes ~150k tokens" }, + { surface: CORE, includes: "**Context backstop** — parent context past ~150k tokens" }, ], }, { diff --git a/tests/odd-routing-contract.test.ts b/tests/odd-routing-contract.test.ts index 3b50b1e6d..cfdf40873 100644 --- a/tests/odd-routing-contract.test.ts +++ b/tests/odd-routing-contract.test.ts @@ -214,7 +214,7 @@ test("mandatory delegation triggers are behavioral in the lazy canonical port an "**Mapping trigger", "**Writer trigger", "**Preparation trigger:**", - "**Long-session backstop", + "**Context backstop:**", "pause and delegate the next bounded unit of work", "**Route declaration:**", "record the chosen route per task", @@ -235,19 +235,19 @@ test("mandatory delegation triggers are behavioral in the lazy canonical port an test("core and lazy canonical trigger lists agree in numbering and semantics", () => { for (const entry of [ - "1. **4-file rule**", + "1. **Evidence-budget rule**", "2. **Multi-file write rule**", "3. **Incident rule**", - "4. **Long-session rule**", + "4. **Context backstop**", "5. **Verification rule**", ]) { assert.ok(core.includes(entry), `always-on core trigger list is missing: ${entry}`); } for (const entry of [ - "1. **Mapping trigger (4-file rule):**", + "1. **Mapping trigger (Evidence-budget rule):**", "2. **Writer trigger (Multi-file write rule):**", "3. **Incident rule:**", - "4. **Long-session backstop (Long-session rule):**", + "4. **Context backstop:**", "5. **Verification rule**", ]) { assert.ok(delegation.includes(entry), `lazy canonical trigger list is missing: ${entry}`); @@ -262,6 +262,32 @@ test("core and lazy canonical trigger lists agree in numbering and semantics", ( } }); +// Evidence-budget rule (gentle-shell#1587, measured in gentle-ai#5139): every +// routing surface states the same numbers, and the retired file-count and +// tool-call triggers are gone from all of them. +test("evidence-budget numbers agree across routing surfaces and retired triggers are gone", () => { + const surfaces: Record = { + "assets/orchestrator.md": core, + "assets/orchestrator-delegation.md": delegation, + "skills/gentle-ai/SKILL.md": read("skills/gentle-ai/SKILL.md"), + }; + for (const [path, text] of Object.entries(surfaces)) { + containsAll(text, [ + "**Evidence-budget rule**", + "at most 3 calls", + "~10k tokens", + "at most ~2k tokens", + "`path:line`", + "~150k tokens", + "Context backstop", + ]); + assert.doesNotMatch(text, /4-file rule|20 tool calls|5 exploratory (?:file )?reads/, `${path} keeps a retired trigger`); + } + for (const agent of ["assets/agents/gentle-ai-explore.md", "assets/agents/gentle-ai-verify.md"]) { + containsAll(read(agent), ["at most ~2k tokens", "`path:line`"]); + } +}); + test("ODD protocol is always-on in the rendered system prompt and runs by default", () => { const orderedClauses = [ "Default workflow: Organic Driven Development (MANDATORY)", diff --git a/tests/orchestrator-budget.test.ts b/tests/orchestrator-budget.test.ts index 466800099..02aaccc30 100644 --- a/tests/orchestrator-budget.test.ts +++ b/tests/orchestrator-budget.test.ts @@ -350,11 +350,11 @@ for (const range of DISPOSITION_MAP) { test("core-alone: load-bearing direct-delegation tokens remain without lazy union", () => { const core = readRealAsset("orchestrator.md"); - assert.match(core, /4-file rule/); + assert.match(core, /Evidence-budget rule/); assert.match(core, /Multi-file write rule/); assert.match(core, /Incident rule/); assert.match(core, /Verification rule/); - assert.match(core, /Long-session rule/); + assert.match(core, /Context backstop/); }); test("core-alone: dynamic Gentle AI ownership replaces package lifecycle instructions", () => { diff --git a/tests/package-manifest.test.ts b/tests/package-manifest.test.ts index 3f80304c6..6e3b46327 100644 --- a/tests/package-manifest.test.ts +++ b/tests/package-manifest.test.ts @@ -1618,7 +1618,7 @@ test("orchestrator routes generic roles without static RDD lens routing", () => assert.match(routing, /`gentle-ai-explore`/); assert.match(routing, /`gentle-ai-worker`/); assert.match(routing, /`gentle-ai-verify`/); - assert.match(routing, /(?:truly local )?read-only check(?:ing)? of (?:known )?1[-–]3 known files|1[-–]3-file read-only check/); + assert.match(routing, /read-only check within the evidence budget/); assert.match(routing, /(?:verification that |verification commands →).*executes? or delegates?|executing\/delegating verification commands/); assert.match(routing, /missing(?: or |\/)unusable[\s\S]*native `Agent`[\s\S]*(?:the )?same read-only/); assert.match(routing, /report (?:the )?fallback/); From 73c00efe2383e88ac028cda1ce3d795ab5c46a40 Mon Sep 17 00:00:00 2001 From: Alan Buscaglia Date: Wed, 30 Sep 2026 17:07:33 +0200 Subject: [PATCH 3/5] docs(odd): record lean delegation context work units Refs #1587 --- odd/tasks/lean-delegation-context.md | 137 +++++++++++++++++++++++++++ 1 file changed, 137 insertions(+) create mode 100644 odd/tasks/lean-delegation-context.md diff --git a/odd/tasks/lean-delegation-context.md b/odd/tasks/lean-delegation-context.md new file mode 100644 index 000000000..7cd490286 --- /dev/null +++ b/odd/tasks/lean-delegation-context.md @@ -0,0 +1,137 @@ +# Lean delegation context and evidence-budget rule (issue #1587) + +Locator: `odd/tasks/lean-delegation-context.md` · Engram mirror: `odd/lean-delegation-context/tasks` (project `gentle-pi`) + +Issue: https://github.com/Gentleman-Programming/gentle-shell/issues/1587 (`status:approved`) +Full study: https://github.com/Gentleman-Programming/gentle-ai/issues/5139 +Branch: `feat/lean-delegation-context` (base `origin/main` 289cee5ba) + +## Objective + +Stop delegated subagents from loading the gentle-ai orchestrator-only managed `AGENTS.md` +blocks while keeping project conventions, and replace the file-count / tool-call delegation +triggers in the Gentle Shell prompt assets with the measured evidence-budget rule. + +## Why (measured) + +- Child explorer prefix: 86.8k tokens with context files vs 49.5k without; ~37k are + orchestrator-only managed blocks. No cross-session cache reuse on claude-bridge. +- Forced delegation per question: ~140k weighted tokens (~89k lean); +46–65% tokens and ~3× + wall time on targeted questions with no quality gain. +- Selective (evidence-budget) delegation: cost-neutral (±2%) with ~27k fewer parent tokens. +- Current rules: 0 delegations in controlled runs, even past the 4-file trigger. + +## Scope and constraints + +- Gentle Shell only. Provider-mirrored and byte-pinned contract text stays untouched. +- Children fail safe: if managed blocks cannot be filtered, keep today's behavior. +- Keep writer and verification delegation; keep human consent, RDD, and native authority. +- Orchestrator prompt stays within its 8 KiB budget. +- Preserve untracked `odd/tasks/herdr-shell-notifications.md` and `odd/tasks/tool-argument-watchdog.md`. + +## Working policy for this feature (the rule being implemented) + +- Inline reads only for one parallel batch (≤3 calls, ~≤10k tokens); larger mapping → one explorer + with a ≤2k-token path:line handoff; 2+ non-trivial files → one writer; bounded command output. + +## Mapped facts (explore muo79eso-6-318q) + +- Pi `before_agent_start` exposes `event.systemPromptOptions.contextFiles` ({path, content}[]) as a + per-turn mutable copy; Pi renders from the mutated options. pi-claude-bridge 0.9.0 keeps a + reference to the same options object and rebuilds its appended prompt from its `contextFiles`. +- Children (`GENTLE_PI_AGENTS_CHILD=1`, lib/agents-runner.ts:207/:454) load the gentle-ai + extension; its `before_agent_start` child path (extensions/gentle-ai.ts ~:9356) leaves + contextFiles untouched today. +- Core triggers in assets/orchestrator.md ("4-file rule", "Long-session rule ~20 tool calls, 5 + exploratory reads") are pinned by tests/odd-routing-canonical-ratchet.test.ts, + tests/odd-routing-contract.test.ts, tests/orchestrator-budget.test.ts (8 KiB, tiny headroom). + fixtures/odd-routing-canonical.md is sha-pinned upstream canon: never edited locally. +- Local node_modules was pi-coding-agent 0.87.1 (lockfile 0.99.1); synced with + `pnpm install --frozen-lockfile --ignore-scripts`. The 4 previously failing local test files + now pass (4/9/17/59). + +## Design decisions + +- T1: in the child path of gentle-ai `before_agent_start`, replace + `systemPromptOptions.contextFiles` (same options object) with copies whose orchestrator-only + managed blocks are removed: `orchestrator`, `sdd-orchestrator`, `sdd-model-assignments`, + `agent-routing` (they bind themselves to the orchestrator). Nested blocks not on that list + (e.g. `remote-authorization` inside `agent-routing`) and all other blocks (`codegraph-guidance`, + `engram-protocol`, unknown names) plus all unmanaged project text are kept. Malformed or + unbalanced markers leave that file unchanged (fail safe). Primary sessions untouched. +- T2: Gentle Shell leads; local trigger anchors are updated intentionally (canon fixture + unchanged, divergence tracked by gentle-ai#5139). + +## Tasks + +- [x] T1 — Lean child context: filter gentle-ai managed orchestrator blocks out of delegated + children's context (+ tests, + measured prefix). Commit `302d0950d`. +- [x] T2 — Evidence-budget delegation rule in Gentle Shell prompt assets (+ contract tests). Commit `14bd23568`. +- [ ] T3 — Single PR closing #1587; merge commit after green CI (user-authorized). + +## Checks + +- Focused `node --experimental-strip-types --test ` per task. +- `pnpm run typecheck`, `pnpm run check:runtime-modules`, `node scripts/verify-package-files.mjs`. +- Measured child prefix before/after with a headless delegation (T1). + +## Delivery + +- Strategy: single PR (user decision). Merge commit after green CI and no conflicts (authorized + 2026-09-30). No release. + +## Progress + +- 2026-09-30: branch created; issue #1587 created and approved; mapping delegated + (explore task muo79eso-6-318q, handoff under the ≤2k contract). +- Baseline child explorer prefix: 86.8k tokens (with context files), 49.5k (--no-context-files). +- T1: route delegated direct (writer trigger: new lib module + extension + tests); risk high + (changes what every subagent sees) → writer self-verification + native review (RDD on) + + live measurement of the child prefix. + +- T1 attempt 1 (writer muo7u2od-7-tw8z): lib/child-context-files.ts + hook in gentle-ai.ts; unit + RED/GREEN 23/23; 111/111 focused; typecheck/runtime/package checks pass. LIVE MEASUREMENT FAILED: + real child still 87.7k. Diagnosis: in the isolated Gentle Shell home the launcher injects the + gentle-pi package into the parent only (-e); children load only settings.json packages, so + gentle-ai.ts never runs in the child (no gentle-ai custom entries in child session). Manual + GENTLE_PI_AGENTS_CHILD=1 session with the package: 94.4k → 56.9k, so the filter logic works. +- T1 redesign (continued writer, task muo82w22-8-f4hz): dedicated extensions/child-context.ts + passed to every child via the existing `extensionPaths` → `--extension` mechanism; gentle-ai.ts + hook reverted. Lesson: unit tests did not model how children load extensions; live measurement + is the acceptance check. + +- T1 redesign result: extensions/child-context.ts + gentle-agents `childContextExtensionPaths` + → `--extension` for every child; gentle-ai.ts untouched. RED: module-not-found / missing export; + GREEN 20/20 + wiring 1/1. Writer: 327/327 focused; typecheck no regressions; runtime and package + checks pass. Parent spot check: child-context-files + gentle-agents 180/180. + LIVE: real delegated explorer children start at 56.9k and 57.9k tokens (baseline 86.8k/87.7k), + −30k (−34%) per delegation; answers correct. + Native review: lineage review-0b23f374b16c0bf8, medium, approved, acknowledged, authority burned + (target sha256:8c82a8b1…). Advisory non-blocking: R3-001 WARNING extensions/child-context.ts:11-20, + R3-002 WARNING tests/gentle-agents.test.ts:4311-4312, R3-003/R3-004 SUGGESTION. + Commit `302d0950d` tree d6b83029… equals the reviewed tree. Authored ~500 lines (tests-heavy). + +- T2 (route: delegated direct, writer trigger: 10 files; risk: medium, prompt text + tests): + - Writer muo8calm-9-n5jm. RED: new cross-surface consistency test failed (`missing contract: + **Evidence-budget rule**`); GREEN 193/193 focused. Full unit suite `tests/*.test.ts`: + 4289 tests, 4255 pass, 0 fail, 0 cancelled (34 skipped). typecheck no regressions; runtime and + package checks pass. Rendered orchestrator prompt 7544 B at a 161-char root (budget 8192). + - Canon fixture and AGENTS.md unchanged; ratchet updates local mirror anchors only (comment cites + gentle-ai#5139). tests/package-manifest.test.ts:1621 regex updated (old wording now false). + - Parent spot check: 54/54 (contract, budget, ratchet). + - Native review: lineage review-dc09524924a05bf6, medium, approved. Advisory: R3 WARNING context + backstop threshold at assets/orchestrator.md:56 (the model cannot observe its own context size + directly → follow-up, relates to #1117); two SUGGESTIONs (ratchet semantics, readme not in the + agreement test). + - INCIDENT (parent): acknowledgement and commit were launched in parallel; STATUS saw an empty + working tree (`empty_candidate_base_ref_required`) and the acknowledgement did not run. Recovery: + `git reset --soft HEAD~1` of the unpushed commit (identical tree dac946aa), bound STATUS offered + the exact acknowledge-approved, executed through the facade → authority burned (target + sha256:0c6f1603…), then recommitted with the same message → `14bd23568`, tree dac946aa… = + reviewed tree. Lesson: acknowledge strictly before committing, never in parallel. + - Authored lines: +65 −34 = 99. +- Running authored total: ~600 lines (T1 ~500, tests-heavy; T2 99). + +## Next step + +T3: commit this document, push, open the PR closing #1587, wait for CI, merge commit if green. From 31a04c25294b3560f2483ec655167447d028129f Mon Sep 17 00:00:00 2001 From: Alan Buscaglia Date: Wed, 30 Sep 2026 17:19:39 +0200 Subject: [PATCH 4/5] test(agents): resolve the expected child-context path with fileURLToPath URL.pathname keeps percent-encoding and yields /C:/ on Windows, so the assertion failed in checkouts whose paths contain spaces. Match the production resolution instead. Refs #1587 --- tests/gentle-agents.test.ts | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/tests/gentle-agents.test.ts b/tests/gentle-agents.test.ts index 6ae83effe..db71a9754 100644 --- a/tests/gentle-agents.test.ts +++ b/tests/gentle-agents.test.ts @@ -4,6 +4,7 @@ import { execFileSync } from "node:child_process"; import { homedir, tmpdir } from "node:os"; import { createRequire, syncBuiltinESMExports } from "node:module"; import { dirname, isAbsolute, join, relative, resolve, sep, win32 } from "node:path"; +import { fileURLToPath } from "node:url"; import { pendingReviewMutation, pendingReviewMutationProfiles, REVIEW_REMINDER_RECEIPT } from "../lib/review-reminder-receipt.ts"; import { SESSION_WORKTREE_ENTRY, SESSION_WORKTREE_CHANGED, resolveSessionWorktree } from "../lib/session-worktree-registry.ts"; import { installSessionChangeCapture } from "../lib/session-change-capture.ts"; @@ -4308,7 +4309,7 @@ test("issue #1162: task-mode subagent_run includes question directly in waiting // isolated Gentle Shell home, so every child receives the child-context // extension explicitly through --extension. test("children receive the child-context extension, and a missing file is omitted", async () => { - const expected = join(dirname(new URL(import.meta.url).pathname), "..", "extensions", "child-context.ts"); + const expected = join(dirname(fileURLToPath(import.meta.url)), "..", "extensions", "child-context.ts"); assert.deepEqual(childContextExtensionPaths(), [resolve(expected)]); assert.deepEqual(childContextExtensionPaths(() => false), [], "a missing extension file fails safe to no --extension"); const extensionArguments = (args: string[]) => args.filter((_, index) => args[index - 1] === "--extension"); From 38385236352fcba879ccbf6ce91f9b0f3405a4d4 Mon Sep 17 00:00:00 2001 From: Alan Buscaglia Date: Wed, 30 Sep 2026 17:20:21 +0200 Subject: [PATCH 5/5] docs(odd): record lean delegation context delivery evidence Refs #1587 --- odd/tasks/lean-delegation-context.md | 18 +++++++++++++++++- 1 file changed, 17 insertions(+), 1 deletion(-) diff --git a/odd/tasks/lean-delegation-context.md b/odd/tasks/lean-delegation-context.md index 7cd490286..c9439d891 100644 --- a/odd/tasks/lean-delegation-context.md +++ b/odd/tasks/lean-delegation-context.md @@ -132,6 +132,22 @@ triggers in the Gentle Shell prompt assets with the measured evidence-budget rul - Authored lines: +65 −34 = 99. - Running authored total: ~600 lines (T1 ~500, tests-heavy; T2 99). +- T3: PR https://github.com/Gentleman-Programming/gentle-shell/pull/1590 (`type:feature`). First CI + run on head 73c00efe2: all jobs green; verify `pnpm test` 4289 tests, 4255 pass, 0 fail. + CodeRabbit (🟡 Minor, same spot as native advisory R3-002): the child-context wiring test built + the expected path with `URL.pathname` (percent-encoding, `/C:/` on Windows). RED reproduced in a + worktree under a path with a space (0/1); fixed with `fileURLToPath` → GREEN 1/1, file 160/160, + typecheck no regressions. Native review lineage review-deab0e498746b6ac approved with no + findings, acknowledged BEFORE committing → commit `31a04c252`, tree 375e42c3… = reviewed tree. + +## Follow-ups + +- Context backstop (~150k) relies on the model knowing its context size; consider a mechanical + signal (relates to #1117). +- Gentle AI generated assets and canon: apply the same rule and child-context scoping + (gentle-ai#5139). +- Parent prefix reduction (duplicated managed blocks, tool schemas) tracked in gentle-ai#5139. + ## Next step -T3: commit this document, push, open the PR closing #1587, wait for CI, merge commit if green. +Wait for CI on the final head, merge commit if green and conflict-free.