Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
21 commits
Select commit Hold shift + click to select a range
ad757f4
feat(playbook): route DAG nodes through generated harness
Sep 20, 2026
3086ced
fix(playbook): isolate stored dags from worker tables
Sep 21, 2026
bfb9433
fix(playbook): preserve legacy backend compatibility
Sep 21, 2026
b9ad5bb
Merge branch 'main' into refactor/participant_full_roster
ypflll Sep 21, 2026
4dffdf6
fix(agent): describe DAG worker labels in schema
Sep 21, 2026
cef2a5d
feat(playbook): wire generated participant functions
Sep 20, 2026
f193274
fix(playbook): keep capability loader out of contracts
Sep 20, 2026
f49cdef
test(playbook): cover harness generation boundaries
Sep 20, 2026
b677b59
feat(playbook): unify reusable harnesses and workflows
Sep 20, 2026
9bc71be
fix(playbook): satisfy typed v2 constructors
Sep 20, 2026
8cf973b
test(playbook): cover unified lifecycle in default suite
Sep 20, 2026
3ba58a9
fix(playbook): persist future-facing persona briefs
Sep 21, 2026
7366faf
fix(playbook): bind durable harness to stored workflows
Sep 21, 2026
a9f27e9
fix(playbook): select workers by capability ownership
Sep 21, 2026
60d44bb
Merge branch 'main' into feat/harness_generation_capabilities
ypflll Sep 21, 2026
a779548
fix(playbook): preserve workflow and turn isolation
Sep 21, 2026
53683a7
chore(playbook): sync harness capability base
Sep 21, 2026
6542e44
feat(playbook): unify task and persona generation
Sep 21, 2026
35c2bd0
fix(rpc): sync playbook mode clients
Sep 21, 2026
a59bd43
test(playbook): cover generation safety boundaries
Sep 21, 2026
59afcad
fix(playbook): stabilize generated persona functions
Sep 21, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
30 changes: 20 additions & 10 deletions CONTEXT.md
Original file line number Diff line number Diff line change
Expand Up @@ -1135,6 +1135,13 @@ read; whether a playbook is offered on this machine is config (the
the model is shown, and `raven playbook run` still resolves it), because the file
is the distribution unit and local state must not travel with it.

Schema v2 separates three concepts. A reusable artifact contains a **Harness**, a
**Workflow**, or both: the Harness is the durable worker table bound into a turn;
the Workflow is the validated DAG compiled from one accepted successful run.
A **Run Record** is the per-execution evidence and metadata kept separately from
the reusable artifact; it is neither discoverable nor loaded as a Playbook.
Schema v1 remains supported through the `dag` and `prompt` modes below.

The two modes differ in where the graph comes from, and therefore in who acts on
a load: `dag` ships it as `nodes`, which the engine fills and dispatches through
`SubAgentDagTool.execute` — the same entry a model-composed graph takes, so one
Expand All @@ -1145,16 +1152,19 @@ model in the room, composes with one call of its own. `mode` is the author's
statement of how completely they specified the procedure, and is deliberately not
in the tool signature.

**Discovery is the model's, not a matcher's.** A playbook is reached through
`load_playbook`, one of the tools a turn can use, alongside `spawn` and
`run_subagent_dag` — there is no pre-turn interception and no LLM gate.
`triggers.keywords` decides which playbooks get *described* in that tool when the
library is larger than `playbooks.router.topK`; the `name` enum stays the whole
library, so a retrieval miss leaves a playbook undescribed rather than
unreachable. What the caller may supply is bounded to `params` and `fills`, and a
`fills` entry aimed at a field the playbook already wrote is refused — so a
playbook can be completed but never edited, and the file in git stays an accurate
account of what ran.
**Discovery has two paths.** Normally the main model chooses `load_playbook`
alongside `spawn` and `run_subagent_dag`. When
`playbooks.agentHarness=generate`, one pre-turn setup-model call receives the
ranked saved candidates and may select a direct match. That selection may bind
the saved durable Harness before the main turn and annotates `load_playbook`;
it never executes the stored Workflow. The main model must call
`load_playbook` to execute a Workflow. This resolver is an LLM decision, not a
passive keyword matcher and not an auto-run path.
`match.keywords` on v2 and `triggers.keywords` on v1 rank what is described
when the library exceeds `playbooks.router.topK`; the `name` enum still covers
the whole library. What the caller may supply is bounded to declared parameters
and fills, and a fill aimed at a field the Playbook already wrote is refused, so
a Playbook can be completed but never edited.
_Avoid_: calling `triggers.keywords` a trigger — a keyword makes a playbook
visible, never run. And avoid describing `confirm` as a playbook-level gate: it is
`SubAgentDagSpec.confirm`, a graph-level parameter the playbook's value is
Expand Down
20 changes: 20 additions & 0 deletions i18n/messages.json
Original file line number Diff line number Diff line change
Expand Up @@ -6139,6 +6139,10 @@
"en": "composed per run",
"zh": "流程每次现搭"
},
"gui.pb.shape_harness": {
"en": "Harness · {n}",
"zh": "Harness · {n}"
},
"gui.pb.more_steps": {
"en": "+{n}",
"zh": "+{n}"
Expand Down Expand Up @@ -6203,6 +6207,10 @@
"en": "Graph",
"zh": "图"
},
"gui.pb.tab_harness": {
"en": "Harness",
"zh": "Harness"
},
"gui.pb.tab_assembly": {
"en": "Assembly guide",
"zh": "编排指引"
Expand All @@ -6223,6 +6231,18 @@
"en": "before dispatch",
"zh": "派发前确认"
},
"gui.pb.f_artifact": {
"en": "artifact",
"zh": "产物"
},
"gui.pb.sec_harness": {
"en": "Harness",
"zh": "Harness"
},
"gui.pb.harness_only": {
"en": "This artifact provides reusable workers and has no stored Workflow.",
"zh": "这个产物提供可复用的 workers,不包含已保存的 Workflow。"
},
"gui.pb.sec_keywords": {
"en": "Trigger keywords",
"zh": "触发关键词"
Expand Down
1 change: 1 addition & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -191,6 +191,7 @@ include = [
"raven/providers/data/*.toml",
# Playbook role-pool defaults and the builtin playbook library.
"raven/playbook/*.yaml",
"raven/playbook/*.json",
"raven/playbook/builtin/**/*.md",
# Tracing dashboard viewer (dependency-free Node server + static client),
# a CLI-launched surface, so it lives beside the CLI and not in the kernel.
Expand Down
4 changes: 2 additions & 2 deletions raven/agent/harness/memory.py
Original file line number Diff line number Diff line change
Expand Up @@ -152,9 +152,9 @@ def _briefed(turn: "TurnContext") -> "TurnContext":
from raven.agent.subagent.charter import current_charter

charter = current_charter()
if charter is None or not (charter.prompt or charter.stop_when):
if charter is None or not (charter.task_brief or charter.stop_when):
return turn
return replace(turn, task_brief=charter.prompt, task_done_when=charter.stop_when)
return replace(turn, task_brief=charter.task_brief, task_done_when=charter.stop_when)

async def shrink(
self,
Expand Down
95 changes: 95 additions & 0 deletions raven/agent/harness_capabilities.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,95 @@
"""Load the checked-in capability switches for generated worker harnesses."""

from __future__ import annotations

import json
from pathlib import Path
from typing import Any, Literal

ModuleName = Literal["memory", "planning", "capability", "action"]
FunctionKind = Literal["self", "participant"]

_PATH = Path(__file__).parents[1] / "playbook" / "harness_generation.json"
_CONNECTED_PARAMETERS = frozenset(
{
("memory", "systemPrompt"),
("memory", "stopWhen"),
("capability", "tools"),
("action", "checks"),
}
)
_CONNECTED_FUNCTIONS = frozenset(
{
("memory", "participant", "intake"),
("planning", "participant", "advise"),
("action", "participant", "judge"),
("action", "participant", "salvage"),
}
)
_CACHE_KEY: tuple[Path, int, int] | None = None
_CACHE_ROOT: dict[str, Any] | None = None


def _entry_enabled(entry: Any, path: str) -> bool:
if not isinstance(entry, dict) or type(entry.get("enabled")) is not bool:
raise RuntimeError(f"invalid harness generation capability file: {path}.enabled must be a boolean")
return entry["enabled"]


def _validate_connected(root: dict[str, Any]) -> None:
for module, body in root.items():
if module == "note" or not isinstance(body, dict):
continue
parameters = body.get("parameters", {}).get("items", {})
for name, entry in parameters.items():
path = f"{module}.parameters.{name}"
if _entry_enabled(entry, path) and (module, name) not in _CONNECTED_PARAMETERS:
raise RuntimeError(f"harness generation capability is enabled but not wired: {path}")
functions = body.get("functions", {})
for kind in ("self", "participant"):
for name, entry in functions.get(kind, {}).get("items", {}).items():
path = f"{module}.functions.{kind}.{name}"
if _entry_enabled(entry, path) and (module, kind, name) not in _CONNECTED_FUNCTIONS:
raise RuntimeError(f"harness generation capability is enabled but not wired: {path}")


def _root() -> dict[str, Any]:
global _CACHE_KEY, _CACHE_ROOT
try:
stat = _PATH.stat()
cache_key = (_PATH, stat.st_mtime_ns, stat.st_size)
if _CACHE_KEY == cache_key and _CACHE_ROOT is not None:
return _CACHE_ROOT
document = json.loads(_PATH.read_text(encoding="utf-8"))
root = document["harnessGeneration"]
except (OSError, json.JSONDecodeError, KeyError, TypeError) as exc:
raise RuntimeError(f"invalid harness generation capability file: {exc}") from exc
if not isinstance(root, dict):
raise RuntimeError("invalid harness generation capability file: harnessGeneration must be an object")
_validate_connected(root)
_CACHE_KEY = cache_key
_CACHE_ROOT = root
return root


def parameter_enabled(module: ModuleName, name: str) -> bool:
"""Whether the generator may emit one module parameter."""
root = _root()
try:
entry = root[module]["parameters"]["items"][name]
except (KeyError, TypeError) as exc:
raise RuntimeError(f"harness generation capability is not declared: {module}.parameters.{name}") from exc
return _entry_enabled(entry, f"{module}.parameters.{name}")


def function_enabled(module: ModuleName, kind: FunctionKind, name: str) -> bool:
"""Whether the generator may emit one module function."""
root = _root()
try:
entry = root[module]["functions"][kind]["items"][name]
except (KeyError, TypeError) as exc:
raise RuntimeError(f"harness generation capability is not declared: {module}.functions.{kind}.{name}") from exc
return _entry_enabled(entry, f"{module}.functions.{kind}.{name}")


__all__ = ["function_enabled", "parameter_enabled"]
54 changes: 52 additions & 2 deletions raven/agent/loop/main.py
Original file line number Diff line number Diff line change
Expand Up @@ -888,6 +888,7 @@ async def run_turn(
"""
session_key = req.conversation or f"{req.source.channel}:{req.source.chat_id}"
flush = True
capture = None
try:
# Pick up a mid-session `deep-research enable` BEFORE the freeze below
# captures the turn's pairs. The promotion re-registers the offer
Expand Down Expand Up @@ -915,23 +916,56 @@ async def run_turn(
# forbids: a turn resolves its pair once and holds it for the whole
# turn tree.
binding = self.binding_for_session(session_key)
delegate_table = await self._write_worker_table(req, session_key, binding)
resolution = await self._resolve_playbook_turn(req, session_key, binding)
delegate_table = resolution.table
if resolution.active:
from raven.playbook.run_record import PlaybookRunCapture

capture = PlaybookRunCapture(
query=getattr(req, "text", "") or "",
disposition=resolution.disposition,
selected_playbook=resolution.selected_playbook,
artifact_name=resolution.artifact_name,
capture_workflow=resolution.capture_workflow,
)
from raven.playbook.run_record import capture_scope

# The charter a dispatch staged for this session, taken for this turn
# only. Both scopes below are None on an ordinary turn, which is the
# path every reader answers to as "no playbook".
charter = self._take_session_charter(session_key)
if resolution.disposition == "artifact" and resolution.artifact_name:
from raven.agent.subagent.charter import Charter

artifact_status = (
f"has generated and saved the reusable Persona Harness {resolution.artifact_name!r}"
if resolution.persisted
else (
f"generated the Persona Harness {resolution.artifact_name!r} for this turn, "
"but persistence failed"
)
)
charter = Charter(
prompt=(
f"The platform {artifact_status}. Do not call load_playbook, search for a persona "
"format, or write Playbook, skill, persona, Harness, Workflow, or run-record files. "
"Do not spawn workers or execute a Workflow. Reply concisely with what the generated "
"Harness is for and whether it was saved."
)
)
with (
use_binding(binding),
self.tools.session_scope_for(session_key),
self.tools.turn_scope(),
delegate_scope(delegate_table),
charter_scope(charter),
capture_scope(capture),
# The participant seats in the hook chain ask this turn's modules,
# so a replaced Action or Planning decides what a plugin's
# judgement does -- bound per turn like the model.
bind_harness(self.harness),
):
return await self._run_turn(
outcome = await self._run_turn(
req,
emit,
drain,
Expand All @@ -940,10 +974,26 @@ async def run_turn(
usage_sink=usage_sink,
text_sink=text_sink,
)
if capture is not None:
await self._finish_playbook_turn(resolution, capture, binding)
return outcome
except asyncio.CancelledError:
flush = False
raise
finally:
if capture is not None and capture.status == "running" and self._playbooks is not None:
from raven.playbook.run_record import RunRecordStore

capture.finish(status="cancelled" if not flush else "failed")
try:
RunRecordStore(self._playbooks.store.root).save(capture)
except Exception: # noqa: BLE001 - cleanup cannot replace the turn's error
logger.opt(exception=True).warning("playbook: failed turn record could not be saved")
loader = self.tools.get("load_playbook")
setter = getattr(loader, "set_preselected", None)
if callable(setter):
setter(None)

# Every way a turn ends passes here, which is what makes this the
# place a foreground DAG run learns its turn is over. A direct chat
# runs on an instance's own lane, concurrently with the main agent's
Expand Down
35 changes: 33 additions & 2 deletions raven/agent/loop/turn_path.py
Original file line number Diff line number Diff line change
Expand Up @@ -1788,12 +1788,40 @@ async def _process_message( # noqa: C901 (cc 47: pre-existing, above the ceilin
# (content, media) reply directly.
#
# Skip the user-inbound hooks for Sentinel / subagent turns (by origin).
inbound_original = content
skip_user_inbound = origin in _SKIP_USER_INBOUND_ORIGINS
if skip_user_inbound:
from raven.agent.harness.participants import Intake, read_intake
from raven.agent.subagent.charter import charter_participants
from raven.contracts.participant import StepView

participants = charter_participants()
if participants:
peeked = self.sessions.peek(msg_session_key)
step = StepView(
session_key=msg_session_key,
iteration=0,
response=None,
transcript=(),
history=tuple(peeked.messages if peeked is not None else ()),
turn_base=0,
question=content,
rollbacks=0,
mode=None,
mode_overlay=None,
phase="user_inbound",
)
answer = await self.harness.memory.ask_intake(content, step, participants)
intake = answer if answer is None or isinstance(answer, Intake) else read_intake(answer, text=content)
if intake is not None:
if intake.reply is not None:
return str(intake.reply), []
content = intake.text

# The turn's one hook-metadata dict, and the words the user actually
# sent: the first crosses every phase group with the turn, the second
# is what history keeps no matter how hooks rewrite the model's view.
turn_hook_meta: dict[str, Any] = {}
inbound_original = content
if len(self.hooks) > 0 and not skip_user_inbound:
_peeked = self.sessions.peek(msg_session_key)
_hook_ctx = AgentHookContext(
Expand All @@ -1805,7 +1833,10 @@ async def _process_message( # noqa: C901 (cc 47: pre-existing, above the ceilin
)
_decision = await self.hooks.before_user_inbound(_hook_ctx)
if _decision.short_circuit_result is not None:
return _decision.short_circuit_result
result = _decision.short_circuit_result
if isinstance(result, tuple) and len(result) == 2:
return result
return str(result), []
if _decision.modified_content is not None:
content = _decision.modified_content

Expand Down
Loading
Loading