Repository navigation
Conversation
…ads (#463) ## Summary `main` is red on three jobs (`python lint`, `unit 2/4`, `unit 3/4`), and has been since #451 merged. All of it is fallout from one arrival: `image_search` came with a second switch beside its credential. `tools.web.search.images` decides whether the tool is registered at all (`raven/agent/loop/wiring.py`), so a lane that never places a picture keeps the tool face it had. The capability table was not told. `is_configured` answers on the vendor key alone, so with a Serper key and the switch left off the table ticked a capability the agent does not hold, and the two tests that drive a real `AgentLoop` caught exactly the disagreement that module exists to prevent. The switch belongs on `is_disabled`, not `is_configured`. That function's own docstring already names the reason: a switched-off tool usually has its credential set, and calling it unconfigured sends the deployer to set a key that is already there. `tools.disabledTools` and this flag are the same kind of answer, so `is_offered` needs no change and the doctor row reads "off" rather than "missing a key". The test configs that mean "everything supplied" now turn the switch on, the way `raven/core/runtime.py` passes it from the config, and a new case pins the combination that was wrong: key present, pictures off, nothing offered, table saying off. Two more pieces of the same arrival: - `ImageSearchTool` copied `WebSearchTool`'s key resolution without its two attribute annotations, so the untyped lambda `wiring.py` hands it left the `api_key` property returning `~AlwaysFalsy | str` against a declared `str`. The type checker pinned in #416 reports it; the annotations are the fix. - `WebSearchConfig` gained the `images` field, so the writer's validated subtree now carries it and two `test_config_update_tools` expectations that pin the written section by value went stale. Not changed here: `get_web_search` still reports `{provider, api_key, max_results}` and does not surface `images`, though `set_web_search` can write it. That is a gap in the read side of the config API, not a CI failure, and it wants its own change. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Run on Windows 11 in a clean worktree cut from `main` (`5a29ea22`): - `uv run pytest tests/test_tool_capabilities.py tests/test_config_update_tools.py -q`: 71 passed. All four tests failing on `main` now pass, plus the new case. - `uv run --frozen --python 3.12 --all-extras ty check raven evolver agents plugins-dist scripts`: the `web.py:516` diagnostic is gone. The diagnostics that remain locally are unresolved imports for optional packages absent from this machine's environment (`boxlite`, `nio`); CI reported exactly one diagnostic before this change, the one fixed here. - `uv run --frozen --python 3.12 --extra dev ruff check raven tests`: all checks passed. - `ruff format --check` on the four changed files: already formatted. - `uv run pytest tests/test_cli_doctor_commands.py tests/test_rpc_console.py -q`, which consume `is_disabled`: the failing-test set is byte-identical before and after this change (8 pre-existing Windows-only failures around file modes and an absent browser package, none of them in the set CI reports). ## Risk `is_disabled` gains a second reason to answer true, so a deployment with a search key and `tools.web.search.images` off now reads as "switched off" in `raven doctor` and the console capability report, where it previously read as offered. That is the state the loop was already in; only the description changes. Nothing about registration, gating, or the tool's runtime behaviour moves. The other two changes are an annotation and two test expectations, with no runtime effect. Rollback is reverting this commit; it touches nothing that other work builds on. ## Related Issues N/A --------- Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
… busy runner (#466) ## Summary Since the idle ceiling became a gate (#428), a unit shard has failed on `main` and on most pull requests with every test passing: the ceiling reads wall clock minus CPU and calls the difference waiting, and a thread that is runnable and not on a CPU spends it the same way. With four workers, the controller and coverage on a four-vCPU runner that queueing reached three to four seconds inside one long test, so the shard died on whichever test was executing when contention peaked, a different one each run. Warming the catalogue (#452) did not stop it. #449 halves the workers per shard, which pays for the same measurement with every shard's wall clock. This changes the measurement instead. Idle now subtracts the time the test's threads spent queued for a CPU, read from the scheduler's own account in `/proc/self/task/<tid>/schedstat`: - every thread of the process is read, not only the one running the test, because the rpc model handlers do their catalogue reads on `asyncio.to_thread`; - a thread writes its own last reading on the way out, through a wrapper on `Thread._bootstrap_inner`, because the per-test event loop is closed at the end of the test and takes its executor threads with it, and `/proc` forgets a thread the moment it ends; - CPU stays on `os.times`, which keeps the CPU of ended threads and of reaped children; the scheduler's per-thread run time would have been lost the same way the queue time was. A sleeping thread is not on a run queue, so a test that waits on purpose is still charged the whole wait. Without `schedstat` (macOS, Windows) the arithmetic is the old one; the strict gate only runs on Linux, and a Linux-only test fails loudly if the runner's kernel stops keeping the statistics rather than letting the gate degrade quietly. `-n 4`, the 3 s ceiling and the four shards are unchanged. If this lands, #449 is not needed. ## Type - [ ] Fix - [ ] Feature - [ ] Docs - [x] CI / tooling - [ ] Refactor - [ ] Other ## Verification Unit tests for the clocks (12 on macOS, 13 on Linux, the last one reads the live kernel): ``` uv run --frozen --all-extras pytest -q tests/test_conftest_idle_ceiling.py macOS: 12 passed, 1 skipped Linux container: 13 passed ``` Full suite on macOS, where the fallback arithmetic runs: ``` uv run --frozen --all-extras pytest -q -p no:cacheprovider 23137 passed, 109 skipped in 291.30s ``` Contrast runs in a Linux container (Docker, linux/arm64, kernel 6.12), the checkout and venv on container-local storage, `main` at 2bdc4bd against this branch. Four workers pinned to two CPUs stands in for the four-vCPU runner; the ceiling is lowered to 0.05 s so the effect shows on one file: ``` docker run --cpuset-cpus=0-1 ... pytest tests/test_rpc_model.py --idle-ceiling 0.05 -n 4 main: 70 unmarked tests over the ceiling, worst 1.67s idle of 2.89s branch: 2 unmarked tests over the ceiling, worst 0.12s idle of 0.50s (0.19s queued) docker run --cpuset-cpus=0-3 ... pytest tests/test_rpc_model.py --idle-ceiling 0.05 -n 16 main: 81 over, worst 6.11s idle of 8.16s branch: 1 over, 0.13s idle of 1.11s (0.73s queued) an unmarked test that does time.sleep(0.5), on the branch: 0.53s idle of 1.70s (0.00s queued) # an honest wait is still charged in full ``` The CI shard command, same container, four workers pinned to two CPUs: ``` pytest -q --shard 3/4 --durations=25 --idle-ceiling-strict -n 4 --cov=raven --cov=raven_everos --cov-branch --cov-report= main: 5131 passed, 10 skipped; "1 unmarked test(s) waited more than 3s (failing the session)" 3.11s idle of 6.04s tests/test_ppt_engine_template_tool.py::test_binding_a_template_by_path; exit 1 branch: 5378 passed, 11 skipped; no idle report; exit 0 branch, shards 1/4, 2/4, 4/4: no idle report in any (three failures, all `git ls-files` exit 128 because the copied worktree's gitdir is not inside the container) ``` - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed (none; the option help text and the summary line say what is now subtracted) ## Risk - Behaviour change is confined to what the ceiling charges a test: less on a busy runner, the same for a real wait. Nothing under `raven/` changes. - `Thread._bootstrap_inner` is private CPython API (present from 3.8 through 3.13); the wrapper is installed once, guarded against a second import of the module, and pinned by a test. If a future interpreter drops it the module fails to import at collection, which is loud. - Rollback: revert the commit; the gate goes back to wall clock minus CPU. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues #428 made the ceiling a gate; #452 warmed the catalogue; #449 halves the workers instead. Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
…458) ## Summary `os.kill(pid, 0)` is the POSIX existence probe. On Windows, CPython implements `os.kill` as `TerminateProcess`, so signal 0 is not a probe there at all, and it fails in two directions depending on what the caller may do to the target: - a process the caller can terminate, which is any child of it, is killed and the call returns cleanly, so the probe reports "alive" about a process it has just destroyed; - a process it cannot, such as the detached `raven web` supervisor, raises `OSError` with `ERROR_INVALID_PARAMETER`, so the probe reports "dead" about a process that is running. The second is why `raven web --stop` did nothing on Windows. `_read_web_state` and `_read_serve_pid` both answered `None`, so `_stop_resident` signalled neither the supervisor nor the gateway, and printed "nothing to stop" while the engine went on serving the page. `raven/utils/pid.py` holds the platform question now. Windows asks `OpenProcess` plus a zero-timeout `WaitForSingleObject`, the pair `raven/updates/upgrade.py` already uses to watch a parent exit: a handle that fails to open for any reason but `ERROR_ACCESS_DENIED` is a dead process, and one that opens and then times out is a live one. The wait rather than `GetExitCodeProcess`, because a process whose real exit code is 259 cannot be told apart from `STILL_ACTIVE`. `_pid_alive` keeps its name in `serve_commands`, since `_gateway_page` imports it and the stop tests patch it; only the answer moves. Four other call sites carry the same probe and are left for their own change. The sharpest is `raven/agent/tools/background_exec.py`, which reaches it on the Windows branch specifically, against the agent's own background tasks. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Run on Windows 11, in a clean worktree cut from `main`: - `uv run pytest tests/test_utils_pid.py tests/test_cli_serve_commands.py -q`: 77 passed, 3 failed. - the same 3 failures reproduce on `main` without this change (`test_the_state_file_stays_owner_only`, `test_sigterm_removes_the_state_file`, `test_sigterm_during_the_prologue_still_removes_the_state_file`): they are pre-existing Windows gaps in the supervisor tests, unrelated to this PR, and this branch adds no new failure. Baseline on `main` is 73 passed, 3 failed; the 4 extra passes are the new `tests/test_utils_pid.py`. The new test drives real subprocesses. Survival is asserted by a wait that must time out, not by a poll taken straight after the probe: termination is asynchronous, and the immediate read still says "running" for a process already on its way down. Forcing the POSIX branch on every platform fails two of the four cases. - [x] Relevant tests pass locally - [ ] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed ## Risk POSIX behaviour is unchanged: the helper keeps calling `os.kill(pid, 0)` there. On Windows, `raven web --stop` starts working and stops silently terminating child processes it only meant to look at, which is the point of the change. Rollback is reverting to the inline `os.kill(pid, 0)` in `raven/cli/serve_commands.py` and dropping `raven/utils/pid.py`. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A --------- Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
…es (#457) ## Summary A Windows checkout of this repo with the default `core.autocrlf=true` rewrites every text file to CRLF, and two byte-exact gates in `ui-web` then fail on a tree that is byte-identical to the one CI passes on Linux: - the boot-order and open-conversation node tests slice their source on `"\n}\n"`, which never matches a CRLF file, so the sliced function comes out empty and 9 tests fail with `ReferenceError`; - `gen-rpc-client --check` compares its own LF output against the file on disk and reports permanent drift in `src/rpc/generated.ts`. `.gitattributes` pins the checkout filter so the working tree matches what the repository already stores. The repository stores LF today, so this adds no renormalization diff: `git add --renormalize .` over the whole tree stages nothing beyond the two files in this PR. Only the checkout filter changes. `ui-web/build.py` is the other half of the same defect. It assembled the page under Python's default newline translation, so `dist/index.html` came out CRLF on Windows even from an LF source tree, and `check-page.mjs` could no longer extract the script payloads (its `<script>\n` anchor needs the LF to be the next byte). That artifact ships inside the wheel, so it must not depend on which platform assembled it. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Run on Windows 11 with `core.autocrlf=true`, in a clean worktree cut from `main`: - `git add --renormalize .` after the change stages nothing (confirms no renormalization diff); - `file ui-web/src/rpc/generated.ts` reports CRLF on `main` and LF after a forced re-checkout on this branch; - `npm run gen:check` (ui-web): `generated.ts matches the contract (178 methods)`; - `npx vitest run` (ui-web): 108 files, 1836 tests passed; - `npm run --prefix ui-web build && python ui-web/build.py`: `dist/index.html` is LF; - `node ui-web/scripts/check-page.mjs`: `OK (1,482,975 bytes, 3 scripts parse)`; - `node ui-web/scripts/count-shared-globals.mjs`, `check-css.mjs`, `check-class-namespace.mjs`: all OK. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed ## Risk No runtime behaviour changes for users on Linux or macOS, where the checkout already produced LF. On Windows the working tree stops being rewritten to CRLF, so a contributor with existing local checkouts may see one large renormalization touch on the next checkout; the committed bytes do not move. The built `dist/index.html` is now LF on every platform, which is what the Linux-built wheel already shipped. Rollback is deleting `.gitattributes` and dropping the `newline=""` argument in `ui-web/build.py`. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A --------- Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
…here (#464) ## What A skill installed through the market had no way to be removed from the page, because `ext.list` decided "was this installed from the hub" by looking for one marker file, and the installer most skills arrive by writes a different one. Closes #456. Two installers put hub skills on disk: | | market module (`raven/skill_hub/hub.py`) | context engine + `use_skill` (`raven/agent/tools/skill_hub.py`, `raven/context_engine/segments/skills.py`) | | --- | --- | --- | | lands at | `<skills>/<name>/` | `<skills>/hub/<slug>@<version>/<name>/` | | stamps | `.skillhub.json` | `.install-meta.json` | `ext.list` tested only for the first, so the second -- the ordinary path -- reported `hub: false`. The page gates its removal control on that flag, so it appeared for the exception and never for the rule. On the machine this was found on, three of four installed skills were the second kind; the one that showed a remove button was the one that came through the market module. ## How Backend only; nothing under `ui-web/` or `ui-tui/` changes. The page already renders the control when `hub` is true, and already skips the market-detail fetch when `hub_id` is empty, so both halves light up on their own once the flag is right. - `raven/skill_hub/audit.py` names the stamp it writes (`INSTALL_META`) so the two readers test for the file the installer wrote rather than a spelling of their own. - `raven/rpc/methods/console.py` reports `hub: true` when either stamp is present. The bundle stamp carries a slug, not a market id, so that skill has an empty `hub_id` -- a state the page already handles. - `raven/skill_hub/hub.py` lets `skillhub.remove` find a bundle by the skill name inside it, resolving the wrapper folder the same way `SkillHubClient._bundle_root` did at install time, and removes the bundle whole: the bundle is the install unit, and the CLI's `skill remove` deletes the same folder. Only a stamped directory is eligible -- a folder placed under `hub/` by hand is left alone, which is the rule `MARKER` already enforced on the other layout. The second change is not optional given the first: reporting `hub: true` without it would put a remove button on a skill the handler then refuses, which is worse than no button. ## Verification - `test_ext_list_calls_a_skill_hub_installed_whichever_installer_stamped_it`: three skills, one per stamp and one with neither, asserting `(hub, hub_id)` for each -- `(True, "uuid-market")`, `(True, "")`, `(False, "")`. The marker names are pinned as literals in the test so a drift between writer and reader is a failure, not a pass. - `test_remove_reaches_a_bundle_the_other_installer_cached`: a stamped bundle is removed whole by its skill name; an unstamped folder under `hub/` is refused and left in place, and so is the neighbouring bundle. - Mutation-checked: dropping the second stamp from `ext.list` fails exactly the first test; dropping the bundle lookup from `remove` fails exactly the second. 139 passed in the two suites on the unmutated tree. - Dry-run against the real workspace: all four installed skills now report `hub: true`; the three bundles resolve to their `<slug>@v0` folders by skill name; the market-module skill and an unknown name resolve to none. - Ruff check and format clean. ## Not done here, on purpose Removal does not stop the context engine re-installing the skill on its next catalogue hit; the CLI documents the same limit and points at `skill block`. Whether the page should offer that beside removal is a UI question and out of scope while `ui-web/` is frozen. Also left as they were: the two installers keeping two layouts and two stamps -- this makes the readers agree with both writers rather than deciding which writer should win. Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com> --------- Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com> Co-authored-by: KT <74288668+0xKT@users.noreply.github.com>
… group (#455) ## Summary A conversation started in the TUI was listed by the served page and then could not be opened: clicking it produced an empty transcript. `SessionManager.session_path` looks for a transcript in this process's own group, and falls back to searching the other groups when it is not there. That fallback was gated on `project_slug`, so only a surface that groups by launch directory ever ran it. The gateway groups by channel and never did -- yet `list_sessions` scans every group, so the page offered a conversation filed under a launch directory's slug and then looked for it under `sessions/tui/`. `session.resume` missed the file, fell through to its fresh-mint branch, and handed the reader an empty transcript under a new id. The write verbs resolve through the same function, so the miss was not only a failed read. A rename, a pin or a turn from the page filed a second transcript under the channel group, splitting one conversation across two files: ``` sessions/-srv-alpha/abc.jsonl -> 2 messages sessions/tui/abc.jsonl -> 1 messages (written by the gateway) ``` Ungating the search is necessary but not sufficient, and the second half is the part worth reviewing. The glob matches on the chat_id, and a chat_id names no channel: `cli:<id>` and `tui:<id>` are two conversations, so an ungated search let a `tui:` save adopt the `cli:` file and append to it. A candidate is now taken only once its own metadata key says it is this session. A transcript carrying no key predates that field and is still adopted on the stem alone, which is the case the search was originally added for. Matching on the stored key rather than on the directory name is also what makes this correct on Windows. `key_from_path` tells a slug from a channel by a leading `-`, which holds only for POSIX paths; a Windows launch directory slugs to `C--Windows-system32`, with no leading separator to read. Root cause was confirmed against a control rather than by inspection: two `tui:` sessions, same handler, same manager, differing only in which group holds the file. The one under `sessions/tui/` resumed with its 26 messages; the one under a slug minted a fresh empty session. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification ``` uv run pytest tests/test_session_project_grouping.py -q 2 failed, 32 passed uv run pytest tests/test_cli_session_commands.py -q 43 passed uv run ruff check raven/session/manager.py tests/test_session_project_grouping.py All checks passed! uv run ruff format --check <same two files> 2 files already formatted uv run ty check raven/session/manager.py All checks passed! ``` The 2 failures in `test_session_project_grouping.py` are present on unmodified `main` in this Windows checkout and are not touched by this change: `test_project_dir_is_not_rewritten_by_a_later_run_elsewhere` and `test_slug_collides_for_a_separator_and_an_underscore` both build fixtures from POSIX paths such as `/srv/alpha`, which `Path` renders as `\srv\alpha` on Windows. Regression scope was measured rather than asserted. Running `pytest tests/ -k "session or transcript or workdir or resume"` (about 1085 tests) with and without the change gives **372 failures before and 372 after, with identical failure sets** -- no test changes state in either direction. The same comparison over the session-adjacent files alone gives 156 before and 156 after. Five tests were added to `tests/test_session_project_grouping.py`, extending the existing file rather than adding a new one. Two of them fail without the fix and pass with it: - `test_the_gateway_opens_a_session_filed_under_a_project_slug` -- the reported bug; - `test_the_gateway_appends_to_the_transcript_it_opened` -- the split-file half; - `test_two_channels_sharing_a_chat_id_keep_their_own_transcripts` -- the guard against over-correcting, and the case that caught the first attempt at this fix; - `test_a_keyless_transcript_is_still_adopted_across_groups` -- pre-key transcripts stay reachable; - `test_the_gateway_still_files_its_own_sessions_by_channel` -- reaching across groups is a fallback for an existing transcript, not a change of where this process writes. The first attempt, which ungated the search without the key check, broke `test_cross_channel_ambiguous_same_id_two_channels` in `tests/test_cli_session_commands.py`. That failure is what produced the key check. Fix verified against the real workspace that produced the report: the affected session now opens with both of its messages, resolving to its actual file under the launch directory's slug. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed No docs change: this corrects the behaviour the surrounding docstrings already described, and coins no domain term. ## Risk User-visible change: a surface can now open a transcript filed under another group. On the gateway this closes a gap rather than widening reach -- `list_sessions` already scans every group, so those sessions were already listed and offered; only opening them was broken. For a slug-grouped surface the search is unchanged apart from the new key check, which makes it strictly stricter: a candidate that would previously have been adopted on its stem alone is now refused when its stored key names a different session. Transcripts written before the metadata `key` field are unaffected -- they answer no key and are still adopted on the stem. `session_path` now reads the first line of each candidate when the direct path is absent. The read is one `readline` on a file the caller is deciding whether to open, and the candidate set for an exact chat_id is normally empty or a single file. Rollback is a revert of the single commit; the change is confined to one function and one new module-private helper, and writes nothing to disk that a revert would strand. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A --------- Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com> Co-authored-by: KT <74288668+0xKT@users.noreply.github.com>
…t moved on (#476) ## Summary The installers stop telling you what to run and run it. `install.sh` ends with `raven web --stop && raven web --foreground`, so a finished install is the page open in a browser rather than a hint to go and type something; first-run setup is walked through by the page's own onboarding. The capability summary, the upgrade-versus-first-run wording and the config.json probe that fed it go with it. `install.ps1` mirrors all of this, with one deliberate divergence: a non-zero page exit only warns there, because under `irm | iex` that is the reader's own interactive PowerShell and Ctrl-C, the ordinary way to end a foreground page, returns non-zero. Two gaps that only bite a source install are closed on the way: - `build_web_assets` compares each artifact against its sources instead of only testing that it exists. An editable install relinks Python and nothing else, so `ui-tui/dist/entry.js` and `ui-web/dist/index.html` kept serving whatever the tree held at first install while every `.py` beside them moved on. The page watches `i18n/` too, since `ui-web/build.py` inlines the catalogue. `node_modules`, `dist` and `.modern` are pruned: `npm ci` rewrites the first on every build, and the other two hold the artifact being judged. - A node packaged without npm now fetches a private runtime that carries one. Debian and Ubuntu split the two packages, so a system node >= 22 satisfied `ensure_node` and the build then found no npm with no private runtime ever provisioned. The download body is split out as `provision_private_node` and called from the build probe as well, subshelled so a failure there is a skipped build rather than a failed install. It stays lazy rather than moving into `ensure_node`, so a wheel install that never runs npm does not pay for a 50 MB download it has no use for. Both READMEs are reordered and the clone install is documented under Quick Start, where it was missing: `./install.sh` and a piped run do different things from inside a checkout, because local mode requires `$0` to be a real file, and `RAVEN_LOCAL_SRC` is the opt-out. The Chinese README also picks up the EverMind ecosystem table the English one had already moved to. One test changed rather than being deleted. `test_readme_quickstart_matches_the_installer_hint` required an `Onboard and run` heading in both READMEs; the installers no longer close with a first-run hint and neither README carries that section, so the pin had no subject left. It is repointed at what the two must still agree on, which is the clone install. `install.ps1` has not been parsed or executed. There is no pwsh on the machine this was written on and no PowerShell step in CI, so every Windows change here is reviewed by eye and pinned by text tripwires only. A real Windows run before release would be worth it. ## Type - [ ] Fix - [x] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification `uv run python -m pytest tests/ --ignore=tests/integration` on the merged tree: 23161 passed, 109 skipped. `uv run pytest tests/test_install_script.py`: 27 passed, including a functional test that extracts the real `is_stale` and `resolve_node_dir` from `install.sh` and runs them under `sh`. `is_stale` is checked across missing artifact, fresh artifact, touched page source, touched catalogue, and build output touched; `resolve_node_dir` across system node with npm, node without npm, node without npm and a failed fetch, and no node at all. Both were mutation-checked: removing the fallback line from `install.sh` turns the new tests red. `sh -n install.sh`, `make check-source-language`, `make check-large-files`, `ruff check`, `ruff format --check`: all clean. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed ## Risk User-visible: a finished install now holds the terminal on a running page instead of returning to the prompt. That also means the installer's exit code is the gateway's, so on a clone where the page build only warned, `install.sh` now exits 1 with "No page is built" where it used to exit 0 with a working TUI. A re-run over a checkout that has moved on will rebuild the frontend, which costs an `npm ci` it used to skip. Rollback is per commit: the launch, the rebuild-on-change and the npm fallback are independent edits to `install.sh` and `install.ps1` and can be reverted separately. No runtime code outside the installers changed. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A --------- Co-authored-by: admin <admin@SH-HuKai.local> Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
## Summary Turn "bug fix to regression case" from a manual chore into a gated pipeline, and give trajectory regressions their own CI gate. - `raven trajectory replay --json`: stable machine-readable report (schema_version 1) with defensive JSON coercion; on a halted replay the full JSON is still emitted before exit code 2. - Case metadata contract: `case.yaml` requires non-blank issue/owner/why/re_record; residual-scan findings pass only with an explicit per-token human sign-off (`reviewed_residuals`: full-token sha256 plus a verbatim reason, nothing exempted automatically -- a machine-produced redaction report is not approval). - `raven trajectory regression validate [CASE_DIR|--all]`: static commit gate -- whole-directory discovery so a broken case cannot hide, both schemas, cassette completeness down to the replay contract (valid JSON is not a usable payload: missing payloads via the extracted replayability check, corrupt inner fields via `validate_recording`), rejection of files the residual scanner would silently skip, and a 256 KiB / 1 MiB size budget; every parse boundary collects problems instead of raising, so one malformed case cannot abort the --all sweep. - `raven trajectory regression init <source> --name <case>`: scaffold from a bundle directory, attempt id, trajectory report tarball, or bug report package (two-layer extraction; both passes accept only regular files/directories with root-contained names). Minimize lands in staging; residual findings are reviewed interactively (the full token is shown locally, files receive digest and reason); publication copies onto the destination filesystem under a `.init-` staging name, re-checks the case name, then renames on one device -- failures clean up only what the run created. - Vacuous-pass guard (`MIN_COMMITTED_CASES`, raised by hand) and a `trajectory` CI job: replay the committed cases, run `validate --all` even after a pytest failure so both gates report, upload per-case failure reports (replay --json schema plus check failures, mode from each case's expect.yaml) as a 14-day artifact. - Docs: the Trajectory Regression Case entry in CONTEXT.md extended in place; `tests/trajectories/README.md` documents the per-case layout and workflow. Also: `config_secrets_loaded` (a field name of redaction.json itself, a guaranteed self-referential false positive in every cassette) joins the residual scanner's benign literals, and both sample cases gain real `case.yaml` metadata. ## Type - [ ] Fix - [x] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification - `uv run pytest tests/ -q` (after rebasing onto latest main): 23149 passed; the 37 failures and 20 errors are environment-bound and reproduced identically on a pristine origin/main worktree without these changes. - `uv run pytest tests/test_trajectory_regressions.py tests/test_cli_trajectory_regression_commands.py tests/test_cli_trajectory_commands.py tests/test_trajectory_replay.py tests/test_trajectory_cassette.py tests/test_trajectory_redact.py -q`: 282 passed (about 90 new tests across the five steps). - `uv run ruff check .`, `uv run ruff format --check .`, `make lint-imports`, `make lint-deps`, `make check-large-files`, `make check-source-language`: pass. `make lint-types` reports one diagnostic in `raven/agent/tools/web.py`, a file this branch does not touch, pre-existing on main. - CI workflow YAML parses (`yaml.safe_load`); the first real run of the new `trajectory` job happens on this PR (see this PR's checks). - CLI assertion width-independence re-verified with `COLUMNS=60`. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed ## Risk - All additive: the default `replay` output path is unchanged; `regression init/validate` are new subcommands; the one scanner change adds a fixed benign literal (any variation still flags). - The new `trajectory` CI job becomes a required-passing gate for PRs; to lift it in an emergency, revert the CI commit (`ci: add trajectory regression gate job and case-count guard`) -- the remaining commits are independent. - Tar extraction accepts only regular files and directories with root-contained relative names on both layers, with `filter="data"` kept as defense in depth. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A --------- Co-authored-by: 江国庆 <guoqingjiang@deepglint.com> Co-authored-by: Claude (claude-fable-5) <noreply@anthropic.com>
## Summary Present Raven's orchestration capabilities and four specialized agents with benchmark figures, and keep the English and Chinese READMEs aligned. - Add the Multi-Agent Orchestration Benchmark at the end of the introduction. Compare Hermes, Claude Code, and Raven with shared 0-1 scales for Node F1, Edge F1, Partial Order Accuracy, and Exact Match Rate. - Move Raven Agents before Benchmarks and add concise Research, Code, Design, and Oncall descriptions with their figures. - Replace the third-party agent table with a centered card image and explain how Raven connects to and orchestrates these agents. Host image files as GitHub attachments. - Standardize primary heading markers and align section order, captions, model settings paths, and the EverMind ecosystem catalog across both languages. ## Type - [ ] Fix - [ ] Feature - [x] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Commands run after rebasing onto the latest main: ```text uv run --no-sync pytest tests/test_readme_scope_canon.py tests/test_living_docs.py -q -o addopts='' # 4 passed uv run --no-sync pre-commit run --files README.md README.zh-CN.md # All applicable checks passed make check-large-files # Passed make check-source-language # Passed env PYTHONPATH=. uv run --no-sync python scripts/check_commit_messages.py origin/main..HEAD # Passed git diff --check # Passed ``` Compared all 27 headings, 12 code blocks, image URLs and sizing, commands, paths, numeric values, and architecture graph structure across both languages. Verified all 24 orchestration scores against the supplied results and confirmed that the uploaded chart matches the local PNG byte for byte. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed ## Risk This changes documentation structure, wording, and linked figures. Runtime behavior is unchanged. Figures depend on GitHub attachment availability. Roll back by reverting the changes to README.md and README.zh-CN.md. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues #462, #465, #474, #477, #479 --------- Co-authored-by: Codex <noreply@openai.com> Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
…the roles (#468) ## Summary The agent loop already names four strategy roles (Memory, Planning, Capability, Action), but two judgements a turn makes still lived in the shell around them. This moves both onto the roles, in two dependent commits, with no behaviour change intended. **1. The charter reads through Memory and Action.** A dispatched sub-agent's playbook (`_meta["raven.playbook"]`) was read from four places inside the loop and the tool registry. `prompt` and `stopWhen` are now read by `DefaultMemory.assemble()`, which fills `task_brief` and `task_done_when` on the `TurnContext` the identity segment renders. `checks` and `code` are read by `DefaultAction.judge()`, which `ToolRegistry.execute` asks once per proposed call, the way it already asks the permission gate. `tools` stays in `withheld_names()` and the registry still decides what a refusal does, so a replaced role withholds nothing it could not already withhold. **2. Mid-turn window shrinking seats on Memory.** The five recoveries the loop ran inline go through one method, `MemoryModule.shrink(messages, pressure, state, model)`, with `WindowPressure`, `WindowState` and `ShrinkResult` as the carriers: - proactive compaction, once the last billed reading crosses the trigger - the standing image window - a provider's overflow error - a picture refused inside a tool result - pictures refused for size The retry mechanism stays in the shell. The six `continue` sites and the iteration rollback did not move, so a hook sees a retried call exactly as it saw the first. `WindowState` is a per-turn object the shell owns and hands to each call, because the role outlives the turn; `image_window` has no default, since 0 is a real value meaning "no picture stays". The pure pieces both sides need move to a new package, `raven/agent/window/`, which the import contracts treat as cargo that may not import the loop shell. `REASONING_EFFORT_LADDER` moves to `contracts/llm_provider.py` because the loop's empty-response descent and the window's head summary both read it and neither package may import the other. Contract papers: `CONTRACTS_VERSION` 26 to 28, `CONTRACTS_LINE_CEILING` 3,040 to 3,200, each bump with its own review paragraph in `tests/test_kernel_budget.py`. A follow-up PR carries the third step, which gives a sub-agent nine verbs in place of the six hook phases. It is kept separate because it touches every bundled plugin. ## Type - [ ] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [x] Refactor - [ ] Other ## Verification Gates, run at each of the two commits: ``` uv run ruff check --no-cache raven evolver agents plugins-dist tests scripts uv run ruff format --check --no-cache raven evolver agents plugins-dist tests scripts uv run lint-imports uv run pytest -q npx commitlint --from origin/main --to HEAD --config commitlint.config.cjs uv run python scripts/check_commit_messages.py origin/main..HEAD uv run python scripts/check_source_language.py origin/main..HEAD uv run python scripts/check_large_files.py origin/main..HEAD ``` Import contracts: 10 kept, 0 broken at both commits. The suite is green at both; the failures this machine shows are the same ones it shows on an unmodified `main` (missing browser and optional packages), and neither commit adds one. The loop's own e2e files, which the default selection leaves out, were run too (charter, playbook, truncation, direct chat, fallback chain, ACP stdio): no failure that `main` does not also show. Behaviour preservation. Each commit was run against a worktree of its own parent under the same scripted provider, outputs normalised for nonces and temp paths, then byte-compared. Four scenarios, eight comparisons, all identical: - a whole turn with no playbook: host turn, spawned sub-agent, tool calls and refusals - a playbook-on worker lane in process, including two charter refusals - six window-shrink scenarios: proactive compaction, the standing image window, overflow, a refused picture, oversized pictures, the effort ladder - the five-plugin hook seam, every phase's decision and metadata per iteration Mutation check on the second commit: breaking each of the five moves inside `DefaultMemory.shrink` turns exactly its own scenario red, so the comparison is not vacuous. Real model, Sonnet 4.5 through OpenRouter, at both commits: - playbook on, in-process lane: both workers receive the charter, write their files and finish - playbook on, ACP fork lane: the charter prompt crosses the process boundary, and a charter check blocks a `write_file` outside the fence - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed ## Risk No user-visible behaviour change is intended, and the comparisons above are byte-identical on every path they cover. For anyone replacing a harness role: `MemoryModule` gains `shrink`, so a replacement written against the old protocol needs it. `DefaultMemory.bind()` names the missing member at assembly rather than failing inside a turn. Rollback is a revert of the merge commit. Nothing persisted on disk changes shape. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A --------- Co-authored-by: yao pengfei <yaopengfei@shanda.com> Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com> Co-authored-by: KT <74288668+0xKT@users.noreply.github.com>
## Summary Synchronize the Chinese README with the updated English overview. The change aligns the Raven introduction and built-in agent section, including the Host Agent positioning, Raven Evolver description, built-in agent wording, and orchestration readiness statement. The existing benchmark charts, links, captions, and the remaining bilingual sections are preserved. ## Type - [ ] Fix - [ ] Feature - [x] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed Commands run: - `uv run --no-sync pytest tests/test_readme_scope_canon.py tests/test_cli_onboard_commands.py -k readme_quickstart -q` (1 passed) - `uv run --no-sync pytest tests/test_docker_runtime.py -k container_exposes_no_engine_selector -q` (1 passed) - `UV_NO_SYNC=1 make check-large-files check-source-language PYTHON_VERSION=.venv/bin/python COMMIT_RANGE=HEAD` (passed) - `git diff --check` (passed) ## Risk - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes This is a documentation-only change. Reverting the commit restores the previous README wording. ## Related Issues N/A Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
## Summary Five upgrade-path and reporting defects in the embedding pin that #414 landed. Three were named as bounded follow-ups in that PR's final review; two came out of driving the acceptance environment afterwards, and both of those were invisible to the unit suite because every fixture already held the new shape. The two found by acceptance: - **A config still holding the retired endpoint shape does not load at all.** The block used to carry the address and the key; `EmbeddingConfig` forbids extras, so `raven doctor` ends in a traceback rather than in a degraded feature, and so does every other command that reads the extension blocks. The migration takes those keys now: the address is adopted onto whichever configured provider answers there, and when nothing answers the whole block goes rather than half of it, with the dropped address named in the log. A model left with no provider is a pin every reader resolves to nothing while every screen reads it as configured. - **A pin could not name a provider this config holds.** The provider half was checked against the registry alone, and Raven carries no spec for every vendor LiteLLM can reach. A pin the wizard had stored and every reader resolved came back from the settings page as a provider that does not exist, leaving the endpoint uneditable on the page that exists to edit it. A section the operator wrote counts as proof the vendor exists; a name nothing holds is still refused as a typo. The three from review: - **Clearing the pin was a no-op.** The picker's inherit option sends both halves empty and the writer dropped empty values before deciding anything, so the call returned applied while the file kept the old pair. - **Doctor reported a move it had not made.** The endpoint only moves when a configured provider answers at the retired address, and the return value was discarded. The unmade move stays in the remaining list now, with a line saying why. - **The canonical SkillForge glossary entry still said `skillForge.everos`** and attributed the local extraction pipeline to an embedded copy of the memory backend. ## Type - [x] Fix ## Verification - `uv run --all-extras pytest -q` -- 23058 passed, 110 skipped, 4 failed; the four are the same pre-existing failures main carries (`test_agent_loop_token_budget`, `test_agents_code_launcher`, `test_agents_oncall_launcher`, `test_ppt_engine_image_search`), matched by test id against a run on the merge base. - `uv run ruff check .` and `uv run ruff format --check .` -- clean. - `make check-commits`, `make check-source-language`, `make check-large-files` -- clean. - Every fix was reproduced before and after against a real config or a real gateway, not only through its test: the unloadable config came from the acceptance home itself and `raven doctor` was run on it before and after; the provider refusal was reproduced over a real WebSocket to a running gateway; clearing and the doctor report were each run end to end. - Each new test was checked by removing its fix and confirming it fails. - Wider acceptance on the running product, with the memory plugin and without it: a gateway with no memory plugin installed built a knowledge base whose width (4096) was measured from the live endpoint and answered a query sharing almost no keywords with the source (score 0.728); the real TUI recalled a fact stored in an earlier turn, with the recall payload in the audit artifact as evidence rather than the model's answer. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed ## Risk Two changes widen what is accepted rather than narrowing it: a provider section counts alongside the registry, and a migration edits a block it previously left alone. The migration only runs for a config carrying the retired keys and is stamped, so it runs once; a config already in the current shape is untouched. Rolling back any single commit is safe -- they do not depend on each other. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ith (#483) ## Summary An EverOS agent case carries three fields: `task_intent`, `approach` and `key_insight`. Recall built `Memory.text` from the first and the last only, so the method a case was solved with never reached the prompt. The reader saw what was attempted and what was concluded, with the steps between them missing. The other two consumers of the same row already keep all three fields: the sub-agent memory record joins intent, approach and insight, and the memory page maps `approach` to a card's body. Only the recall path that feeds the model dropped it. This joins the three fields in field order. Two constraints shaped how: - Intent stays on the first line. The skill router takes this text's first line as a hit's display name and as its line in the search listing, so moving intent off line one would rename every recalled case. - Empty fields drop out, so an EverOS that answers without an `approach` still reads as prose instead of carrying a blank paragraph. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification ``` uv run --frozen pytest tests/test_everos_backend.py tests/test_skill_router_everos_source.py -q 182 passed in 17.69s make lint-python lint-types ruff check: All checks passed! ruff format: 1960 files already formatted ty check: All checks passed! make check-source-language clean, no findings ``` The case test asserts the exact joined text rather than substring membership, and it was run against the pre-change conversion first to confirm it fails there (AssertionError on the missing approach). A second test covers a row that arrives without an `approach`, which is the shape that would otherwise produce a blank paragraph. No user-facing doc states the recall text shape. The CONTEXT.md passage that lists all three case fields describes the sub-agent memory record, which is a different code path and is unchanged. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed ## Risk Each recalled case now carries one more paragraph, so the agent track's share of a prompt grows by roughly the length of one `approach` per case hit. Nothing downstream parses this text: the skill router reads the first line for a display name and passes the remainder through untouched, and that first line does not move. Rollback is a revert of the single commit. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
## Summary Synchronize the Chinese README with the updated English README. The change aligns the Raven introduction, benchmark captions, built-in agent copy, third-party agent positioning, and the removal of the obsolete benchmark summary section. The change also adds the updated wording for data analysis, PowerPoint slide generation, and unattended AI4S benchmark captions while preserving all chart links and remaining bilingual sections. ## Type - [ ] Fix - [ ] Feature - [x] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed Commands run: - `uv run --no-sync pytest tests/test_readme_scope_canon.py tests/test_cli_onboard_commands.py -k readme_quickstart -q` (1 passed) - `uv run --no-sync pytest tests/test_docker_runtime.py -k container_exposes_no_engine_selector -q` (1 passed) - `UV_NO_SYNC=1 make check-large-files check-source-language PYTHON_VERSION=.venv/bin/python COMMIT_RANGE=HEAD` (passed) - `git diff --check` (passed) ## Risk - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes This is a documentation-only change. Reverting the commit restores the previous README wording. ## Related Issues N/A --------- Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
…ch (#485) ## Summary Enabling an entrance from the page (weixin, for one) builds and starts the adapter on the spot through `ChannelManager.start_one`. The gateway command hands every channel its spine dispatch in one loop at launch, and that loop has already run by then, so a channel started this way has no `intake._submit`: it logs `no spine dispatch wired` once per message and drops every one, while the page draws it connected and the adapter's own login has succeeded. The outlet half of the same hazard was fixed earlier: `channels.on_started` registers an outlet so a hot-started channel can be replied to. The intake half stayed a launch-time loop, and `start_one` fires `on_started` and nothing else, so nothing after launch ever calls `set_submit` for a channel born later. The fix composes the intake onto the hook the outlet already uses. A module-level `_wire_channel_intake(channels, dispatch)` sets the submit on each channel the manager holds and wraps `channels.on_started` so a channel started later gets the outlet it already got plus the intake; the command calls it at the point where `_inbound_dispatch` is born. Channels present at launch are wired exactly as before. The helper is module-level rather than a closure in the command body so that it can be executed by tests, which the command body cannot be. Nothing under `ui-web/` or `ui-tui/` changes. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification - `.venv/bin/python -m pytest tests/test_cli_gateway_commands.py -q` on this branch rebased onto current `main`: 50 passed. Four new cases drive `_wire_channel_intake` through fake channels and a fake manager: every channel present at launch gets the dispatch; a channel the manager starts later gets it through `on_started` while the outlet hook that was already there still runs; a manager with no outlet hook still wires the intake; and the command body is pinned by source to call the helper with `_inbound_dispatch` and to have lost the launch-only loop. - Mutation-checked both ways: dropping `wire(ch)` from the composed hook fails exactly the two cases that start a channel late; dropping the outlet call fails exactly the outlet assertion of the hot-started case. - `python scripts/coverage_gate.py diff --base-ref upstream/main --threshold 90` after the file ran under `--cov=raven --cov-branch`: 91.67% (11/12 executable changed lines). The one uncovered line is the call inside the `gateway` command body, which no unit test executes, the same as every other line of that body. - `.venv/bin/ruff check` and `.venv/bin/ruff format --check` on `raven/cli/gateway_commands.py` and `tests/test_cli_gateway_commands.py`: clean. - The drop was reproduced on a live gateway before the fix (2026-09-14): enable, disable and re-enable weixin through the page's own RPC puts the adapter on the `start_one` path, and every inbound message was logged as `no spine dispatch wired` and dropped. - No user-facing docs change; the behaviour was always the documented one. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed ## Risk - Security: none. No new input path; the dispatch a hot-started channel now receives is the same one launch-time channels already had. - Backward compatibility: channels present at launch are wired by the same call as before; the only change is that a channel started later is wired too. - Rollback: revert the squash commit; the hook falls back to the outlet-only lambda. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
…es (#469) ## Summary Follows #468 (merged as `869ad78a`), which moved the charter reads and the window's mid-turn shrinking onto the harness roles. This branch is rebased onto `main` and carries only its own three commits. A sub-agent's product logic was an `AgentHook` subclass: six phase methods over a fourteen-field context it could write, with the turn's own state kept in the context's free-form `metadata` dict. Surveyed before this change, the five bundled plugins had 71 phase methods, 42 of them doing real work, reading fourteen context fields and using six decision fields. Every one of them fits a small set of verbs. This PR makes those verbs the contract, in three commits. **1. A sub-agent writes a conduct instead of six hook phases.** New factory-loop-tier paper `contracts/agent_conduct.py`: `AgentConduct` with nine optional verbs (`intake`, `select_tools`, `advise`, `system_addendum`, `review`, `salvage`, `outbound`, `archive`, `observe`), answered against a read-only `StepView` of fourteen fields, with `Intake` and `Verdict` (Accept / Resample / End) as the answers and a `ConductFactory` because the host builds one conduct per turn. It sits beside `loop_hooks` rather than replacing it. The six phases remain the loop's *timing* contract: when the loop asks, in what order, what a refusal does. This is the *judgement* contract: given one step, what does this agent say about it. `ConductHook` seats one in the other, so a plugin that still ships a plain `AgentHook` keeps working and the composite's ordering, merging and rollback are untouched. All five bundled plugins are ported: design-engine, ppt-engine, code-flow, oncall-flow and research-flow. Research keeps its six-phase gate chain: the gates now subclass a plugin-private `Gate` and read a `GateCtx` the conduct builds from its `StepView` and its own turn-private facts, so six thousand lines of gate bodies did not move and its wrappers still see the cross-gate state they read. The turn's conduct and the system-message addendum it splices live in the turn's own `metadata` dict, not on the hook: one hook instance serves every turn a process runs, and a system turn can overlap a user one on the same chain. **2. The conducts seat on the roles.** `MemoryModule.intake`, `PlanningModule.advise`, `ActionModule.review` and `ActionModule.salvage` each take this turn's conducts, so a replaced role decides what a plugin's judgement does rather than the seat deciding it. The composition rules live in `raven/agent/harness/conducts.py`: the first non-Accept verdict wins, advice joins in order, the first salvage wins, an intake threads the text through each conduct. `run_turn` binds the harness for the turn the way it binds the model, and the seat falls back to the same rules when nothing is bound, which is how a plugin test drives a hook directly. Contract papers: the conduct paper is factory-loop tier, so the contract-tier digest and `CONTRACTS_VERSION` do not move with it. `CONTRACTS_LINE_CEILING` 3,200 to 3,420, each bump with its own review paragraph in `tests/test_kernel_budget.py`. **3. What the review found.** Eleven threads, all resolved in the third commit. The ones that changed behaviour rather than prose: `CodeFlowHook` kept its `rolls_back_iterations = False` (inheriting `True` from the new base had it hold every reply behind the draft gate instead of streaming); `StepView` now publishes deep-frozen nested values, since freezing only the outer tuple left a conduct able to write through to `ctx.messages`; a turn's conducts are collected at the role boundary, so a replaced role sees one batch of the turn's judgements rather than one call per seat; an addendum is located by its own text rather than a saved offset, which two conducts appending system text used to shift out from under each other; `Planning` and `Action` gained binders, so a replacement role missing a new member is named at assembly rather than swallowed by the composite's `except Exception`. `CONTEXT.md` carries the Agent Conduct term and the Agent Hook entry no longer describes what this PR replaces. ## Type - [ ] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [x] Refactor - [ ] Other ## Verification Gates. CI runs these at the branch tip (17 pass, 1 skipped -- the skipped row is the disabled Claude Mention workflow). Locally, at each of the three commits: ``` uv run ruff check --no-cache raven evolver agents plugins-dist tests scripts uv run ruff format --check --no-cache raven evolver agents plugins-dist tests scripts uv run lint-imports uv run pytest -q npx commitlint --from origin/main --to HEAD --config commitlint.config.cjs uv run python scripts/check_commit_messages.py origin/main..HEAD uv run python scripts/check_source_language.py origin/main..HEAD uv run python scripts/check_large_files.py origin/main..HEAD ``` Import contracts: 10 kept, 0 broken at each of the three commits. The suite is green at each, and no commit adds a failure that `main` does not already show on this machine (this checkout has the heavy extras uninstalled, so the local run carries failures `main` carries too -- CI is the authority, and it is green). The loop's own e2e files, which the default selection leaves out, were run too (charter, playbook, truncation, direct chat, fallback chain, ACP stdio): no new failure. Behaviour preservation. The oracle for this change is the hook seam itself: all five plugins driven through every phase of a multi-iteration turn under the same scripted provider, recording each phase's decision and the observers it files. Outputs normalised for nonces and temp paths, then byte-compared. Re-run after the review fixes at the branch tip `81e6f423` (all three commits), against `main` at `c65c2059`: | scenario | bytes | | --- | --- | | five-plugin hook seam | 23,840 == 23,840 | | a whole turn with no playbook (host turn, spawned sub-agent, tool calls and refusals) | 12,821 == 12,821 | | six window-shrink scenarios | 133,371 == 133,371 | These cover the tip, not each commit separately. The seam comparison is the one that matters most for the third commit: its first fix changed `rolls_back_iterations` for code-flow, which decides whether replies stream or are held, and the seam oracle records the decision at every phase, so an unintended change there shows as a diff rather than as silence. A unit test proves the seat is a real decision point rather than a pass-through: with a lenient Action role bound, a conduct's Resample is not applied, and Planning's advice and Memory's intake are what the loop receives. Real model, Sonnet 4.5 through OpenRouter, at the branch tip: - playbook on, in-process lane: two workers, each receives the charter (3 checks), writes its file, finishes; 26 recorded model calls - playbook on, ACP fork lane: the charter prompt crosses the process boundary (the fork answers ZEBRA where the task text says PONG, so the brief reached the forked model through Memory), and a charter check blocks `write_file` to `notes/hi.txt` against a `pathPrefix: out/` rule - Raven-Research, a fork running the full gate chain: `plain-first: escalated_tool`, then `sufficiency-gate: released on the pages at iteration 4 after 2 searches / 2 pages` - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed ## Risk No user-visible behaviour change is intended, and the hook-seam comparison is byte-identical on every phase it covers. For plugin authors: the six-phase `AgentHook` is unchanged and a plugin shipping one keeps working. `ConductHook` is an adapter that seats a conduct in that same chain. For anyone replacing a harness role: the role protocols gain `intake`, `advise`, `review` and `salvage`, so a replacement written against the old protocol needs them. Rollback is a revert of the merge commit. Nothing persisted on disk changes shape. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A --------- Co-authored-by: yao pengfei <yaopengfei@shanda.com> Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
…res (#410) ## Summary Absorbs the first of two remaining commits from the research fork before it stopped iterating. This one lands instruments only: nothing here changes what a shipped run does, and the research flow label does not move because no model-visible text changed. `raven/agent/loop/dead_end.py` (new) asks whether a finished turn ended on the model's own prose. The test is structural, with a textual sub-label for the tail shape. Three decisions are worth review: - It lives under `raven/agent/loop/` rather than in the research plugin, because the loop is what would re-run a dead turn, and a host module cannot import a plugin without breaking `make lint-imports`. - The ask labeller is injected (`ask_kind`) rather than imported, because only the product that writes the harness asks knows how they are worded. With no labeller the failure mode is still recognised structurally and reported as `harness_ask_unknown`, so the boolean never depends on wording. - The fork marked its injected turns with a key of its own. This loop already had an equivalent in `_HOOK_INJECTED_KEY`, so the predicate reads that rather than adding a second marker for the same fact. The conditional rerun that reads the predicate arrived on this branch afterwards, in #418, along with the turn's wall clock. Both are supplied as data a hook writes onto the turn's metadata rather than as config the loop reads, so an agent whose hooks leave no budget runs exactly as it did before either existed: the rerun is unreachable for it and the clock is never consulted. The rerun starts from the original question rather than salvaging the failed attempt, because salvage was measured on 25 real duds to produce content 23 times and the right answer none of them, turning a detectable zero into a confident wrong answer. Three defects in that work were found by review and fixed on this branch, each with a reproduction that is red on the commit before it. The loop numbered episodes from an iteration count that restarts, so a rerun emitted indices [0, 1, 2, 0] and the TUI, which keys episode rows and their fold state by that index, had the rerun's first step collide with the first attempt's. Each attempt took its own clock reading, so a first attempt ending just under the limit handed the rerun a fresh full budget and a ten-second turn ran to eighteen. And the research hook opens its per-turn ledger at iteration 1, which a rerun reaches again, so the second open replaced the path the close reads and left the first attempt's file unreachable on disk. Reading that back afterwards turned up three more, none of them raised in review. The turn's end record carried the attempt's iteration count beside the turn's elapsed time with nothing saying which was which, so it gained ``attempt``. The hook contract described that record as "status and iteration count" while it carried five fields, which hid ``stopped_by`` from anyone writing a terminal gate against the contract. And the rerun's progress line named research in a loop that serves every agent, and called itself the first attempt when the budget allows more than one. A verbatim body sink (`RAVEN_VERBATIM_SINK`) records search and fetch results at three chokepoints. It is a second file rather than more columns on the ledger because a 250 KB row would leave no complete line in the 64 KB tail window `_resolve_session_seq` reads, which would silently reset the re-run split. Off unless the environment names a file, and a write failure is logged rather than raised. `DigestOutput` lets a digest return a write-only ledger annotation beside its text. The annotation is stripped from the tool result and merged into the fetch row, so a digest can never label its own output where the model can read the label. Nothing emits one yet; this is the pipe a later digest sidecar uses. `is_hard_tool_failure` was blind to the whole reader path. `web_fetch` never returns a bare string, it returns a JSON envelope, so a spent quota, a rejected URL and a failed validation all read as success and reset the streak `web_search` had accumulated. Transient markers are still tested first, so a 429 inside an envelope stays retryable. Forced finalize now books its own salvage spend. The phase fires on every run that produced no answer and burns up to `max_tokens` doing it, and it was counted nowhere. The Chinese refusal opener lives in `raven/i18n/zh_lexicon.py`, which is the exemption zone this repo keeps Chinese forms in, and the predicate imports it from there rather than carrying a literal outside the zone. Two defects found in the first review round are fixed on this head, both reproduced before being changed. The verbatim sink borrowed the ledger's `_resolve_session_seq`, which recovers the re-run split by reading the last 64 KB of the file and parsing the last complete line. That works for the ledger because its rows are small; this sink stores the large ones, so the first record wider than the window leaves nothing parseable and the sequence resets to 1, making two runs indistinguishable. Runs are now told apart by a tag minted in memory once per process and never recovered from disk, which removes the read a large row defeats. The argument for keeping this file separate from the ledger was already in its docstring; it applies to the sequence too, and now says so. `is_hard_tool_failure` started counting JSON envelopes while `failure_class` still read them as undifferentiated text. The streak key is `(tool, failure_class)` and the break threshold is two, so a blocked URL followed by a reader's HTTP refusal fired the stop-repeating nudge at a model that had changed both its cause and its approach. Both predicates now share one envelope parse, and an envelope is classified by its own `error` string rather than by a textual ladder that cannot see into one. Coarseness is for model prose, which varies without meaning anything; a tool writes its error from a small fixed vocabulary, so equal strings are the same failure and different ones are different. The parts that vary with the page travel in the envelope's other keys, so they cannot split a streak. The last commit removes an `isinstance` guard in the envelope reader that the diff-coverage report flagged as the only uncovered changed line. JSON has one shape that opens with a brace, so the branch was unreachable; removed rather than covered. ## Type - [ ] Fix - [x] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Run at this head (95c975b), on `origin/main` with nothing to rebase onto. ``` uv run --frozen --python 3.12 --all-extras pytest -q tests/ 22132 passed, 43 skipped, 109 warnings in 388.67s ``` Integration and e2e are deselected by default (`-m "not integration and not e2e"`), so the run above does not include them. This change adds no integration surface. ``` make coverage make coverage-ratchet line current=88.66% baseline=87.30% delta=+1.36pp branch current=81.59% baseline=79.43% delta=+2.15pp Coverage ratchet passed with 0.05pp tolerance. make coverage-baseline-check Coverage baseline update is monotonic. COVERAGE_BASE_REF=origin/main make coverage-diff Diff coverage: 100.00% (126/126 executable changed lines) ``` Diff coverage reported 99.21% on the commit before last, with one uncovered changed line: the sentence the rerun says to the reader, which no test drove because none passed an `on_progress` into a retrying turn. The last commit pins it rather than waiving it. ``` make lint-python All checks passed! / 1946 files already formatted make lint-imports Contracts: 10 kept, 0 broken. make lint-deps Success! No dependency issues found. make check-source-language clean make check-large-files clean make check-commits clean ``` The kernel budget gate caught one thing worth naming: documenting the turn-end record in the hook contract put `raven/contracts` at 2808 lines against a ceiling of 2791. The per-field detail moved to the write site in `turn_path`, which is not budgeted, and the paper now names the fields and says where the rest is. `raven/contracts` is at 2788. Each of the three review fixes was proved against the commit it fixes rather than reasoned about: a detached worktree at that commit, the new tests copied in, red there and green here. They reproduce the reviewer's own numbers -- the episode list `[0, 1, 2, 0]`, the eighteen seconds, and the stranded ledger file. The strengthened comparison in `test_the_retry_starts_from_the_question_not_from_the_wreckage` is a tightened guarantee rather than a regression test and passes either way, which is why it is called out here. `agents/raven-research/.env` was moved aside for every suite run: it is gitignored and absent in CI, and two launcher tests that delete a key from the environment otherwise find one in the file. This machine has no search-provider key, so no research brief can actually be run here. The retrieval-path changes are verified by unit tests only, which is why the three chokepoints are pinned by tests that assert both the presence and the absence of a row rather than by an observed run. The same limit applies to the rerun: the clock and the dead-end trigger are driven by a deterministic clock in tests, never by a real turn. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed No user-facing surface changed, so there is nothing to document. The one operator-visible knob, `RAVEN_VERBATIM_SINK`, is an instrument that is off unless set and is documented in the module that reads it. ## Risk The `is_hard_tool_failure` change is the only unconditional one, and it reaches every agent rather than just research: any tool returning a JSON envelope with a non-empty top-level `error` now counts toward the tool-failure streak that triggers the change-approach nudge. That is the intended correction, and the direction is conservative because transient markers are still tested first. Reverting it is a single branch. Everything else is inert by default. The body sink writes nothing unless the environment names a file. `DigestOutput` has no emitter in this change, and a digest returning a plain string produces the same envelope it always did, which `test_a_bare_string_digest_still_works` pins. The dead-end predicate now has a caller, and it reaches only an agent whose hooks write a budget onto the turn's metadata. No agent on this branch does: `agents/raven-research` ships `wallClockSeconds` and inherits the rerun from its plugin's class default, and every other agent leaves the key unwritten and is bounded exactly as before. The turn's overrun is one budget plus at most one iteration, because the clock is read between iterations and never mid-generation: cancelling a call in flight would discard a finished generation and leave no answer at all. Rolling the rerun back is the budget key; rolling the clock back is the same. Rollback is the whole commit; nothing here is depended on by anything already on main. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A --------- Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com> Co-authored-by: KT <74288668+0xKT@users.noreply.github.com>
…es (#472) ## Summary After a conversation switches model from the page, the agent keeps introducing itself as the model it left. The switch itself works -- the session record, the request and the usage ledger all carry the new id -- but the system prompt's `You are running on model: ...` line was read from `agents.defaults.model`, so the agent was told it still runs on the configured default and repeated that back. From the page it is indistinguishable from a switch that never happened. The fix reads the id from the turn's active `ModelBinding` instead. The loop already opens `use_binding(binding)` around every turn (`raven/agent/loop/main.py`, around `_run_turn`), and that binding is the session's own `/model` pick when it has one (`binding_for_session` in `raven/agent/loop/wiring.py`), else the default. `render._resolved_model_id` now takes its model id from `active_binding()` when one is set and from config outside a turn, and runs either through the same storage-to-wire conversion (`providers.wire`) it always ran the default through. One source file changes; nothing is threaded through the assembly path. This is the approach the review suggested over the earlier revision of this PR, which carried the id down as a new contract field, and it is better on every point that revision was faulted for: - **Right spelling.** The earlier revision forwarded `session.metadata["model"]`, the storage spelling, and rendered it as-is; on Codex, Azure and every gateway install that is not the id the request carries. Here the bound id goes through `wire_model` exactly as the default does, so `openai-codex/gpt-5.1-codex` renders as `gpt-5.1-codex` and a gateway install keeps its `openrouter/` prefix, in both branches. The conversion stays on the one path that owns it, and a lazily built provider's identity `wire_model_id` is not relied on. - **The estimation path agrees by construction.** `ContextBuilder._get_identity` (`raven/agent/context/builder.py`) sizes the system prompt for the token budget and calls the same renderer with no arguments; it is invoked from `_assemble_context_messages`, inside the binding, so it now names the same model the assembled prompt does. A field handed down the assembly path reached the segment and not this caller. - **No contract change.** `AssemblyContext` and `TurnContext` are untouched, so there is no field position to argue about, no producer (the curator's `TurnContext` rebuild) that can forget to copy it, and `CONTRACTS_VERSION` stays where `main` has it. - **`stable = True` stays true.** A binding is per session, constant within it; the line changes once per switch, which is what a switch is. A router's per-message pick is not reflected in the line (the router is opt-in and off by default); that is the one case a contract field would have covered, and it is also the case that made the segment's stability claim false. - **No empty-string hole.** `ModelBinding` refuses an empty model id at construction, so the line cannot be deleted by a blank pick. A fallback hop further down a model chain still reads a prompt naming the primary: context is assembled once, before the call. Not changed here and not a regression -- the sibling `can_see_images` field documents the same caveat. Nothing under `ui-web/` or `ui-tui/` changes. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Five new tests in `tests/test_segments.py::TestIdentityNamesTheBoundModel`: - `test_a_switched_conversation_is_told_the_model_it_is_bound_to` -- inside `use_binding` on a gateway config, `_resolved_model_id()` names the bound model with the gateway prefix and never the configured default. - `test_the_bound_id_goes_through_the_same_storage_to_wire_conversion` -- a codex-form id bound to the turn renders as the wire spelling the client sends, not the stored one; this is the assertion the earlier revision could not fail. - `test_outside_a_turn_the_configured_default_still_answers` -- no binding, config default, as before. - `test_the_segment_names_the_bound_model` -- the whole segment, through `IdentitySegmentBuilder.build`. - `test_the_estimation_prompt_and_the_turn_prompt_agree_on_a_switched_model` -- `ContextBuilder._get_identity()` equals the segment's text inside a binding and names the switched model. Commands and results: - Mutation-checked: replacing the binding read with the config read fails exactly the four tests that bind a model (4 failed, 46 passed in `tests/test_segments.py`); the outside-a-turn test and the three pre-existing wire-form tests stay green. - `.venv/bin/python -m pytest` over every test file that names `identity_text`, `_resolved_model_id`, `IdentitySegmentBuilder` or `ContextBuilder(`, plus `tests/test_contracts_two_tier_ledger.py`, `tests/test_kernel_budget.py` and `tests/test_context_stable_prefix.py` (11 files): 375 passed. - `python scripts/coverage_gate.py diff --base-ref upstream/main --threshold 90` after those suites ran under `--cov=raven --cov-branch`: 100% (3/3 executable changed lines). - `.venv/bin/ruff check`, `.venv/bin/ruff format --check` and `.venv/bin/ty check` on `raven/context_engine/segments/render.py`: clean. - The contracts gates pass unchanged: this revision touches nothing under `raven/contracts/`. - No user-facing docs change: the prompt line already documents itself as naming the model the turn runs on. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed ## Risk - Security: none. The binding's model id is already in every request and in the usage ledger; nothing new is read or sent. - Backward compatibility: outside a turn (no binding) the renderer behaves exactly as before. Inside a turn the line names the bound model instead of the configured default; for a conversation that never switched they are the same id, so the rendered prompt is byte-identical there. No wire, config or contract change. - User-visible change: after a `/model` switch the agent names the switched model in its identity line from the next turn on. The identity prefix changes once per switch, so a prompt-cache prefix keyed on it is invalidated once per switch, not per turn. - Rollback: revert the squash commit; one file, nothing else to reconcile. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues Fixes #470 Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
## Summary Update the hosted AI4AI (Nanochat 50M Pretraining) chart reference in both bilingual READMEs. - Replace the old image attachment URL in README.md and README.zh-CN.md. - Use the updated chart with the Bits Per Byte (BPB) label and the Lower Is Better annotation above the BPB values. - Keep the chart hosted as a GitHub attachment; no image binary is added to the repository. ## Type - [ ] Fix - [ ] Feature - [x] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed Commands: - `git diff --check` - `UV_NO_SYNC=1 make check-large-files check-source-language PYTHON_VERSION=.venv/bin/python COMMIT_RANGE=HEAD` - `UV_NO_SYNC=1 uv run --no-sync pytest tests/test_readme_scope_canon.py tests/test_cli_onboard_commands.py -k readme_quickstart -q` - Downloaded attachment SHA-256 matches the local PNG. ## Risk - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes This is a documentation-only URL update. Reverting the commit restores the previous chart. ## Related Issues #465 Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
…ber like the fleet (#493) ## What Two defects in the shipped `agents/raven-code` profile, both inherited from the retired vendored twin's comparison hygiene and both silent at runtime. **The skill route was advertised but dead.** Trunk's skills are pull-discovery: a name+description menu rides the user envelope, and its footer tells the model, verbatim, *"Read one with read_skill(id); search differently with find_skill"* (`raven/context_engine/scent.py`). raven-code kept `read_skill`, `use_skill` and `find_skill` in its disable rows, so the menu taught a route that answered "unknown tool" -- while the router still ran, retrieval still happened, and every body it fetched was dropped unread, with nothing logged to say so. All three tools register off the skill registry alone (`raven/agent/loop/wiring.py`); a Hub endpoint only adds their `hub/` branch, and `hub` itself stays disabled. **Memory was double-off.** design, oncall, ppt and research all ship `memory.backend: "everos"`; raven-code alone shipped backend `null` **plus** the `everos-memory` plugin opt-out -- the one shipped agent that could not remember. The identity was already `raven-code` in all three places (`memory.userId/agentId`, the plugin slice, `subagent.json`); only the two off-switches go. The redundant `skillForge` block goes too: its two knobs (`rewriteEnabled`/`llmGateEnabled`) are only built under push discovery (`factory.py`), which no shipped agent runs, and the other four agents carry no such block. **Review follow-up (second commit):** enabling everos-memory also offers the plugin's `understand_media` tool through the entry-point lane, which the hermetic face fixture cannot see -- an unledgered 14th tool in production. Same remedy as the playbook rows that board past the fixture the same way: the disable row is the pin, and `PLUGIN_LANE_WITHHELD` puts the lane on the ledger. raven-research holds the same line; design, oncall and ppt serve the tool deliberately. ## Fleet posture after this change | | everos | read/use/find_skill | skillForge block | |---|---|---|---| | design / oncall / ppt / research | on | design serves all three | none | | **raven-code (before)** | **double-off** | **all three disabled** | explicit false/false | | **raven-code (after)** | on | all three served | none | ## Tests The tool face moves on its ledger, not by luck (`tests/test_agents_code_launcher.py`): `SKILL_LANE_TOOLS` joins the face arithmetic, `find_skill` leaves `TRUNK_NEW_WITHHELD` (five stay withheld), `PLUGIN_LANE_WITHHELD` pins the plugin lane, and the everos posture test is rewritten from *stays-factory-off* to *ships-on*. The face itself is hermetically rebuilt from the render: 13 tools = the previous 10 + the three skill tools, verified by name and schema. - `tests/test_agents_code_launcher.py`: 70 passed - reviewer's set (`test_agents_code_launcher` + `test_agent_loop_skill_tools` + `test_context_scent` + `test_everos_plugin_discovery`): 110 passed - ruff check + format: clean --------- Co-authored-by: litong <238663200+TongLi31@users.noreply.github.com>
## Summary
Add `.github/CODEOWNERS` with a single entry mapping `/raven/agent/` to
`@LivXue`, so a
pull request that changes a file under that directory automatically gets
a review request
from the directory's owner.
This only requests a review. The branch ruleset keeps
`require_code_owner_review` off, so
no approval requirement changes and no contributor's merge path is
affected. Making it a
merge gate would be a separate ruleset change, deliberately not made
here.
Two limits are inherent to the mechanism rather than gaps to close
later:
- GitHub never requests a review from the author of a pull request, so
one opened by the
owner does not assign them. The API answers that case with HTTP 422.
- A draft pull request gets no automatic request. It is sent when the
draft is marked
ready for review.
CODEOWNERS is read from the base branch, so this governs pull requests
opened after it
lands and does not apply retroactively. Ten of the sixteen pull requests
open at the time
of writing touch `raven/agent/`. Those were assigned manually, apart
from the one the
owner authored, which cannot take the assignment for the reason given
above.
## Type
- [ ] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [x] Other
## Verification
Gates run in a clean worktree branched from `origin/main`, each exit
code read rather than
inferred from empty output:
```
make check-large-files exit 0
make check-source-language exit 0
make check-commits exit 0
```
Behaviour checked against GitHub rather than by reading the file. The
CODEOWNERS validator
on the pushed branch returns no errors, which is what proves both that
the pattern parses
and that the named owner resolves to an account holding write access:
```
gh api "repos/EverMind-AI/Raven/codeowners/errors?ref=..."
{"errors":[]}
```
Not run, and why: `make lint` and `make test-python` analyse Python,
TypeScript and
JavaScript, and this diff adds one ten-line text file containing no
code, so neither
measures anything about the change. No test in the repository exercises
CODEOWNERS
resolution, which is why the validator above is cited instead of a test.
Separately
confirmed that nothing reads the directory this file lands in as a set:
of the three tests
mentioning `.github/`, one validates backticked documentation paths
carrying a known file
extension, one names a single workflow file, and one globs
`.github/workflows` only.
- [ ] Relevant tests pass locally
- [ ] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed
## Risk
Nothing changes for any contributor: an added review request does not
gate a merge, and
CODEOWNERS confers no permission of its own. The owner handle named here
is already public
on every commit and pull request in this repository, so the file
discloses nothing new.
Rollback is deleting the file, or the one entry in it. Nothing depends
on it.
- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes
## Related Issues
N/A
Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
## Summary Replace the first image in README.md and README.zh-CN.md with the updated Raven banner, "One raven, a whole flock of specialists." Both versions use the same GitHub-hosted attachment. ## Type - [ ] Fix - [ ] Feature - [x] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed - `uv run --frozen --extra dev pytest tests/test_readme_scope_canon.py -x -q`: 2 passed. - `uv run --frozen --extra dev pre-commit run --files README.md README.zh-CN.md`: all applicable hooks passed. - `git diff --check origin/main...HEAD`: passed. - `make check-commits check-large-files check-source-language`: passed. GitHub confirmed the attachment upload. Download timeouts prevented a complete comparison with the supplied PNG. Browser rendering at desktop and mobile widths has not been checked. ## Risk - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes This changes the README banner only. The image is hosted as a GitHub attachment. Reverting the two URL edits restores the previous banner. ## Related Issues Fixes #126 Co-authored-by: Codex <noreply@openai.com>
#507) ## Summary Stands up a bilingual MkDocs site under `docs-site/`, published to GitHub Pages by a workflow that builds on every pull request and deploys only from `main`. Fourteen topics ship in English and Chinese: the quick start, self-hosting, Docker, the WebUI, the sandbox manual, the runtime architecture, the command reference, the repository layout, the developer workflow, the proactivity design and reference pair, the self-evolution map and the tracing API. This pull request is additive. Both READMEs keep their prose and gain one link to the site in the header row; trimming them is deliberately a second pull request, so the site can be confirmed live before anything depends on it. Key decisions: - **`docs-site/`, not `docs/`.** `docs/README.md` defines `docs/` as design notes and dated records, and `docs/plans/` is an archive explicitly not held to today's layout. A user manual has the opposite lifecycle. - **Full bilingual parity**, via the `mkdocs-static-i18n` suffix structure (`page.md` beside `page.zh.md`). Parity decays immediately without a guard, so a test fails when a page has no twin. - **Anchors are pinned on the Chinese side.** A CJK heading slugifies to nothing, so Markdown numbers it `_1`, `_2`, and inserting one heading renumbers every anchor below it, silently breaking links readers already hold. Each Chinese heading therefore states the anchor its English twin gets for free: 133 heading anchors, identical across the two languages, none positional. - **A left rail, not a top bar.** The header carries the brand, the search box and the navigation tree down the left; the table of contents on the right is one continuous polyline clipped to the headings currently on screen. - **No template overrides.** Material customises through `theme.custom_dir` with Jinja partials, those partials are `.html`, and `scripts/check_large_files.py` holds `.html` in `BLOCKED_ASSET_EXTENSIONS`. The skin is therefore one stylesheet remapping Material's own custom properties, plus one script for the table-of-contents rail. The brand mark is a `data:` URI for the same reason. - **An owner-signed `docs-site/` exemption zone** in AGENTS.md section 1.3. The i18n config carries the Chinese navigation labels, and `mkdocs.yml` is not `*.md`, so the blanket suffix exemption does not reach it. A contract test holds the AGENTS.md list and `EXEMPT_PREFIXES` equal, so both moved together. - **The site never links out to a file it does not host.** Documents that stay in the repository are named as plain paths rather than linked, and a guard fails a page that links to an external webpage or a repository-relative path. Two defects found and closed here rather than after publication. The Chinese site rendered an English sidebar: navigation labels live in `mkdocs.yml`, not in any page, so comparing the two languages' Markdown finds nothing and only the rendered output shows it. And a control could be present, sized, animated and answer `querySelector` while being unclickable, because a popup Material sized for a top bar landed outside the viewport once the header became a rail. ## Type - [ ] Fix - [x] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification All commands run at the pushed head, rebased onto `origin/main`. ``` make check-large-files exit 0 make check-source-language exit 0 make docs-build exit 0, 0 errors docs-site/site/index.html and docs-site/site/zh/index.html both present mkdocs_static_i18n: Translated 17 navigation elements to zh uv run pytest tests/test_docs_site.py tests/test_source_language_check.py \ tests/test_readme_scope_canon.py tests/test_docker_runtime.py \ tests/test_living_docs.py 44 passed uv run pytest tests/integration/test_docs_site_controls_e2e.py 13 passed pre-commit run --files <changed files> all hooks pass git diff --check clean ``` Anchor parity was measured on the built HTML rather than assumed: 14 page pairs, 133 heading anchors compared at h2 through h6, every one identical between the two languages and none positional. The controls are hit-tested rather than queried. `document.elementFromPoint` at the exact point a reader clicks must return the control itself, which is the only check that catches an off-viewport popup; the language switcher, the search panel, the repository link and the table-of-contents rail each pass it. The brand lockup is measured, not eyeballed. A failing test was written first and reported the wordmark sitting 4.00px below the mark's centre line; it is 0.00px after the fix, and a 4x screenshot puts the residual optical offset at 0.25px. Every guard this branch adds was confirmed to fail first. The brand-class guard was mutated from both sides in turn: renaming `.em-eyebrow` in the stylesheet turned it red, renaming `class="em-standfirst"` in `index.md` turned it red, and it is green again with both restored. The anchor guard was shown to catch both failure modes, a missing anchor and a wrong one. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed ## Risk No behaviour change to any shipped code path. Nothing imports `docs-site/`, the new Makefile targets are additive, and the only change outside `docs-site/` and its tests is one link added to each README's header row. Rollback is deleting the branch: the site does not exist until Pages is enabled with `Source: GitHub Actions`, and every other change is inert without it. Disclosed residue, checked and deliberately not changed: - The README header now links to `https://evermind-ai.github.io/Raven/`, which 404s until Pages is enabled. It is the same URL the Documentation section already carries. - The `deploy` job is skipped on pull requests by design, so a green check here does not prove deployment works. The first real evidence is the `deploy` job's own conclusion on the `main` run after merge. - GitHub Pages is not enabled on this repository. Measured: `GET /repos/.../pages` returns 404. Under `Deploy from a branch` the uploaded artifact is silently ignored while the workflow still reports success, so the source must be set to `GitHub Actions`. - `tests/integration/` is excluded by `norecursedirs`, so the browser suite is a local guard and never runs in CI. - The self-hosting page's Makefile target, docker path and port claims are unguarded. They were equally unguarded in the README they came from, and the owner ruled against adding a guard for them. - The self-evolution SOP and the benchmark contract stay in the repository and are named as plain paths from the pages that reference them. The SOP is a translation of an upstream project's internal document and carries redaction placeholders, which is a publication decision rather than a formatting one. - `docs/plans/` still shows the pre-restyle `mkdocs.yml`. CONTEXT.md defines that directory as an archive of the tree as it stood, with an explicit rule against updating an old plan to match a later change. - The Chinese pages drop the brand's italic serif accent word. Cormorant Garamond carries no CJK glyphs, so it would fall back to another face without any sign. - The code font is Material's default Roboto Mono. The brand names four typefaces and no monospace, so none was invented for it. ## Related Issues N/A --------- Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
… states (#503) ## Summary The overview says the four built-in agents are "Powered by the **Raven Evolver** engine". The architecture section of the same file says `evolver/` is a separate tool that consumes Raven as a library and that the runtime does not import Evolver, and `pyproject.toml` enforces that with an import-linter contract literally named "the runtime does not import the evolver". A reader gets two incompatible answers to whether Evolver is a runtime engine or an external development tool, and the contract settles which one is true. This was gloryfromca's blocking review finding on #486. The thread was resolved when that PR merged, but the text was never changed, so the contradiction is live on `main` today in both language mirrors. Both mirrors now put Evolver where it actually runs: refining the shared harness from outside, against benchmarks, developing the agents rather than running inside them. That is the reading the review offered for the case where the intended claim was that Evolver was used to develop the agents. The overview also takes the canonical spelling of harness self-evolution from the architecture section, as the review asked. Scope note: only the Evolver relationship changes. The SOTA sentence in the same paragraph is untouched - it was not part of the finding, and what it should say is a product call rather than a consistency fix. ## Type - [ ] Fix - [ ] Feature - [x] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Prose only: two lines, one per language mirror, no code and no new files. - `git diff --stat` -> `README.md | 2 +-`, `README.zh-CN.md | 2 +-` - The two gates that apply here are satisfied by construction rather than by a command: `check-source-language` exempts `*.md` by suffix (AGENTS.md 1.3), and `check-large-files` sees no added or resized file. - Read back against the two places the claim has to agree with: the Evolver row of the agent table, and the harness self-evolution bullet in the architecture section. Both mirrors now match both. - [ ] Relevant tests pass locally - [ ] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed ## Risk Documentation wording only. No runtime behaviour, no interface, no build. Rollback is a plain revert. The claim removed was the one the import-linter contract forbids being true, so nothing downstream depends on it. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A. Lands the finding from #486 that was resolved without a change. Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
) ## Summary `failure_class` keys the loop's tool-failure streak on an envelope's `error` string alone. That works only while the string comes from the small fixed vocabulary its producers are meant to keep, and it is a convention the module cannot enforce - the docstring added in #410 names exactly this failure mode: one cause counting as many. Every media tool broke it. `_format_http_error` interpolated up to 400 characters of vendor body (which carries a request id) into `error`; the unreadable-reference handlers interpolated the path; the four catch-alls put `str(e)` there, which spells the host. So a dead endpoint produced a class per call, `(tool, failure_class)` never repeated, the streak never reached `_LOOP_BREAK_THRESHOLD`, and the stop-repeating nudge was unreachable for `image_generate`, `text_to_speech` and `video_generate` alike. The URL gate in front of `web_fetch` had it too, in both copies of the tool: most reasons `validate_url_target` composes name the address they refused, so a model walking an internal range got one class per address. #410 fixed the handlers behind that gate and left the gate itself. `error` now carries only what a reader could enumerate. The variable part moves to `detail`, which keeps it in front of the model and inside `is_hard_tool_failure`'s transient-marker scan. Where the split gave the single call a further key, the batch reducer forwards that too: `hint`, which names the proxy a 403 can be routed through, and `url`, which names the reference a guarded fetch refused. A batched picture is told what the single call is told, which is the property the split has to preserve. Distinct causes stay distinct, which is the failure mode this trades against: two HTTP statuses, two exception types and the four vendor refusals each keep their own class, and the tests assert that direction too. Found by sweeping for the shape rather than by reading the code: every `"error":` whose value is not a literal is a candidate, and that grep is the cheapest guard against this recurring. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification - `uv run pytest tests/ -q -p no:randomly` -> 23443 passed, 109 skipped - `make lint-python` -> ruff check and format clean - `make lint-types` -> all checks passed - `make lint-imports` -> 10 contracts kept, 0 broken - `make check-source-language`, `make check-large-files` -> clean - Revert-to-red, per fix, restoring the source from a backup afterwards: reverting `media_gen.py` alone turns the four new class tests red and leaves `test_two_different_statuses_stay_two_classes` green; reverting either copy of `web.py` turns only that copy's gate test red and leaves the two #410 streak tests green. For the batch fold, each key is pinned on its own: reverting it to `detail` alone turns both new batch tests red, restoring `hint` alone leaves the reference test red, restoring `url` alone leaves the 403 test red. - After rebase onto 45203a4: `uv run pytest tests/test_media_gen_tool.py tests/test_security_web_ssrf.py -q -p no:randomly` -> 100 passed; `make lint-python`, `make lint-types`, `make check-source-language`, `make check-large-files` -> clean. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed ## Risk Tool output shape changes: a caller reading the reason out of `error` now finds it in `detail`. The readers in this repo are the model itself and the tests updated here; `tests/test_security_web_ssrf.py` is the one behavioural assertion that moved, and it still pins that a private address is refused for being private. Rollback is a plain revert of this commit - no state, no migration, no config. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A. Finishes the envelope work started in #410, which fixed the reader's handlers and left the gate in front of them and every media tool. --------- Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
…its (#490) ## Summary `raven gateway` exited with SIGSEGV (-11, or 139 in a shell) on every SIGTERM, so a systemd or docker stop recorded a crash instead of the clean stop it was. `raven tui` and `raven agent` had the same defect. Root cause: `LocalSkillCatalog` auto-starts `SkillFileWatcher`, a daemon thread that parks inside the Rust `watch()` of watchfiles. Daemon status does not make process exit safe while the thread sits in native code -- CPython runs `Py_FinalizeEx` under that call. Every Python shutdown step succeeds first, so the crash lands after the graceful chain is complete and the real exit code is masked. With faulthandler armed all three faults report `<no Python frame>`, i.e. no Python code is on the stack. Every owner now retires it: - the gateway's shutdown chain, for the generation still bound at exit; - the TUI's teardown; - the one-shot agent's teardown; - `RavenRuntime.dispose`, for a generation retired at a swap. `build_runtime` mints a fresh `AgentLoop` per generation, so each swap starts a new watcher and abandons the previous one; without this a reloaded gateway still segfaults on stop. The generation-organ roster grows the same entry, so the two halves of that contract stay in step; - the gateway's shutdown again, for a candidate staged by a reload that the serving loop never consumed. That generation is not the bound `agent`, so nothing else reaches it; - the gateway's shutdown once more, for a candidate `take()` handed out whose binding a cancelled unbind never finished. `_serve_generations` holds that candidate in a local across `await _unbind_generation()`, so a shutdown cancelling the await unwinds the coroutine and drops the local, leaving the generation staged nowhere and bound nowhere. `SwapCoordinator` now keeps it: `take()` records the candidate it hands out and `release()` clears it, which spans exactly handover to the moment the caller's own binding points at the new loop. The serving loop needed no new line, because it already calls both ends of that span. The replay runner already stopped the same watcher for the same reason; these were the other owners. The stop sits in the generation contract and in each command's teardown rather than in `AgentLoop.stop`, whose other caller is the generation swap -- that path keeps the process running and must keep skill auto-refresh with it. The gateway's sweep now lives in a module-level `_retire_generation_watchers` rather than inline in the serve closure. `run()` is a 550-line closure with no import seam, so the inline form could only be pinned by reading its source text -- which never executes it, and so earned no coverage on the lines that matter. The helper drains all three seats and keeps going when one stop raises, since a sweep that stops early still leaves a watcher parked in native code. Two comments describing a removed guard go with it: one in `agent_commands` pointing at an exit chokepoint in `raven.cli.commands.run`, and the test-side scaffolding that patched `os._exit` for an exit the command no longer makes. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Process-level, faulthandler armed, throwaway home and a dummy key nothing calls: ``` before after raven gateway SIGTERM -11 0 (-11 on 5 of 5 runs) raven gateway SIGHUP then SIGTERM -11 0 (2 reloads after) raven tui PTY, then signal -11 0 raven agent ordinary one-shot exit 139 1 (1 is the turn's own auth failure, which 139 had been masking) ``` Controlled pair isolating the cause, on the loop factory the TUI uses -- build the loop, let the interpreter finalize: ``` watcher left running exit 139 stop_file_watcher() exit 0 ``` Suites, from the branch tree: ``` uv run pytest tests/integration/test_gateway_sigterm_smoke.py -q -n 0 -m integration 1 passed (this is the guard test that was failing) pytest tests/test_cli_agent_loop_parity.py tests/test_cli_gateway_commands.py \ tests/test_cli_gateway_health.py tests/test_cli_tui_commands.py \ tests/test_cli_agent_commands.py tests/test_gateway_spine.py \ tests/test_generation_swap.py tests/test_plugin_services_lifecycle.py \ tests/test_rpc_control.py tests/test_rpc_system.py \ tests/test_skill_watcher_cleanup.py tests/test_trajectory_replay.py \ tests/test_utils_asyncio_runner.py -q 307 passed, exit 0 ruff check / ruff format --check over the changed files clean ty check over the changed files no diagnostics on the changed lines scripts/check_large_files.py origin/main..HEAD exit 0 scripts/check_source_language.py origin/main..HEAD exit 0 commitlint --from origin/main exit 0 ``` Second round, after review found the in-transition window; rebased onto `origin/main` at 6d8c363: ``` pytest tests/test_cli_gateway_commands.py tests/test_core_runtime_swap.py \ tests/test_generation_swap.py tests/test_cli_tui_commands.py \ tests/test_cli_agent_commands.py -q 171 passed, exit 0 every test file importing raven.core.runtime 131 passed, exit 0 real gateway: boot, SIGTERM, exit code 0 on 3 of 3, no fault text ruff check / ruff format --check clean make check-source-language / make check-large-files exit 0 coverage_gate.py diff --threshold 90 95.83% (23/24), passed ``` Revert check on that round: removing the `_in_transition` record, the in-transition seat in the sweep, or the sweep's `try/except` each fails its own test and nothing else. The cancellation window itself is reproduced against real `SkillFileWatcher` threads, counted by thread name, and is red without the fix. Diff coverage on changed lines was 50% (4 of 8) before that round, because the gateway-side tests read the chain's source instead of running it. The one line still uncovered is the call site inside `run()`, which cannot execute without booting a real gateway; the wiring to the seam is pinned by source inspection. Each of the four regression tests was confirmed red before its fix. The TUI and agent stops were additionally re-confirmed afterwards by removing both again and watching those two tests go red, then restoring them. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed ## Risk User-visible change: these three commands now exit with their real status instead of a segfault. Anything reading the exit code sees a different value -- 0 where it used to see 139, and for a failed one-shot the turn's own non-zero code rather than 139. Rollback: revert the commits. The watcher is a cache-invalidation helper, so removing the stops restores the previous behaviour exactly, crash included. Checked and deliberately not changed: - The stops are not wrapped in try/except. `ContextBuilder.__init__` assigns `skills` unconditionally, `stop_file_watcher` is documented safe when no watcher was started, and the watcher's own contract is that neither `start` nor `stop` raises. This matches the existing production caller in the replay runner, which is also unwrapped. - `raven acp` reaches finalization the same way but calls `os._exit(0)` after a real session, so the fault is masked there rather than fixed. Left as is; making it consistent is a separate change. - No central gate was reinstated. A future command that builds a loop and exits normally would reintroduce this, and the source-inspection tests here only guard the retirement seats that exist today. - The watcher starts in `ContextBuilder.__init__`, which is construction-time work, and `RavenRuntime.discard` is contract-bound to stay call-free because a call there would mean an organ was started before FREEZE. So a staged or in-transition candidate is retired by the gateway's sweep rather than by `discard`, and the docstring records why. Moving the watcher out of construction would let `discard` own it; that is a separate change. - The shutdown-before-consumption window was not reproduced by timing: three runs sending SIGTERM 0.2s, 1.0s and 2.0s after a reload all completed the swap first and exited 0. The staged-candidate retirement is driven by forcing the state in a test rather than by racing for it. Known gap in the evidence: the two `raven tui` runs were ended with SIGTERM after Ctrl+C went unanswered, because the TUI never finished loading against a throwaway home. The ordinary-exit case is evidenced by the agent one-shot, which exits by itself with no signal involved. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A --------- Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
…ked again (#467) ## Summary Two regressions in the hand-built deck lanes (medium and high on Raven-Design), and the stream failure that made every run of them a coin toss. **The deck lane draws its own pages again.** PR #451 pointed the hand-built lane at the engine's helper modules (`ppt_layout`, `ppt_theme`, `ppt_shapes`), which were written for the max lane's script backend. Their defaults became the deck: `heading()` put a white surface band bled across every page, `card()` became the way anything was placed, text frames were top-anchored with the lower third empty, and a build script named three colours. Measured on the same request, the accent colour appeared 8 times across 20 pages. The skill, route note and assets reference go back to the pre-#451 text: the lane draws with python-pptx and takes only the icon set, `add_formula` and `math_runs` from the engine. On the reverted lane the accent appears 54 times, the deck carries 12 real pictures and the layouts differ page to page. **Colour is named beside the type sizes.** The skill listed sizes for every role and said nothing about colour, so a lane with the right palette still drew in one colour. One paragraph now says what colour is for; the layouts note no longer reads as one accent per page, which a lane took as a ceiling. Where `image_search` is not offered, the fallback no longer tells the lane to give the pictures up: `web_search` for the page, `web_fetch` it, and download where the fetch backend keeps image links. **A stall in a silent think is asked again.** Across two measured runs the lane's turn failed 28 times on `Upstream idle timeout exceeded`, every time in the one round that writes the whole build script into a tool argument. `stream_llm_call` refuses to retry after output a watcher has seen, and counted reasoning deltas and half-built tool-call fragments as that output, though neither is ever shown as the reply. So the configured retry ladder was never reached. `rendered()` now asks what a watcher could have seen; every retry path resets the buffers first (a slot left standing merged with the next attempt into a call the model never made); a stall handed back as an error response carries the usage it spent. The in-process builtin backend, whose comment claimed a retry ladder and passed none, now receives the deployment's `llmErrorRetryDelays` and `llmRetryAfterOutput` through the SubagentManager. No packaged product takes that path; all five run as ACP agents on the loop's own call site. **Two bundled templates edited in place**, by the maintainer's word. `beige_geometric_general_report` loses its page 20, a dark photograph cut by a circle under pale rounded cards; 23 pages become 22 and the two reference pages the lane borrows from it (17, 19) keep their numbers. `black_circuit_tech_launch` loses the 27 percent circuit texture on its master and the photograph on its Section Header layout, so its body and section pages sit on the master's plain `#000000`; cover and closing are unchanged. `templates.manifest.json` is repinned; `fetch_templates.py --verify` holds all ten pins. The registry archive was not re-cut. No configuration default changes. The only behaviour every lane sees: a stream that fails before any reply text reached a watcher is retried on the configured ladder instead of failing the turn. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification ``` uv run --frozen pytest tests/test_agent_loop_llm_error_patience.py tests/test_agent_loop_stream.py \ tests/test_subagent_manager.py tests/test_subagent_routing_backend.py -q 265 passed in 6.85s (after rebase onto main) uv run --frozen pytest tests/test_litellm_provider_stream.py tests/test_litellm_provider_stream_end.py \ tests/test_provider_stream_fallback.py tests/test_providers_stream_idle_budget.py -q 102 passed uv run --frozen ruff check raven/agent/loop/main.py raven/agent/subagent/backends/raven_loop.py \ raven/agent/subagent/manager.py raven/providers/streaming.py All checks passed! grep -nP "[^\x00-\x7F]" agents/raven-design/route-raven-ppt.md -> no matches (the note is pinned ASCII) ``` Both new tests were checked against the unfixed code: `test_a_stall_during_a_silent_think_is_asked_again` fails with the old `emitted` test, and `test_what_the_failed_attempt_left_behind_is_dropped` fails with the buffer resets removed. The failure itself was reproduced outside the runtime: the same 37-message history, six pictures and tool table, sent through litellm 1.101.0 to the same vendor, succeeds on a short answer every time and fails mid-stream one time in three when asked to write the full build script; healthy long streams show a maximum inter-chunk gap of 1.9s over 234s, so the cut is a dropped connection, not a slow model. Decks built on the changed lane, same request as the #451 acceptance run (glm-5.3-flash via OpenRouter): medium 21 pages with 12 pictures and 7 text colours, 48 gate findings (4 `excessive_whitespace`, down from 18); high 13 pages with 13 pictures and 9 text colours, 37 findings. The high run's page count came from the host agent's forwarded brief, which stated 10-14 pages the user never asked for; that is a separate issue. Template edits: `tests/test_ppt_engine_templates.py tests/test_ppt_engine_default_templates.py tests/test_ppt_engine_template_bands.py tests/test_ppt_engine_template_theme.py tests/test_ppt_engine_measure_geometry.py` -> 152 passed, 1 skipped; `make check-large-files` -> clean; both templates rendered and read page by page. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed ## Risk - [ ] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes User-visible: on every lane, a streamed call that fails during reasoning or mid tool-call, with no reply text yet shown, waits out the configured ladder (default 15/30/60s) and asks again instead of ending the turn. A call that already streamed reply text behaves as before. Reasoning tokens of the abandoned attempt are spent again and are now recorded in usage. Rollback: revert the `providers` commit alone to restore the old failure; the skill text and the builtin-backend plumbing are independent of it. ## Related Issues N/A
## Summary Space on an attempt row in the `raven trajectory` browser previously showed two truncated per-turn preview lines. This PR replaces it with an interactive full-conversation viewer. It re-lands the content of PR #382, which was auto-closed by the 2026-09-12 history linearization without being merged; the branch is rebased onto the linearized main, with one extra commit expressing the CJK display-width test fixtures as unicode escapes to satisfy the source-language gate that main gained since. - Data layer `raven/trajectory/conversation.py`: rebuilds an attempt's span snapshot into a causally ordered stream of labeled records (User input, LLM input/output, Tool input/output, skill, memory, and subagent events). Event-time semantics (inputs at span start, outputs at span end) plus nesting depth from parentSpanId keep the causal order. Full content comes from artifact files (paths restricted to the trace store, streamed 512 KiB cap per file); span previews are only a visible fallback. ERROR spans and expected-but-unreadable payloads always yield a placeholder record. LLM inputs omit only the item-by-item verified common message prefix against the previous call (omissions and rewrites are announced); tool_calls of unexecuted tools keep their full arguments. - Rendering: kind-colored labels aligned to one column, hanging-indent bodies with word-aware CJK-safe wrapping, degraded/error/meta note lines, single-pass turn separators (interleaved traces are never regrouped), and a stacked layout on narrow terminals that never drops characters. - Interactive viewer (prompt_toolkit full-screen, entered for every preview on a terminal): scrolling, a muted bottom bar with an 'h for help' hint, filter/collapsed state and the scroll range (laid out by display width, never overflowing any terminal), a scrollable help page reachable in full on small screens, a global collapse (s: bodies over five display lines fold to five plus an ellipsis; error and degradation notes never fold), and a label filter (%: real input line, case-insensitive, empty clears; filtered-out turn groups lose their separators). Resizes re-wrap and re-judge folding live; the content caret is hidden (only the filter input shows one). q/Esc returns to the list; Ctrl+C cancels the whole browser from any viewer state. A rebuild failure falls back to the legacy per-turn preview; without a terminal the preview degrades to direct printing. - Every dynamic field passes the sanitization gates (control characters, rich markup, and ANSI never reach the terminal). Adds the Conversation Record term to CONTEXT.md. ## Type - [ ] Fix - [x] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification - `uv run pytest tests/test_trajectory_conversation.py tests/test_cli_preview_viewer.py tests/test_cli_trajectory_browse.py tests/test_trajectory_store.py -q` -> 298 passed (viewer assertions run against really rendered screens) - `uv run pytest -q` (full, on the rebased branch) -> 23463 passed, 125 skipped; 33 failed + 20 errors reproduce with identical node ids on a clean origin/main worktree - environment baseline, not introduced here (passed-count delta of 96 equals the tests this branch adds) - `make check-source-language`, `make check-large-files`, `make lint-types`, `make lint-python`, `make lint-imports` -> all pass - Manual: rendered 6 real local attempts at widths 72/120/30 with zero over-width lines (largest attempt 7060 lines in 0.35s); drove the real TUI in a pty at normal and 44x12 sizes (viewer, help paging, collapse, label filtering, Ctrl+C cancel) - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed ## Risk - Read/display side only; no persisted format changes. The untrusted input surface (artifact content, typed filter input) is contained by the store-prefix path check, the 512 KiB streamed cap, and per-field sanitization. - User-visible change: the Space preview is now a full-screen viewer; the legacy rendering remains as the automatic fallback when the rebuild fails, and non-terminal environments print directly. - Rollback: revert this PR, or restore the previous `_preview_screen` call site to get the old preview back. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A --------- Co-authored-by: Claude (claude-fable-5) <noreply@anthropic.com>
…e shell (#516) ## Summary The documentation site built one search index for both languages, so a reader on the English pages was offered Chinese ones: results they cannot read, behind links that leave the language they chose. A build hook now splits the merged index in two and points the Chinese pages at their own copy, which the theme also resolves every result link against. Where that split runs is the whole fix. The merge lands in a plugin's own `on_post_build`, so the split has to be the last thing that event does. Running it on `on_shutdown` looks equivalent and is not: `mkdocs serve` never shuts down, so the preview went on serving the merged index however often it rebuilt. The new guard makes the same assertion against `mkdocs build` and against `mkdocs serve`, because passing one says nothing about the other. The rest settles the shell against the reference layout: - the rail's foot is its own box the navigation cannot show through, and the tab icon is the EverMind mark; - the footer keeps to the content column instead of covering the fixed columns once the reader reaches the page foot; - the contents column reaches the screen foot with its scrollbar hidden behind a fade, and follows the reader by centring the first lit entry; - search opens as a centred dialog over the page. The form itself travels into that dialog, so what stays in the rail is drawn from the field's own metrics and its shape does not change on a click. Neither state animates: the theme transitions the two shapes into each other, and leaving search played that out in the corner of the page after the reader had moved on. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification - `uv run pytest tests/integration/test_docs_site_search_e2e.py` -> 2 passed. Without this change the preview case fails: the Chinese index is HTTP 404 and the English index carries 146 Chinese pages. - `uv run pytest tests/integration/test_docs_site_controls_e2e.py` -> 25 passed - `uv run pytest tests/test_docs_site.py tests/test_source_language_check.py` -> 31 passed - `make check-commits`, `make check-source-language`, `make check-large-files` -> exit 0; `ruff check` on the added files -> no findings - `make docs-build` -> the root index holds 149 entries and no Chinese page; the Chinese index holds 146, none of them addressed from the site root Driven in a real browser at 1560x820 in both languages: a word that appears in both trees returns 10 results and 0 from the other language, and the corner of the page is pixel-identical to its never-opened state at every frame after leaving search. Every page of the built site was checked at the page foot (28 pages, 0 faults) for a footer covering a column, for a column whose last entry cannot be reached, and for the rail's foot icons being clickable. No page content changed, so nothing else needed updating with it. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed ## Risk User-visible: search no longer returns pages from the other language, and the search control, the footer and the two fixed columns change shape. The field left in the rail while the dialog is open is a drawn replica, and every rule behind it sits above the theme's desktop breakpoint, so narrow viewports keep the theme's own full-screen search untouched. Rollback: revert the commit. The hook is one file, and dropping it from the `hooks:` list in `docs-site/mkdocs.yml` restores the previous merged index on its own. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A --------- Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
) ## Summary An EverOS role points a section at a provider: the settings card writes the model and asks `lend_provider_credentials` for that provider's key and address to go beside it. The address was copied only when there was one to copy -- the reasoning being that a section with no lender address may be holding one the reader typed by hand. A section being pointed at a new lender is holding the OLD lender's address, though, and leaving the field out of the answer keeps it. Moving the memory role from an OpenRouter model to a DeepSeek one left DeepSeek's key at OpenRouter's address, which the far end refuses. DeepSeek is not a corner case here: every vendor LiteLLM routes carries no `default_api_base`, so the lend answers with no address for all of them. The lender answers about the address either way now, empty included. An empty value reads where it lands the way an unset one does: the EverOS writer sets the field, and the section stops carrying somebody else's endpoint. Found while running the settings page's acceptance case for the EverOS roles against a real gateway. The lend rule is from #414, not from the settings work that surfaced it. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification - `uv run pytest tests/test_config_update_providers.py tests/test_rpc_settings.py tests/test_cli_onboard_commands.py tests/test_rpc_console.py -q` -- 641 passed on the branch this was written on; 191 on the two closest suites here. - The new case fails without the change: the lender answers `{"api_key": ...}` where the section needs `{"api_key": ..., "base_url": ""}`. - `uv run ruff check` and `ruff format --check` over both files -- clean. - Real gateway, temp `RAVEN_HOME`, EverOS plugin present: `settings.everosSet` with `borrow_from: "deepseek"` writes the model and DeepSeek's key and leaves no `base_url` behind; the same call with `borrow_from: "openrouter"` writes OpenRouter's address back. Before the change the first of those left `base_url = "https://openrouter.ai/api/v1"` under DeepSeek's key. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed ## Risk - Two callers, both of which point a section at a provider: the EverOS card (`settings.everosSet`) and the CLI onboarding reuse step. For a lender that has an address, nothing changes. For one that has none, the borrowing section now stops carrying the previous lender's -- which is the fix. - A reader who typed an address by hand into an EverOS section and then borrowed from an addressless provider loses that typed value. That is the same event as the bug, read the other way; the section is being pointed somewhere else, and the card shows the resolved address. - Rollback: revert the commit. ## Related Issues N/A --------- Co-authored-by: gloryfromca <23442919+gloryfromca@users.noreply.github.com> Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
…ction (#557) ## Summary Open the Showcase with a Raven-Code run that picked, built and judged an FPS game over 42 rounds, and give README.zh-CN.md the Showcase section it never had. The entry sits in its own table above the existing ones: the left cell carries the run's task graph and a 30s clip of the result, and the right cell is left empty for a second sample. It is a separate table rather than a row in the deck grid, because a playable game is not a deck and an empty cell inside that grid would read as a broken row. The clip is a GitHub attachment referenced by a bare URL on its own line. That is the only form that renders a player: a video element is removed by the markdown sanitizer, which was confirmed by rendering both forms through the markdown API before writing either one. The same rendering also shows that the player header displays the name the file carried at upload, so the clip was re-uploaded under its intended name rather than left under a working name. README.zh-CN.md previously had eight sections against the English nine, and the missing one was Showcase. Rather than hand-copying it, the section was derived from the English one and its entry titles translated, so the two files stay structurally aligned; alt text is left in English, matching every other alt attribute in that file. Both files now render the same structure: one player, seventeen images, fourteen cells, one empty cell. The section's opening sentence is reworded in both files while the section is being touched: it now says Raven generates the orchestration rather than chooses it, and the Chinese copy names the panel as the multi-agent orchestration graph. Checked and deliberately not changed: - The task-graph alt text describes three nodes running "across 42 game rounds". The panel actually shows round 42 of a capped loop, with those three nodes being that one round's pick, do and judge steps, so the wording is loose. The author chose to keep it; it is recorded here rather than silently altered. - The first upload of the clip is still attached to its hosting issue under the earlier filename. Nothing references it, and attachments are not removable once posted, so it stays as upload history. ## Type - [ ] Fix - [ ] Feature - [x] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Rendering, through GitHub's own markdown API rather than by eye, on the exact section text of each file: gh api -X POST /markdown --input - # mode gfm, repo context Both files returned one video element, seventeen images, fourteen cells and one empty cell, with the player header reading the intended filename and the first entry title reading as intended in each language. Byte equality of every uploaded asset against its local source, fetched back over HTTP: curl -sL <asset-url> | sha256sum All matched. A HEAD request is not usable here: the attachment host answers HEAD with 403, which reads as a failed upload. Tests that actually read these two files, found by grepping for them rather than assumed absent: uv run --frozen --all-extras pytest tests/test_cli_onboard_commands.py tests/test_docker_runtime.py -q 344 passed Repository gates: make check-commits check-large-files check-source-language All passed. The two file gates were also run against a bare revision, which compares the working tree, because a commit-range form reports nothing while an edit is still uncommitted. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed No test asserts anything about Showcase content, asset URLs or section parity between the two READMEs. The 344 above cover one heading lookup and one whole-file read; nothing would fail if an entry were malformed. The rendering check above is what stands in for that, and it is manual. ## Risk - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes Documentation only; nothing executable changes. The assets are public attachments on a repository that is already public and carry no credential or private detail. The clip is 53 MB, so the Showcase section is heavier to load than before, and the player appears only in the rendered file view: the diff and compare views show the bare URL as text. Rollback is a revert of these commits; the assets stay hosted either way. ## Related Issues N/A --------- Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
## Summary - Align English and Chinese documentation for Docker/Compose, sandbox execution, and the in-tree tracing API with verified runtime behavior. - Separate proactivity usage guidance from developer design and implementation details, with configuration, troubleshooting, cost, and safety boundaries. - Update navigation and remove redundant section separators. Keep page URLs unchanged. - Preserve published proactivity section and title anchors in both languages. Use native aliases for same-page content and visible migration links for moved sections, with regression tests. - Clarify Sentinel delivery fallback and menu-selection requirements, tracing-on/off argument handling, headless CLI menu output, and viewer descriptor lookup. Add focused regression coverage without changing runtime behavior. - Clarify the initial frozen baseline and improve Evolver wording in the bilingual site and repository mapping. Keep its structure, navigation, implementation terms, and technical claims unchanged; broader Evolver documentation work is excluded. - No runtime changes. ## Type - [ ] Fix - [ ] Feature - [x] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Passed on the rebased revision: ```bash uv run --frozen --no-sync pytest tests/test_docs_site.py tests/test_tracing_api.py tests/test_sandbox_unit.py tests/test_sentinel_runner.py tests/test_sentinel_planner.py tests/test_sentinel_fast_path.py tests/test_nudge_policy.py tests/test_proactive_spawn.py tests/test_core_sentinel_stack.py tests/test_core_cron_stack_ledger.py tests/test_cli_sentinel_commands.py tests/test_decision_router.py -q -n 0 -x # 395 passed uv run --frozen --no-sync --with 'mkdocs-material>=9.5' --with 'mkdocs-static-i18n>=1.0' --with 'mkdocs-git-revision-date-localized-plugin>=1.6.0' --with 'jieba>=0.42.1' mkdocs build -f docs-site/mkdocs.yml --strict --site-dir /tmp/raven-pr551-push.9ZfOEi/site # Strict English and Chinese builds passed UV_NO_SYNC=1 make lint-python PYTHON_LINT_TARGETS='tests/test_docs_site.py tests/test_tracing_api.py tests/test_decision_router.py tests/test_cli_sentinel_commands.py' UV_NO_SYNC=1 make check-large-files check-source-language COMMIT_RANGE=origin/main..HEAD PYTHONPATH=. uv run --frozen --no-sync python scripts/check_commit_messages.py origin/main..HEAD npm exec --yes --package=@commitlint/cli@21.1.0 -- commitlint --from origin/main --to HEAD --config commitlint.config.cjs git diff --check origin/main...HEAD # All passed ``` - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed Documentation checks also covered bilingual examples and the proactivity JSON configuration schema. A built-HTML audit verified all 110 historical proactivity section fragments and four title fragments, including same-language migration destinations, and found zero broken internal fragment links across 29 HTML pages (672 links). Replaying pre-fix content made all four page compatibility cases fail as expected. The new descriptor-directory test failed on all three documents before their stale paths were corrected. Added coverage exercises tracing on/off, bare-number versus explicit menu selection with incomplete classifier configuration, and discovery-menu dispatch through the real CLI stack's headless sink. Runtime type checks and the full repository test suite were not run; no runtime code changes. No browser click test, live LLM, real-VM, or container end-to-end run was performed. ## Risk Documentation-only, with regression tests. Proactivity navigation changes, but page URLs and historical heading fragments are preserved. An old link to content moved between pages lands on a visible migration entry and requires one click to reach the new section; no JavaScript redirect is needed. Existing queued-nudge policy and external-agent isolation limitations are described, not fixed. Roll back by reverting the documentation and associated test changes; no data migration is involved. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A --------- Co-authored-by: zhao.wang <270284818+userName20260323@users.noreply.github.com> Co-authored-by: Codex <noreply@openai.com>
## Summary A 26 MB deck dropped on the composer was refused. The upload ceiling stood at 25 MB, a number it never chose: it was the file viewer's, adopted when `fs.upload` landed beside it. The viewer picked 25 MB because no renderer in the page does anything useful past that size, which is not a reason that applies to an upload. An upload answers a different question. It only has to land a path in the workspace -- a non-image attachment is named to the model and never read into the message -- so the bytes exist to reach the disk, not a renderer. This raises `MAX_UPLOAD_BYTES` to 100 MB and leaves `MAX_VIEW_BYTES` at 25 MB. A file between the two can therefore be attached and not previewed; the constant's comment now states that instead of implying the two move together. `frame_ceiling_for_upload` derives the WebSocket frame ceiling from this constant, so the base64 expansion that once capped attachments at roughly 3 MB cannot come back from raising the number. The page's mirrored copy moves with it. The ceiling that remains is memory, not policy: one upload at this size costs the gateway the frame, the JSON string parsed out of it and the decoded bytes at once, and the parse holds the event loop while it runs. A much larger limit wants the streamed HTTP upload path, not a bigger number here. That is left for its own change. This also sweeps four comments that stated 25 MB as the current ceiling; they name the constant now. The three describing the old 3 MB frame bug keep their number, because they report what was true then. Deliberately out of scope: the composer and the subagent page measure an attachment after base64-encoding it rather than before, so an oversized file is read whole and thrown away while the tab blocks. `uploadRefusalBySize` already exists for the pre-encode check and the knowledge page uses it; wiring the other two callers crosses the composer source contract, so it belongs in its own change. ## Type - [ ] Fix - [x] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Server side, end to end at the size that was refused (26.2 MB), through the real aiohttp WebSocket and the real method: - a frame carrying that file as base64 comes back answered, not disconnected - `fs_upload` accepts it and the bytes land in `uploads/` at the right size - `MAX_UPLOAD_BYTES + 1` is still refused, with `100 MB` in the message Those three ran as a throwaway probe and are not added to the suite. The committed guard is `test_the_frame_ceiling_can_carry_the_largest_upload_the_method_accepts`, which reads the constant rather than a literal and so keeps the two limits coupled at any value. Gates run locally, after rebasing onto github/main at 82a51d8: uv run pytest tests/test_rpc_transport.py tests/test_rpc_files.py tests/test_rpc_console.py -q 172 passed make lint-python All checks passed / 1625 files already formatted make lint-imports clean make lint-deps Success! No dependency issues found. make lint-types All checks passed! npm run gen:check --prefix ui-web generated.ts matches the contract (178 methods) npm run type-check --prefix ui-web clean npm test --prefix ui-web 108 files, 1837 tests passed node ui-web/scripts/check-page.mjs OK node ui-web/scripts/count-shared-globals.mjs OK node ui-web/scripts/check-css.mjs OK node ui-web/scripts/check-class-namespace.mjs OK Page side, in a real browser against a `raven serve` built from this branch, on an isolated RAVEN_HOME: - a 26.1 MB .pptx dropped on the composer attaches: the chip goes solid at "26.1 MB" with no failure note, and the file lands in `<workspace>/uploads/` at exactly 27,400,000 bytes with its non-ASCII name intact - a 105 MB file is refused, and the note names the new ceiling ("105.0 MB", "100.0 MB"), so the limit was raised rather than removed and the message follows the constant The drop was a synthesized DragEvent: it exercises the page's own drop handler and the whole upload path, but not the operating system's part of the drag. No user-facing docs name this ceiling. The one doc mention of 25 MB is the viewer's render cap in `docs/specs/2026-08-26-desk-deliverables-shelf.md`, which this change does not touch. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed ## Risk Behaviour change: an attachment up to 100 MB is now accepted where 25 MB was the wall. Per upload the gateway holds the frame, the parsed JSON string and the decoded bytes at once -- roughly 370 MB peak for a maximal one -- and the JSON parse blocks the event loop for a second or more. Single-user desktop use absorbs that; a shared gateway taking concurrent uploads would feel it. Security: the route is unchanged. `fs.upload` is reachable only over an authorized socket, so the larger allocation is available to the signed-in local reader and to nobody new. What grew is how much memory that reader can ask for in one call. A file between 25 MB and 100 MB can be attached and not opened in the viewer, which still refuses past `MAX_VIEW_BYTES`. Rollback: revert the two constants. Nothing persists, no schema or contract moved, and `frame_ceiling_for_upload` recomputes from whatever the constant says. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A Co-authored-by: arelchan <204152633+arelchan@users.noreply.github.com> Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
) ## Summary Rebuilds the Showcase layout in both READMEs, and replaces one image. **Why the old layout did not hold.** The section had grown to nine cases across three tables of mixed shape. A heading that wrapped to two lines pushed its column's images down. Paired task graphs of unequal height left the artifacts below them starting at different heights. And GitHub's stylesheet stripes every second row while a case spanned three, so the pattern drifted until some headings sat on white and others on the grey. **What it is now.** Each case spans three rows - heading, task graph, artifact - so cells in a row share a top edge and a long heading cannot cascade into the images. Each pair gets its own table, which restarts the stripe and puts the usual gap between pairs. The game run keeps the opening slot and now takes the full width instead of leaving half a row empty. Image widths are set per pair so both sides render to one height, and the narrower image is centred. The divider sentence that separated the old tables is gone, since the headings already say it. **A new comparison board.** The Frameworks artifact is replaced with a redrawn plate: each framework carries its vendor mark, and the orchestration style became a three-way tag - explicit graph or canvas, declared task flow, dynamic at runtime - where the old plate split the six in two. Its alt text said "versus" and named two categories; it names three now. The render is 2000x1547 against the old 2000x1332, so its width is set to 86 percent to stay level with the sweep chart beside it. ## Type - [ ] Fix - [ ] Feature - [x] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification - `make check-large-files` -> exit 0 - `make check-source-language` -> exit 0 - The replacement board was fetched anonymously from its attachment URL: HTTP 200, 2000x1547, and its md5 matched the local source file byte for byte. - Rendered the section through GitHub's own markdown endpoint under the published markdown stylesheet and read back the computed row backgrounds. All five heading rows now report one background; before the split they reported two. - Measured every image box in that render. All four pairs match on top edge, and within 2px on height: | pair | left | right | | --- | --- | --- | | Song / Greece | 379x129 and 379x718 | 379x129 and 379x718 | | Pop / Abstract | 379x127 and 379x713 | 371x127 and 379x713 | | Frameworks / Sweep | 379x126 and 326x252 | 379x126 and 379x252 | | Poster / Explainer | 314x131 and 379x348 | 379x131 and 367x347 | - Confirmed against the repository-rendered README that the game run's bare attachment URL is served as a video player, and kept it on its own paragraph inside the cell so it still is. - [ ] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed ## Risk Documentation only, in both READMEs. One of the seventeen attachments is replaced, with its alt text rewritten to match the new plate's three-way legend; the rest are untouched, and no file is committed to the repository. Rollback is reverting the commit. - [ ] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues Follows #547 and #557. Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
…rounds directory (#553) A campaign asked the shared rounds directory how much had been spent, and the directory answered for every job in it, including jobs its sibling campaigns had submitted. Rounds that keep one directory therefore bill each other, and the bill is cumulative: round three opens already charged for rounds one and two. Measured on a run of 2026-09-11. Five campaigns declared the same rounds directory. Round two had a budget of 130 and had actually used 34.6, but the eighteen jobs the earlier rounds had left in the directory brought the reading to 125.6, so the remaining 4.4 would not cover one more trial at 5.9 and the round stopped. 142 GPU-minutes went unspent and that round's multi-seed conclusion was lost. A sibling run that gave each round its own directory was not affected, which is why this reads as a directory-hygiene problem until you look at the meter. Two changes. Spend is now counted from the campaign's own ledger, so a job a sibling submitted is not billed here; the filter sits above the existing timeline arithmetic, so same-device overlap is still charged once and two cards in parallel are still charged twice. And declare refuses a rounds directory a live sibling campaign already keeps, naming the sibling, so the two rounds do not land in one directory in the first place. Type: fix Verification: - All oncall tests: 775 passed, 1 failed. The one failure (test_the_products_tool_face_equals_the_forks_config_intent, a missing optional a2a_client dependency on this machine) also fails on 0464aa1, so this adds six passing tests and no new failures. - The six new ones pin: a sibling's job is not billed; the campaign's own two trials overlapping on one device are still billed 20 minutes once; a job this executor just submitted is billed before the ledger has caught up (no refund window); billing_only leaves a foreign backend untouched; declare refuses a directory a live sibling keeps; a concluded sibling, and the campaign itself, do not count as a clash. - The first test also pins the measured number: the same fixture directory reads 125.6 unfiltered, which is 34.6 of its own plus 91.0 of its siblings. Risk: low for the meter, moderate for declare. The meter can only ever bill less than before, and a run that already gives each round its own directory reads the same. Declare now refuses something it used to accept, so a caller that deliberately shares a directory between live campaigns will see an error naming the sibling; concluded siblings and re-declaring the same campaign are both still fine. --------- Co-authored-by: xiaotian.luo <xiaotian.luo@thetahealth.ai> Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
…rail groups by it (#542) ## Summary A conversation can now be started in a folder of the person's choosing, from the page, the way a code editor's folder switcher works: on the new-task screen the composer carries a folder chip that opens No folder / Recent / Open folder..., and the folder picked rides on `session.create` as its `workdir`. The engine already pinned a session's turns to that directory (shell, filesystem, deliver and media tools follow the binding); the page had no way to set it. The choice is made once. After the first message the chip only reports the conversation's folder and cannot be pressed, which is exactly the engine's semantics: the pin is taken at creation and every later turn reads it. The rail splits the rest of the list on the same fact: conversations pinned to a folder sit under their own collapsible heading, each row wearing the folder's name (full path on hover); conversations on the policy default keep the old heading, renamed to say what it is once the other group exists. Pinned and cron rows are untouched. Two small backend additions carry it: - `session.list` items gain an optional `workdir`, the stored override or absent, which is what the rail groups on. - `fs.dirs` lists the subdirectories of one absolute directory (home when omitted), directories only, dotfiles omitted, capped at 500, and marks each entry -- and the listed directory itself -- with whether `session.create` would accept it (`validate_override`: not the agent home, not its ancestors, not its memory/skills/sessions trees). The picker greys those out instead of offering a folder the create would refuse. "Recent" is derived from the rail's own rows, newest activity first, so nothing new is stored and every browser sees the same list. The offline demo page installs neither hook; the chip works and "Open folder..." says it needs a gateway. ## Type - [ ] Fix - [x] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification - `uv run pytest tests/test_rpc_console.py tests/test_rpc_session.py tests/test_rpc_contract_shapes.py tests/test_rpc_schema_match.py tests/test_rpc_registration.py tests/test_ui_language_repaint.py`: 661 passed. New: six `fs.dirs` tests (subdirectories only, dotfiles omitted, sort, `ok` on entries and on the listing, home default, root has no parent, relative/file/missing/unreadable refused, a child vanishing mid-listing), one `session.list` test (the override or null), and the `fs.dirs` round trip against its declared result model. - `scripts/coverage_gate.py diff --base-ref upstream/main`: 97.83% of changed executable lines (45/46). - `ruff check` and `ruff format --check`: clean. - ui-web: `npm run gen:check` (179 methods, in sync), `npm run type-check`, `npx vitest run`: 109 files, 1849 tests passed. New: `shell/workdir.test.ts` (draft/locked states, recents deduped, stage and unstage, the browse walk with greyed entries and a disabled use button, the offline and refused-path notes), three rail tests (the two groups, the tag on either separator, folding), two `open-conversation` tests (the staged folder rides the create and lands on the row; none staged creates with none). - `python ui-web/build.py`, `check-page.mjs`, `check-css.mjs`, `check-class-namespace.mjs`: OK. - ui-tui: `npm run gen:rpc`, `npm run gen:i18n`, `npm run lint:rpc`, `npm run lint:i18n`, `npm run type-check`: in sync and clean. - On a real gateway, in the served page: opened the picker on the new-task screen, browsed home into a demo folder, picked it; the chip named the folder; promoted the draft; the chip locked with the fixed-once-started note and no longer opened; `fs.list` for the new session rooted at the demo folder; a turn asking the agent to run `pwd` answered with that folder; `session.list` carried it as `workdir`; the rail showed the row under the folder group with its tag, and the other rows under the no-folder heading. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed ## Risk - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes `fs.dirs` lists directory names on the gateway's machine to an authenticated page. The page already lists and reads files under a session's working directory, and the agent's shell tool reaches the whole machine, so this widens what the page can see by directory names only; dotfiles stay hidden and the agent's own data is marked unusable, matching `session.create`'s refusal. A page that never presses the chip behaves exactly as before: `session.create` is called with no `workdir`, and a `session.list` consumer that ignores the new optional field sees the old shape. Reverting the commit restores the previous page and contract; sessions already pinned keep working, since the pin was the engine's feature already. ## Related Issues N/A --------- Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
## Summary Two commits. The first fills the Launch WebUI section of both READMEs, which carried three `Screenshot placeholder` lines and no images, with two figures each: the new task page and the subagents page. The second lands the Raven mark those figures show. On the figures: - Two, not three. The third placeholder asked for a task graph, which cannot be captured from stored data: the web DAG panel is driven by live `dag.*` events and its disk restore is gated on a localStorage note that only a live event writes, so showing one would have required a real `run_subagent_dag` turn. That slot is dropped rather than filled with something its caption does not match. - They come from an isolated instance with its own `RAVEN_HOME` and an empty session store, so no conversation titles, memory entries or workspace file names appear. The A2A block was removed from its config, and the config carries no channels and no cron, so the throwaway instance had no outward side effects. Its home, which held a copy of real provider keys, was deleted after the run. - Each README gets its own language, captured with `language` set to `en` and to `zh`. Following the convention already in these files, `alt` text stays English in both while the caption follows the prose language. On the mark: - The new file is a dark bird on an antique-gold halo. Against the dark theme's surface that ink measures 1.07:1, and no tone filter can lift it because it is not one colour to invert, so it carries an opaque plate rather than the `prefers-color-scheme` rule the previous file used. Ink against plate measures 18.29:1 on both themes. - The plate is a card, so the mark fills its frame the way `miromind.svg` already does. Keyed by file rather than by `data-agent`, which carries the preset and raven's own rows have none. - The halo stroke goes from 1.6 to 3.2. Against a 64 unit canvas the old width lands under one device pixel at the two smaller slots the mark renders in, 0.45 at 18px and 0.65 at 26px, so the ring washed out wherever the roster draws it. Reviewer note: the figures were captured before the stroke went to 3.2, so the halo in them is thinner than what this branch now renders. Everything else in them is current. Re-capturing would mean a second upload for a difference of half a device pixel, so they are left as they are. Assets are hosted as a comment on issue #505 rather than committed, per the repository asset rule. ## Type - [ ] Fix - [ ] Feature - [x] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification ``` .venv/bin/python -m pytest tests/test_ui_agent_marks.py -q -p no:randomly -> 7 passed .venv/bin/python -m pytest tests/test_scope_canon.py tests/test_docs_site.py \ tests/test_docker_runtime.py -q -p no:randomly -> 27 passed npx --prefix ui-web vitest run --root ui-web \ scripts/agent-mark-css.test.mjs src/shell/agent-mark.test.tsx -> 2 files, 18 passed make check-large-files -> exit 0 make check-source-language -> exit 0 make check-commits -> exit 0 ``` The gates were run over a committed range rather than an empty diff, and the mark suite was re-run afterwards because `make check-commits` can rebuild the virtualenv under it. The mark change was rendered and measured rather than eyeballed: at 1440px the mark is 28px, and the two candidate files differ only in a sub-pixel halo and a plate that is near the tile colour in the light theme, so the two are not visually separable at full size. The renders were compared by pixel signature against standalone renders of each file. Every uploaded asset was fetched back with a GET and compared by sha256 against the local render; all four are byte identical. A HEAD request is not a valid check here, because user-attachments URLs answer 403 to HEAD and 200 to GET. Both README sections were rendered through the GitHub markdown API to confirm the figures resolve, the alt text survives and each file references the assets for its own language. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed ## Risk Two user-visible changes. Both READMEs render two figures in Launch WebUI where three placeholder lines used to be. The Raven mark in the subagents roster and in every other slot it draws in is a different image, larger in its frame, and no longer answers the theme from inside its own file. Rollback is a revert of either commit independently; they touch disjoint files. The hosted assets can stay either way. Backward compatibility is not applicable: no interface, config key or stored value changes. - [x] Security impact considered - [ ] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues #505 --------- Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
Restore the missing icon mappings for the overview cards in both locales.
## Summary The four Launch WebUI figures were flat rectangles butted against the page. They now carry a 24px corner radius and a soft drop shadow, so each reads as a window floating on the page rather than a slab of pixels pasted into it. - The capture is untouched. The same 1440x900 content sits inside a 1588x1048 frame; the extra 74px on each side is the room the blur needs, and nothing in the interface was re-shot. - The canvas is RGBA. Both the rounded corners and the shadow are transparent, so GitHub composites them correctly on the light theme and on the dark one. On the dark theme a black shadow against `#0d1117` is nearly invisible by construction, and only the rounding reads; a tinted shadow would fix that and would look dirty in the light theme, so the shadow stays neutral. - `width` stays at 90 percent. Because the padding is inside the image, the interface itself now renders about nine percent smaller than before. That is the space the shadow needs to read as depth, so it is kept rather than compensated for. - Both language pairs are replaced: English in `README.md`, Chinese in `README.zh-CN.md`. The corner mask is drawn at 4x and resampled down, because PIL's `rounded_rectangle` is aliased at 1x and leaves a visible stair-step on a 1440px edge. Shadow parameters, for anyone reproducing them: radius 24, Gaussian blur 34, offset 16 down, 30 percent black. Assets are hosted as a comment on issue #505 rather than committed, per the repository asset rule. ## Type - [ ] Fix - [ ] Feature - [x] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification ``` .venv/bin/python -m pytest tests/test_scope_canon.py tests/test_docs_site.py \ tests/test_docker_runtime.py -q -p no:randomly -> 27 passed make check-large-files -> exit 0 make check-source-language -> exit 0 make check-commits -> exit 0 ``` The gates were run over a committed range, not an empty diff. Each uploaded asset was fetched back with a GET and compared by sha256 against the local render; all four are byte identical. A HEAD request is not a valid check here, because user-attachments URLs answer 403 to HEAD and 200 to GET. Both README sections were rendered through the GitHub markdown API, which confirms each file now references its own language's new asset and that the alt text survived. The rounding and the shadow were inspected at 1:1 on a white and on a `#0d1117` backdrop before upload, to confirm the corner is clean and the edge carries no light fringe on the dark theme. The diff is four URLs, each appearing twice as the link target and the image source. Alt text, captions, width and layout are unchanged, and no stale asset id survives in either file. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed ## Risk Presentation only. Both READMEs render the same four captures with rounded corners, a shadow, and slightly smaller interface content. No code, no build output, no behaviour change. Rollback is a revert of the single commit; the previous assets are still live on issue #505, so the old figures would come back intact. Backward compatibility is not applicable: no interface, config key or stored value changes. - [x] Security impact considered - [ ] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues #505 Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
…#583) ## Summary Refreshes the Raven-Design visual design figure in both READMEs with the resupplied Opus-5 evaluation. Raven-Design and Claude Code moved on all three benchmarks in that figure; the GPT5.6-Luna rows and the Hermes column carry over unchanged. | Benchmark (Opus-5) | Raven-Design | Claude Code | Hermes | | --- | --- | --- | --- | | ArtifactsBench Dashboard | 64.4 -> 66.3 | 63.6 -> 62.8 | 64.4 (unchanged) | | ArtifactsBench SVG | 83.5 -> 86.2 | 82.7 -> 83.4 | 83.0 (unchanged) | | GDPVal | 89.4 -> 91.9 | 88.7 -> 89.2 | 89.1 (unchanged) | Two consequences of the update are worth naming, because they change what the figure shows rather than only the bar lengths. Raven-Design now leads Hermes on all three benchmarks, where ArtifactsBench Dashboard was previously a tie at 64.4. And Claude Code now leads Hermes on ArtifactsBench SVG and GDPVal while trailing it on ArtifactsBench Dashboard, having previously trailed on all three. The chart was re-rendered with the same renderer, the same palette and the same layout as the figure it replaces, so only the six Opus-5 bars and their value labels differ. The asset is hosted as a comment attachment on #477, per the repository rule that image files are not committed. Both READMEs reference the same asset, so the diff is one URL swap per file; the alt text, the anchor, the width and the caption are untouched. ## Type - [ ] Fix - [ ] Feature - [x] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Chart data, before rendering: - Built the eighteen expected plotted values by hand from the supplied numbers, then asserted the renderer's readout equal to them. Diffed against the previous value set: exactly six cells differ, all of them Opus-5 Raven-Design or Claude Code. - The renderer's own guards passed: `18 values verified; no overlapping labels`, plus its assertions that no text escapes the 2000px canvas and that the output is 2000px wide on a white background. Uploaded asset, after rendering: - `curl -sL <asset-url> | sha256sum` against the local render: identical, 141368 bytes, `e175bc24e7d6208ab26d2aaaab213431585398c0de9705f058dfd4092023924f`. Checked with GET, not HEAD, because user-attachments answers HEAD with 403. - `POST /markdown` on the changed line returns both `href` and `src` pointing at the new asset. Repository state: - Asserted the retired asset id occurred exactly twice per README (anchor plus image) before replacing, and zero times after. `git grep` over the whole tracked tree finds no remaining reference to it. - `git grep` for the six retired score values across every markdown, text and YAML file: no hits, so no prose quoted the numbers the figure carries. - `uv run --frozen --all-extras pytest tests/test_living_docs.py tests/test_docs_site.py`: 18 passed. `tests/test_cli_onboard_commands.py -k readme`: 1 passed. These are the tests that read the repository-root READMEs. - `make check-large-files`, `make check-source-language`, `scripts/check_commit_messages.py origin/main..HEAD`: all exit 0. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed ## Risk User-visible change is the Raven-Design visual design figure in the English and Chinese READMEs. No code, no configuration and no dependency is touched, so there is no runtime behaviour change and nothing to be backward incompatible with. Rollback is a one-line revert per README: the retired asset `061b5818-595b-4392-b9b6-53a1effb8347` stays hosted on #477 and keeps serving, because replacing a README reference does not delete the attachment it pointed at. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues #477 Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
) ## Summary `raven skill block` could leave a config file that Raven then refuses to load. `set_skill_blocked` wrote `data.setdefault("skillForge", {})` unconditionally. Dict keys match byte for byte, but the schema treats `skillForge` and `skill_forge` as one field (`alias_generator=to_camel` with `populate_by_name`), so a config file spelling the block `skill_forge` gained a second block instead of having its own one patched. `RavenConfig` forbids extra inputs, so `load_raven_config` binds the alias and rejects the leftover field name: `raven agent`, `raven gateway` and the TUI RPC server all then refuse to start. `raven status` keeps working throughout, because the base loader pops extension keys before validating. That asymmetry is what makes the defect quiet: the command most people reach for to check a config is the one that cannot see it. The same misread has two softer faces, both closed by the same change. `raven skill unblock` became a silent no-op on such a config, reporting success while the skill stayed blocked; and blocking an already-listed skill reported it as newly blocked. Both read the blocklist out of the empty block the call had just created. The fix probes the spelling the file already carries and reuses it. That is not a new shape: `loader.py` already does exactly this for the `skillRouter` and `everosSkillLight` migrations, and `set_sentinel_nudge_quota` does it for `sentinel.nudge_policy`. Among the top-level blocks a config writer creates, `skillForge` is the only one whose camelCase and snake_case spellings differ at all. The second commit closes what the first one opened. Making the written key conditional made two lines that report the write false on a snake_case file: the log line, and the result line of `raven skill block` and `raven skill unblock`, both of which named `skillForge.blocklist` unconditionally. Measured on a `skill_forge` config, the write landed on `skill_forge.blocklist` while both messages still said `skillForge.blocklist`, sending an operator to grep their own config for a key it does not carry. They now report the setting rather than the file key, which is what `set_sentinel_nudge_quota` already does: it respects the same two spellings and logs `sentinel nudge quota patched` with no key path at all. The `--help` text and the module docstring keep naming `skillForge.blocklist`, because they describe the command rather than the write. Checked and deliberately not fixed, with reasons: - A config already corrupted by the old code does not heal. Running `raven skill block` again now selects the camelCase block and leaves the snake_case one in place, so such a file still needs a hand edit. Merging the two blocks automatically is lossy and ambiguous, and belongs in a migration argued on its own terms rather than inside a fix whose job is to stop producing the state. - `update_cron_config` writes `to_camel(key)` unconditionally, which is the same shape one level down, on field keys rather than block keys. Its failure is milder and different: `CronConfig` does not forbid extras, so a file carrying both `default_timezone` and `defaultTimezone` loads with the camelCase one winning. Measured directly. Nothing breaks, but a later hand edit of the snake_case key is silently dead. Pre-existing, a different failure mode, and in another config surface. - The three operator-blocklist refusals in `raven/agent/tools/skill_hub.py` name `skillForge.blocklist` in text the model reads. These were already inaccurate for a hand-written `skill_forge` config before this branch, since the loader has always accepted both spellings on the read side. Not introduced here, and fixing them widens the diff into the agent tool surface. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Run on Linux with Python 3.12.13, on a branch rebased onto the current `main` (which now carries the A2A work). The rebase conflicted only in `tests/test_config_update.py`, where both sides appended a section at the end of the file; both sections were kept. The rebased commit's diffstat is byte for byte the pre-rebase one, +70/-1 over two files, so the resolution neither absorbed nor dropped anything. - `uv run --frozen --all-extras pytest tests/test_config_update.py` -> 47 passed - `uv run --frozen --all-extras pytest tests/ -k "config or skill_forge or skill_block"` -> 1709 passed, 33 skipped, 0 failed - `make lint-python` -> ruff check clean, 2008 files already formatted - `make lint-imports` -> 10 contracts kept, 0 broken - `make lint-deps` -> no dependency issues - `make lint-types` -> all checks passed - `make check-commits`, `make check-large-files`, `make check-source-language` -> exit 0 Neither commit's tests are theatre, and each was checked on its own. Reverting only the first commit's production hunk turns 4 of its 5 tests red: duplicate key, `load_raven_config` ValidationError, unblock no-op, and blocklist read from the wrong block. The fifth covers camelCase behaviour that predates the fix and correctly stays green. Restoring the message wording the second commit replaced turns its new CLI regression red on the `skillForge` assertion, and restoring the fix turns it green. End to end with an isolated `RAVEN_HOME`: a config carrying the duplicate block makes `raven agent` exit with `ValidationError: skill_forge Extra inputs are not permitted`. The same starting config driven through the fixed `raven skill block` gets past config validation and reaches the expected missing-API-key error, with the user's existing `auto_install` setting preserved beside the new blocklist. Malformed blocks were checked too. `skillForge: null` beside a `skill_forge` table used to raise `AttributeError` and now works. `skillForge: null` alone, and `skillForge` set to a string, still raise `AttributeError`, unchanged from before. File mode was measured rather than assumed, because the new A2A onboarding narrows `config.json` to 0600 and this branch writes to the same file: after `initialize_a2a_server`, a `raven skill block` leaves the file at 0600, since `atomic_update` preserves the mode it finds. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed `make lint` also runs `lint-ui`, which fails in a fresh worktree with `ERR_MODULE_NOT_FOUND` because `ui-web/node_modules` is not installed there. This change touches no JS or TS file, so the Python lint targets above are the gates that cover it. No user-facing docs were needed; the helper's docstring states the casing contract. ## Risk Behaviour changes only for a config file that spells the block `skill_forge`. For a file using the camelCase spelling, which is what onboarding writes, the bytes written are identical to before. Two changes are visible to a snake_case user, and both are the point of the fix: the blocklist is written into their existing section rather than into a new one, and `raven skill unblock` now actually removes an entry instead of reporting success and doing nothing. One change is visible to every user regardless of spelling: the result line of `raven skill block` and `raven skill unblock` now reads `Skill blocklist = [...]` where it read `skillForge.blocklist = [...]`, and the corresponding log line drops the key path the same way. This is wording only; no value or exit code changes, and the `--help` text still names `skillForge.blocklist`. Rollback is a revert of the two commits. No migration runs, so nothing has to be undone on disk. Reverting only the second leaves the first in place with the messages inaccurate on a snake_case config, so revert both or neither. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes No permissions, credentials or file modes are changed; the 0600 measurement above is the check behind that sentence. ## Related Issues N/A --------- Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
…527) ## Summary `restrict_to_workspace` is an operator's promise that the agent cannot reach outside its workspace. It kept that promise only for paths spelled literally. `ExecTool._check_workspace_restriction` extracted candidate paths from the command as typed, then called `os.path.expandvars` on what it had extracted. Extraction wants a `/` sitting on a word boundary, and `$HOME/` puts a letter there, so a variable spelling produced no candidate at all and the expansion never saw the text that hid the path. The expansion sat downstream of the extraction it needed to feed. An unset name is the sharpest spelling: the shell drops it, so `$NOPE/etc/shadow` is `/etc/shadow`, and there is no literal form to fall back on. Nothing else was looking either, because `cat` is read-only and `default_tier` answers `allow` for a read-only command without asking; for a read the fence was the only thing in the way. Measured with the fence on, before this branch and after it: | command | before | after | |---|---|---| | `cat /etc/shadow` | refused | refused | | `cat $HOME/.ssh/id_rsa` | ran | refused | | `cat "$HOME/.ssh/id_rsa"` | ran | refused | | `cat $NOPE/etc/shadow` | ran | refused | | `cat ${NOPE:-/etc/shadow}` | ran | refused | | `cat ${UNSET:-$HOME/secret}` | ran | refused | | `cat ${PWD%/*}/outside.txt` | ran | refused | | `cat ${PWD%$PWD}/etc/passwd` | ran | refused | | `cd /; cat etc/shadow` | ran | refused | | `cd -- /; cat etc/passwd` | ran | refused | | `sh -c 'cat $HOME/secret'` | ran | refused | | `env sh -c 'cat $HOME/secret'` | ran | refused | | `command cd /; cat etc/shadow` | ran | refused | | `cat $PWD/notes.txt` | ran | ran | | `echo 'keys go in $HOME/.ssh'` | ran | ran | | `cd subdir; cd ..; ls` | ran | refused | | `cd missing; cd ..; ls` | ran | refused | | `cd subdir \|\| cd ..; ls` | ran | refused | | `cd subdir \| cat; cd ..; ls` | ran | refused | | `cd subdir & cd ..; ls` | ran | refused | | `pushd ..; ls` | ran | refused | | `test -d subdir && cd subdir && echo r; cd ..; ls` | ran | refused | | `cd subdir && cd .. && ls` | ran | ran | | `cd subdir && cd ..; cat notes.txt` | ran | ran | The fix expands first, with the shell's own rules, then scans. Each rule is taken from the shell rather than invented: - an unknown name expands to nothing, as the shell does; - the environment that decides is the child's allowlisted baseline, not `os.environ`, so a name the command will never be given reads as unset; - single quotes expand nothing, so `echo 'keys go in $HOME/.ssh'` stays a sentence; - an escaped `$` yields the bare character, which this level does not expand but a nested shell handed it does; - a brace body goes back through the same pass, so a parameter inside a fallback word or a trim pattern is read on the same terms as one anywhere else; - `%NAME%` is read only where cmd.exe is the shell, and there without regard to case, because that is what cmd does and `%VAR%` means nothing to `sh`; - `PWD` resolves to the directory the command runs in, because a shell sets it from there whatever this process inherited. A nested shell's payload is scanned in the shell that runs it, recursively, through the policy's own embedded-shell helper and depth bound. Both readers of a segment unwrap `env`, `sudo`, `command` and leading assignments first, through the policy's own helper, so the fence and the deny list agree on what a segment runs rather than keeping two lists. Stepping out is the same escape as reaching out, so a `cd` out of the workspace is refused as well. Only the destinations the scan cannot see needed this: `/` has nothing after it, `..` is not absolute, and a `cd` with no argument names `$HOME` by saying nothing. `cd /etc` was already refused, because `/etc` is a path like any other. Each `cd` starts from where the last one landed, but only where a separator proves it got there: `&&` runs its right side because the left one returned zero. After anything else the move may not have happened, so both readings are held and a later step that leaves from either is refused. ### The posture Bypasses of one family kept appearing: everything the fence did not model was allowed, so every gap was silent and the next one would be too. The contract is now stated on `_check_workspace_restriction`: **A parameter expansion is resolved faithfully or the command is refused.** The set of spellings is finite, so the residue is closed. Substitution, case folding, indirection, offsets and a brace this cannot parse each refuse with the construct named. A refusal is visible and can be argued with; the failure it replaces could not be. **Command substitution is a declared limit, not a gap.** It holds an arbitrary program, so no textual guard can resolve it, and refusing it would refuse `echo "built at $(date)"` along with everything else. Paths written literally inside one are still scanned. Two constructs are exempt from the refusal for a reason that does not depend on modelling them: `${#NAME}` and `$((...))` yield numbers, and a number cannot name an absolute path. The same rule decides what the directory walk may carry. A `cd` the shell may not have run cannot be carried, and the separator is what says whether it ran: a bracket or a pipe puts it in a subshell, `||` runs it only when the one before failed, `;` continues whether it failed or not. Only an `&&` chain is knowable, and the chain is the unit rather than the step -- it stops at its first failure, so what follows one inherits the position after any prefix of it. The walk carries those prefixes and unions them where the chain ends. The directory walk took three rounds, and the second and third findings were both in code written to close the one before. Recorded because the shape repeats: the first version advanced the walk on every `cd`, so a `cd` into a directory that does not exist left the walk a level deeper than the shell and the `cd ..` after it read as a return. The second read the separator after each `cd`, which closed that and four more spellings of it (`||`, a pipe, a background `&`, a bracket) but treated `&&` as proof the `cd` ran -- and a condition earlier in the same chain can skip it. The third makes the chain the unit: it stops at its first failure, so what follows inherits the position after any prefix of it. Every one of these was measured against bash rather than reasoned about, in a workspace with and without the directory the command names, because which branch leaks depends on that. `pushd` was found the same way and is the one finding here that came from neither review round. Widening the fence from "no outside path is named" to "and the command does not walk out" is the one deliberate contract change beyond the posture. Without it the rest is bypassed by `cd /etc; cat shadow`. Checked and deliberately not changed: - **The permission tier.** A read-only command is tiered `allow` and asks nobody. That is why this defect mattered, but it is a separate control that applies to every install including the ones with the fence off, so changing it is not in scope for making the fence keep its own promise. - **A symlink inside the workspace pointing outside.** `cat link` names no absolute path, so there is nothing for a textual scan to refuse. Both sides of the comparison already resolve physically, so a named path cannot be laundered through a symlink. - **A path computed at run time**, such as one decoded from base64 in a substitution. The sandbox executor is the boundary for both of these, which is what the code says about itself. Blast radius: `_check_workspace_restriction` returns before any of this when `restrict_to_workspace` is off, which is the default, so an install that did not raise the fence runs none of the new code. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Branch scoped with three dots against the rebased base: `git rev-list --left-right --count origin/main...HEAD` reports `0 14`. - `COLUMNS=200 uv run --frozen --python 3.12 --all-extras pytest -q` 2 failed, 23826 passed, 112 skipped in 636.80s under coverage. A second run of the same head gave 1 failed, 23803 passed. - The one failure is `test_install_script.py::test_resolve_node_dir_answers_each_case_it_exists_for` ("node without npm provisions a runtime"). It is not introduced here: it fails identically with this change removed, and it fails the same way when its file is run alone, so it is neither this branch's nor an ordering artifact. - A second failure is intermittent and its attribution is left open rather than claimed: `test_agents_research_tools.py::test_a_thin_dataset_card_falls_back_to_the_metadata_the_table_needs` (`KeyError: 'served_url'`). It appeared in two of three full-suite runs of this branch and not in the one full-suite run taken with this change removed, which is a single baseline sample and not enough to rule the branch out. Against that: the test passes when its file runs alone both with and without this change, and it passes when its file and the new one run together in a single process, so the new tests do not reach it directly. The mechanism that does fit is in that file: it selects a process-wide `ContextVar` 29 times and never restores the token, so its state leaks to whichever tests share its xdist worker, and the default `--dist load` decides that per run rather than by a fixed split. Adding tests changes that distribution without changing any code the test touches. Not changed here, since the isolation defect is that file's own. - `uv run --frozen --python 3.12 --all-extras pytest tests/test_shell_workspace_fence.py -q` 43 tests, 116 cases, all passing. Every case added for a reported bypass was watched failing before its fix; the rest are over-blocking guards, which pass on both sides by design and are there to fail if a fix reaches too far. - Reverting only the first commit's hunk in `raven/agent/tools/shell.py`, with the new tests left in place, turned 13 of the then 38 cases red: every case written for the defect, and none of the over-blocking guards. The tree was restored afterwards and `git diff HEAD` confirmed empty, since a three-way reverse-apply stages what it applies. - `make coverage-diff` (the gate that failed on the first revision): `Diff coverage: 97.89% (186/190 executable changed lines)`. - Recursion through brace bodies terminates because a body is strictly shorter than the text holding it and a value is substituted rather than re-read; checked at 200 levels of nesting and 500 parameters in one body, both answer without a `RecursionError`. - Fence verdict against a real bash, fourteen shapes, each run in a workspace with and without `subdir`: `holes: 0 over-blocks: 0`. A hole is allowed-and-leaks; an over-block is refused-while-neither-branch-leaks, which is why both fixture states are needed to claim either. - The walk's separator rule is held from both sides. Never ending an `&&` chain turns 10 cases red, every escape class; ending it at every separator turns 6 red, the proven walks that must keep running. Dropping `pushd`/`popd` recognition turns 3 red; recognising `popd` without its worst-case reset turns 1. - `pytest tests/test_*shell*.py tests/test_permissions*.py tests/test_*exec*.py tests/test_sandbox*.py` -- 958 passed. - `ruff check` and `ruff format --check` clean on both changed files. - `scripts/check_commit_messages.py`, `scripts/check_large_files.py` and `scripts/check_source_language.py` over `origin/main..HEAD`: all pass. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed The docs box is unchecked because nothing documents the fence's path semantics. The only mention outside the code lists `restrictToWorkspace` as a config field, which this does not change; there was no published claim describing the old behaviour. Not reproduced here: the end-to-end refusal on Windows. A Windows path is not absolute to a POSIX `Path`, so a verdict test would turn on the host running the suite rather than on the expansion. The percent-expansion tests assert the expanded text and drive the Windows branch through a class attribute instead. ## Risk With the fence on, commands that name an outside path through a variable, that `cd` out of the workspace, or that use a parameter expansion the fence cannot resolve are now refused where they previously ran. The first two are the control doing what its name says. The third is the posture, and it is the one that can refuse a command that would have been harmless: `echo ${NAME^^}` names no path and is refused because nothing here can prove it does not. That trade is deliberate. An over-block is visible to the operator and can be argued with or reported; an under-block is silent, and eight of them were found across two review rounds on this branch before it was ready. One shape that was refused mid-branch now runs: a `cd` sequence chained with `&&` that returns to the workspace, such as `cd subdir && cd .. && ls`. The `;` spelling of the same walk is refused, and correctly: when the directory does not exist the `cd` fails, the `;` carries on from where the shell stood, and the `cd ..` leaves the workspace. Measured against bash, it reads the parent. Nothing changes for a default install, where the fence is off and this code does not run. Over-blocking was the failure mode guarded against throughout, and one instance of it was introduced and then removed within this branch: reading `%VAR%` on POSIX refused `echo '%HOME%/notes'`, which `sh` prints literally. Workspace-relative `$PWD` use, single-quoted text containing a `$`, `cd` inside the workspace and back, `cd` into an operator-added extra directory, `cd -- subdir`, a wrapper in front of inside work, a nested payload that stays inside, a fallback that resolves inside, a trim that resolves to a basename or matches nothing, `${#NAME}`, `$((...))`, `date +%Y%m%d`, `printf '%s%s'` and `cut -d/` style delimiter arguments all still run, each with a test. Rollback is reverting these commits; the fence returns to its previous behaviour with no migration and no persisted state involved. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A --------- Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
…544) ## Summary At teardown the store pipeline counted every unsettled record as lost, including one a worker had already handed to the memory service. Cancelling that worker does not cancel its request, so the service finished the write anyway, and the host told the user the turn was gone while it was being indexed. Measured on three one-shot CLI runs against a live EverOS: the client gave up at its budget and `atomic_facts_extracted` landed for the same session 32 to 48 seconds later, every time. The message was wrong twice over -- the turn was written, and the service was not unavailable. Review found that the first version of this fix replaced that error with its mirror. The mark is set on entering `MemoryBackend.store()`, which is not a handoff: a backend that persists only after an await, cancelled mid-await, leaves `DrainOutcome(lost=0, in_flight=1)` with the backend's persistence list empty, and the notice then said the turn had reached long-term memory. A slow connection cancelled before its request body goes out has the same shape. `MemoryBackend` declares no delivery guarantee, so no evidence from the transport is available without extending the protocol and asking every backend to implement it. This branch takes the other route and preserves the uncertainty instead. What changed: - `StorePipeline.drain` returns `DrainOutcome(lost, in_flight)` instead of one count. A record is marked in flight around the backend call, and the mark is read before the collect await that unwinds the cancellation, because that unwinding clears exactly what is being counted. - `AgentLoop.drain_backend_stores` returns the outcome and logs the two apart: a warning for a turn this side watched fail, and for an unsettled one a line saying the drain stopped waiting and cannot see whether the service finished it. - `report_dropped_memory_writes` becomes `report_memory_write_outcome`, since it no longer reports only what was dropped. An unsettled turn reads "were still being written when shutdown stopped waiting; whether the memory service finished them is not known here" -- the outcome, not a verdict on it. - `DrainOutcome` is published on the memory engine's face beside `StorePipeline`, so callers reach it the way `test_memory_engine_face` requires. Its docstring now carries both cases, with the EverOS measurement scoped to the one it measured. - A test drives the `raven agent` teardown that renders the outcome. It had no coverage at all: the shared harness stubs `maybe_build_memory_backend` to `None`, so the branch guarding those two lines never ran. `in_flight` keeps its name deliberately. This repository already uses the word for started and not settled -- `swap_in_flight`, `outbound.in_flight`, a DAG node in flight -- with no delivery sense, so renaming it would detach the field from a vocabulary that is already right. The delivery claim lived in the prose, and that is what changed. Deliberately unchanged: the 2.0s drain budget. It is far shorter than the write it waits on -- one store measured 16.98s end to end, of which the flush leg was 12.1s -- and raising it to cover that would park every CLI exit for as long while buying nothing, since the service completes a request it received either way and nothing needs to recall the turn that just ended. That reasoning sits beside the constant, because the gap between the two numbers reads like a defect and the next reader would otherwise close it. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Run on this branch's head, Python 3.12.13, all extras synced, rebased onto the current default branch. - `pytest tests/test_agent_loop_backend_dispatch.py tests/test_providers_factory.py tests/test_cli_agent_commands.py tests/test_memory_engine_face.py tests/test_cli_gateway_commands.py` -- 171 passed, 0 failed. - Full suite: CI on this head runs the suite in four shards: 23742 passed, 117 skipped, 0 failed. A local full run reports one failure, `tests/test_install_script.py::test_resolve_node_dir_answers_each_case_it_exists_for`, which was run again in a clean tree extracted from the default branch at the commit this branch is based on and fails there with the same assertion and the same value; the CI runner, which is not root, passes it. 0 introduced either way. The one failure is `tests/test_install_script.py::test_resolve_node_dir_answers_each_case_it_exists_for`. It was run again in a clean tree extracted from the default branch at the commit this branch is based on, and fails there with the same assertion and the same value. 0 introduced. - Diff coverage against the default branch: `Diff coverage: 93.94% (31/33 executable changed lines)`, measured by the `coverage gates` job on this head. The ratchet moved with it: line +1.81pp, branch +2.74pp. The remaining uncovered pair is the same teardown inside `gateway_commands.run`, a 634-line function whose harness would cost far more than the one in `agent_commands`. - `make lint-python`, `make lint-imports`, `make lint-types` -- all exit 0. - `make check-commits`, `make check-large-files`, `make check-source-language` -- all exit 0. - The rebase was checked with `git range-diff`: every pre-existing commit replays `=`, and none of the commits the base gained touches a file this branch touches, so the changed-line set the coverage gate measures is unchanged. Neither round's tests are theatre, and each was checked on its own: - First round: the tests were watched failing on the missing attribute and the missing function. Two mutants were then run against them -- moving the in-flight read below the collect await reddens two of them, which is the constraint the comment beside that read now names; moving it below `cancel()` alone leaves them green, correctly, because `cancel()` only schedules. - Second round: the notice test was watched failing on the delivery claim before the wording changed, and the end-to-end test was watched failing on the log line. Removing the `report_memory_write_outcome` call from the `agent` teardown turns the new CLI test red on its stdout assertion; restoring it turns it green with `git diff HEAD` empty. One thing that would have gone unnoticed: `capsys` captures nothing from loguru here, because the sink holds the real `sys.stderr` from when the handler was added, so the first version of the end-to-end test passed while reading an empty string. It adds its own sink now and asserts the sink saw something, so the log half cannot pass unread. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed No docs change: the only user-facing text is the shutdown notice itself, and no document quotes it. ## Risk User-visible change is one line of CLI output at exit, and it now states an unknown outcome rather than a good one. A user who reads it has nothing to redo -- re-sending the turn could duplicate a write the service did complete -- so the line is informational by design, not an alarm. The exit path is otherwise untouched: the budget constant is byte-identical to the one on the default branch, and the diff of that file contains no non-comment line, so exit timing does not move. `drain_backend_stores` changes its return type, which is internal. Five callers: three render the notice and are updated here, two await it and ignore the value as they already did. Rollback is a revert of the five commits on this branch. Nothing persists and no format changes. Reverting only the last two would restore the delivery claim while keeping the split, which is the state review rejected, so revert all or none. Checked and deliberately not changed: - An existing assertion changed meaning and is called out rather than left in the diff. The test pinning the wording "still finishing in the background", whose docstring stated the write had reached the service, is renamed and rewritten; its two negative assertions were right and are kept, and a new sibling pins the other direction. - The memory plugin still prints its own `final-flush sweep hit its 5.0s budget` line to stderr at shutdown. It belongs to a different layer and is factually correct; folding it into this notice would widen the diff into the plugin. - The same teardown in `gateway_commands` stays uncovered. Its enclosing function is 634 lines against 60 in `agent_commands`, and the gate passes without it. - No domain term is coined. `StorePipeline` appears in `CONTEXT.md` only as a passing mention with no entry of its own, and `DrainOutcome` is a module-local type rather than a concept the map tracks. ## Related Issues N/A --------- Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
## Summary The parameter sweep case shipped a two-panel matplotlib plot. The same run also produced a dashboard, and it carries what the plot could only assert: the sixteen measured cells as a grid with recall and latency in each, the best cell ringed at top_k 10 and chunk_size 1024, and a footnote naming the single warm-up sample behind the 256/top_k=3 outlier. This swaps the plot for the dashboard and rewrites the alt text to describe that grid rather than the two lines. **Language.** The dashboard was authored in Chinese. Both front pages share one image set, and the comparison board beside this cell is already English, as is the alt text on both pages, so what lands is an English render built the same way that board was. The Chinese original is left untouched on disk and nothing in the repository refers to it. **Widths.** The new render is 2000x1568 against the board's 2000x1547, so the two sit almost exactly on one aspect. The board goes back to full width and the dashboard takes 99 percent; the board only needed 86 percent against the old plot's 2000x1332. ## Type - [ ] Fix - [ ] Feature - [x] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification ``` git diff --check origin/main...HEAD clean COMMIT_RANGE=origin/main...HEAD make check-large-files exit 0 PYTHONPATH=. uv run --frozen --python 3.12 --extra dev python scripts/check_source_language.py origin/main...HEAD exit 0 make check-commits exit 0 PR_TITLE="docs: replace the sweep chart with the run's dashboard" make check-pr-title exit 0 git merge-tree --write-tree HEAD origin/main clean ``` The attachment was fetched anonymously before it was referenced: HTTP 200, 2000x1568, md5 fcd434b3000cbcda9339e98c22bf974b, matching the local render byte for byte. Rendered the section through GitHub's markdown endpoint under the published markdown stylesheet and measured every image box. All eight pairs share a top edge and match within 2px on height; the Frameworks pair is now 379x293 against 375x294, where before this change it was 326x252 against 375x294. Both front pages were compared after the edit: each carries the identical set of attachment ids, and neither still references the old plot. No test suite is relevant to a change that edits two markdown files. The python lint job was not run locally for the same reason; CI runs it on the head. - [ ] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed ## Risk Documentation only, in both READMEs. One attachment reference is replaced and one width attribute is restored; no file is committed to the repository. The old plot remains reachable at its own attachment URL, so a revert of the commit restores it with no other action. - [ ] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues Follows #555. Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
…575) ## Summary The ppt-engine plugin's outbound hook still wrote a PDF rendering next to every published deck and announced it on a second `MEDIA:` line. A delivery therefore carried two files where the user asked for one, and the transcript showed a PDF tile beside the deck. The earlier change that stopped the build stage from producing the preview (D64) did not reach this hook. The hook now publishes the deck and names it once. The web UI renders its own preview from the deck on demand, so nothing is lost; the two helpers that placed and named the copy are removed. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification - `uv run --frozen pytest tests/test_ppt_engine_plugin.py -q` -> 52 passed (after rebase onto main). Four tests that asserted the PDF copy and the second MEDIA line are flipped to assert their absence. - `ruff check` / `ruff format --check` on the two files -> clean. - Observed on a full gateway before the fix: a template-engine delivery arrived as `deck.pptx` plus `deck.pdf` and two transcript tiles. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed ## Risk - User-visible: a deck delivery has one file, not two. Anyone who relied on the PDF beside the deck now opens the preview in the web UI or converts the deck themselves. - Rollback: revert the squash commit. - [ ] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A
## Summary `raven web` cannot start at all on a Windows host whose reserved port ranges cover 8765..8784. The gateway asks `pick_port(8765)` for the control-plane port; `pick_port` probes twenty ports forward and raises `OSError` when every one of them is refused. That raise reached no handler, so the whole gateway died, the supervisor retried five times and gave up. All twenty can be refused at once on Windows. winnat reserves whole hundred-port blocks (`netsh interface ipv4 show excludedportrange protocol=tcp`), a host whose dynamic port range starts low collects those blocks in the 8000s, and a bind inside one fails with WinError 10013 while netstat shows the port unused. On the host this was found on, the reserved blocks were 8636-8935 and 9014-9580, which swallow the probe span whole. The fix takes an OS-assigned port when the span is exhausted, rather than dying on a port nobody asked for. Nothing needs this port to be predictable: local clients read it from the lock payload, and `ControlPlaneServer.start` already reads the bound port back off the socket so that port 0 works. A warning names what happened, so an operator who cares can still see that the historical port was unavailable. Scoped to the control plane. `pick_port` itself is unchanged, because its other two callers serve the page, where the port is what an open browser tab is already pointed at and moving it strands that tab. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Run on the Windows 11 host that exhibits the failure (reserved ranges 8636-8935 and 9014-9580, so 8765..8784 is entirely unbindable). Unit tests: ``` uv run pytest tests/test_cli_gateway_commands.py tests/test_rpc_control.py tests/test_cli_gateway_page.py -q 71 passed ``` Both new behaviour tests were mutation-checked rather than assumed to bite: - restoring the inline `pick_port(8765)` at the call site: `test_the_gateway_takes_its_control_port_from_the_fallback` fails; - removing the `except OSError` fallback from the helper: `test_the_control_plane_falls_back_to_an_os_assigned_port` fails. - changing `pick_port`'s exhaustion raise from `OSError` to `RuntimeError`: `test_pick_port_raises_the_class_the_fallback_catches` fails, and is the only case that does (1 failed, 49 passed). Added in review: the two cases above replace `pick_port` with a stub that raises `OSError` itself, so neither reaches the real function, and the cross-module class contract the fallback rests on went unpinned. End to end on the same host, with an isolated `RAVEN_HOME` and every channel disabled: ``` uv run raven gateway --config <isolated config> --port 18795 WARNING raven.cli.gateway_commands:_control_plane_port - control plane: no free port from 8765; taking an OS-assigned one INFO raven.rpc.control:start - control plane: listening on ws://127.0.0.1:9594/ws OK Control plane: ws://127.0.0.1:9594/ws OK Health: http://127.0.0.1:18795/health ``` The published endpoint carries the port it actually got (`control_port: 9594` in the lock payload), and both control-plane clients reach it there: ``` uv run raven gateway status -> gateway pid 712, generation 1, page mounted at http://127.0.0.1:18792 uv run raven gateway stop -> gateway stopping, both ports released ``` Before this change the same command on the same host ended in `OSError: no free port in 8765..8784`. Lint: ``` uv run --extra dev ruff check raven/cli/gateway_commands.py tests/test_cli_gateway_commands.py All checks passed! uv run --extra dev ruff format --check raven/cli/gateway_commands.py tests/test_cli_gateway_commands.py 2 files already formatted ``` Not claimed: `make ci` was not run in full. `tests/test_gateway_lock.py` has 4 failures on this host, all of them POSIX file-mode assertions; they reproduce identically on unmodified `origin/main` (verified by stashing this change), so they are pre-existing and Windows-only, not introduced here. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed ## Risk Behaviour is unchanged on any host with a free port in 8765..8784: the fallback is reached only where the gateway previously refused to start. Where it is reached, the control plane answers on a different port each boot, which is already the case for the token beside it and is why both are published together rather than assumed. The control plane is loopback-only and token-gated on its first frame, and neither property depends on the port number, so an OS-assigned port is no weaker than the historical one. Rollback is the single commit; nothing is persisted that outlives the process beyond the lock payload, which is rewritten on every boot. ## Related Issues N/A --------- Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
…571) ## Summary Two halves of one incident, on 2026-09-20. A Raven-Design run (an ACP sub-agent) stalled on its model stream; the parent conversation was told `turn_failed (code -32099)` and nothing else, diagnosed it as an npm crash and restarted the run. Behind that, the same run had written 58,675 `agent_thought_chunk` frames, one per reasoning token, averaging 5.7 characters. Every frame is a synchronous write plus flush on the client's pipe, so when the client (the gateway, frozen by a synchronous `find`, PR #566) stopped reading, the server blocked on that write too. **Thought chunks are coalesced.** `UpdateTranslator` (the ACP server's outbound sink, `raven/acp/updates.py`) now holds consecutive `thinking.delta` updates per session and sends them as one `agent_thought_chunk` when: - a 200ms window closes (a loop timer, `THOUGHT_COALESCE_WINDOW_S`), or - the held text reaches 512 characters (`THOUGHT_COALESCE_MAX_CHARS`), or - anything that has to follow them arrives: a chunk of the answer, a tool call, a latch or an ending event, the session's release, the turn slot being dropped, or the connection closing. Order on the wire is kept: held text always precedes whatever released it, and a delegated agent's thought (a different `_meta` target) never joins the main agent's held text. The thresholds are measured, not chosen: replaying the incident's frame log, 200ms/512 gives 4,333 frames (13.5x fewer); a 100ms window gives 6,210; a 1,024-character cap changes nothing (4,331), so the window is what does the work. The median gap between consecutive thought frames in that log was 0.34ms, the p90 92ms. `agent_message_chunk` (the answer itself) is not coalesced: a reader watches it stream, and its frame count was 79. **A failed turn names what failed.** (Second round: the class is prefixed only when the message does not already open with it, so an SDK error that names itself does not stutter; and `StreamIdleTimeoutError` is classified by type, as `stream_idle_timeout`, before any substring branch can read the budget in its message as a 429 or a 500.) `Lane._run_turn` built `TurnFailed(error=str(exc))`, and `str()` of asyncio's bare `TimeoutError` is empty, so the ACP translator's `_error` had only the wire message to show. `describe_failure(exc)` now carries the exception class and its message when it has one (`ValueError: boom`, `TimeoutError`). And the stall itself is named at the source: the per-chunk idle `wait_for` in `LiteLLMProvider.chat_stream` raises `StreamIdleTimeoutError(idle=...)`, a `TimeoutError` subclass (so every `except TimeoutError` arm and retry verdict keeps applying, the same argument `FirstByteTimeoutError` makes) whose message reads `the model stream sent nothing for 180s (bound stream_idle_timeout=180s)`. Not changed: the frame writer is still a blocking write on the loop. Fewer frames make the stall far less likely, they do not make it impossible; moving the write off the loop would change the ordering guarantee `UpdateTranslator`'s docstring states and is not attempted here. The other providers' idle watchdogs (the Responses and Anthropic SSE paths) still raise a bare `TimeoutError`; they now surface as `TimeoutError` through `describe_failure` rather than as nothing, and giving them the same message is a follow-up. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification At head b6097ea: ``` uv run pytest tests/test_acp_updates.py tests/test_spine_scheduler_lane.py tests/test_litellm_provider_timeout.py -q -n 0 # 233 passed uv run pytest <24 test files across acp, spine, providers, rpc spine, subagent acp> -q # 1048 passed (before the last test was added) uv run pytest -q --ignore=tests/integration # 23791 passed, 110 skipped at head b6097ea, in a throwaway worktree with COLORTERM=truecolor make lint-python lint-imports lint-types lint-deps # all pass make check-source-language check-large-files # exit 0 (COMMIT_RANGE=origin/main..HEAD) ``` Revert-to-red, each production file (pair) reverted to `origin/main` on its own with the tests kept: ``` raven/acp/updates.py reverted 6 failed raven/spine/scheduler.py reverted 2 failed raven/providers/{first_byte,litellm_provider}.py reverted 1 failed ``` Mutation, each mutant applied alone, control restored between runs: ``` no flush before an update that must follow the thoughts 3 failed window timer never armed 2 failed 512-character cap disabled 1 failed different speakers merged into one chunk 1 failed release_session keeps the held text 1 failed close keeps the held text 1 failed end_turn keeps the held text 1 failed (0 before the second commit added its test) describe_failure drops the class name 2 failed idle stall re-raised as a bare TimeoutError 1 failed control (nothing mutated) 230 passed ``` Second round (267 tests across `test_acp_updates.py`, `test_spine_scheduler_lane.py`, `test_error_classification.py`): ``` the three files reverted to the first-round head 4 failed idle stall classified by substring 3 failed describe_failure always prefixes 1 failed settle_turn keeps the held text 1 failed sliding window (timer re-armed per chunk) 1 failed control 267 passed ``` The two guard clauses the first round's mutation pass showed could not decide the flush (`or result.latch`, `or _ends_the_stream(event)`) were dropped rather than tested; the comment says why. One existing test changed: `test_reasoning_arrives_as_a_thought_chunk` asserted a lone thought went out immediately; it now sends two thoughts and an answer chunk and asserts one thought frame precedes the answer, which is the ordering rule this PR adds. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [x] User-facing docs or screenshots are updated when needed (CHANGELOG `Unreleased / Fixed`, two entries) ## Risk User-visible behaviour changes: - An ACP client sees reasoning text arrive in bursts of up to 200ms or 512 characters instead of per token, so the first thought token of a turn arrives up to 200ms later than before. Answer text, tool calls and endings are unchanged and never reordered relative to the thoughts that preceded them. - Every failure event (`error` on the wire, `turn_failed` in ACP) now starts with the exception class, e.g. `ValueError: boom` where it read `boom`. A client matching the old text exactly would need to match the suffix. - The mid-stream idle stall in the LiteLLM path raises `StreamIdleTimeoutError` instead of a bare `TimeoutError`. It is a `TimeoutError`, so `except TimeoutError` behaves as before; `classify_error` names it `stream_idle_timeout` (retryable, fallback) where it used to say `network`, and `_NON_AUTH_HINTS` has no entry for it, so the CLI renders its error line alone, as it does for `first_byte_timeout`. Rollback: revert the squash commit. No stored data or wire schema changes; the coalescer is process-local state. - [x] Security impact considered (held text is the model's own reasoning, already on this channel; `redact` is applied where it was before, in the failure text) - [x] Backward compatibility considered (frame shapes unchanged and schema-validated in the tests; exception hierarchy preserved) - [x] Rollback path is clear for risky changes ## Related Issues N/A --------- Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
## Summary The agents page offered a Connect that could not be trusted, in two separate ways, and this closes both. A press held for as long as the server's readiness gate takes -- one real prompt through that agent's own backend, measured at 16s against a live adapter -- and nothing on the row said it had begun. A reader with nothing to look at pressed again. The second press then killed the first: the connection pool keys a connection on its launch arguments, cwd among them, and cwd falls back to the calling task's workspace, so a ping's key never matches a held connection and acquire retires it. The first press came back "sub-agent did not answer a test message", naming the agent for a failure raven had caused. Reproduced on a live gateway: press one failed at 3.0s, exactly the click interval; press two succeeded at 17.1s. Two changes, either of which alone leaves a hole. The page holds a row while its own write is in flight, the way the Test button beside it always has -- the writes only, since test_cancel has to reach the server while its test is still running. And a readiness ping now runs on a connection pool of its own, closed on the way out, which also stops a ping fired while the agent is serving a real run from taking that run's connection with it. The same A/B afterwards: both presses succeeded, 15.0s and 13.2s, with no relaunch between them. The second way was the button itself. Four of the eight third-party agents this machine can offer would refuse the press and always would have: three have no binary on PATH, and one answers the handshake and then declines to open a session without a credential. Each carried a Connect identical to a working row's. The page now says which it is and offers nothing -- "Not Installed" for a command that does not resolve, "Unauthorized" for an agent that asked to be signed in. Neither is pressable, because neither remedy is on this page; the way back is the card's own Test, which re-measures and replaces the recorded verdict, plus a restart, which re-measures a recorded refusal rather than trusting it. A connected row is untouched, so an expired credential cannot cost it the Disconnect it still needs. needs_auth is measured, not inferred, and it takes the refusal's own words. verify_agent already read a verdict off the handshake's auth methods and the error text, spent it on choosing a status, and dropped it; the advertisement half is not evidence about any one refusal, because an agent that works advertises auth methods too -- one measured here names four -- so needs_auth takes only the credential-shaped refusal while the coarse status keeps the reading it always had. The snapshot then carries the verdict and nothing downstream pattern-matches an agent's prose. Two supporting changes follow: the snapshot backfill now records a credential refusal as well as a pass, and covers shipped presets nobody has configured, which is exactly the set whose Connect was about to mislead. Costs and residue, all deliberate: - The backfill now handshakes the shipped acp presets that are on this machine, not only the configured rows: 9 instead of 4 here. Sequential, backgrounded, no model tokens. A preset whose command does not resolve is skipped, so nothing is launched to learn what the free probe already answered. - Recording a non-ready snapshot is narrowed to credential refusals. A timeout or a crashed adapter stays unrecorded, because that is a fact about one minute and freezing it would label a working agent broken until somebody pressed Test. - The snapshot store gains at most one row per shipped preset. - The TUI roster reads the same rows and calls the same subagents.add / subagents.toggle. It does not read needs_auth, so its button is unchanged -- not a regression, and not fixed here because it would widen the diff into another surface. Named so a reviewer does not have to find it. - Recording a preset's snapshot also means a preset row's `stateful` and its probe status are read from a measurement rather than defaulted. More accurate in both cases, and named here because it is a second consequence of one line. - A Test records a snapshot and does not rebuild the agent table. That is unchanged and pre-existing for configured rows; a preset row has no roster entry to rebuild. - The pre-submit sweep found three defects in this branch's own code and each is fixed in its own commit: the page's leading comment still declared "two verbs on a row, and only two" after a third state was added; the preset filter resolved commands against this process's PATH while _probe_acp answers the row on screen from the login shell's; and the scheduler's early return had to go, because whether any shipped preset is on this machine cannot be answered from the synchronous side. ## Type - [x] Fix - [ ] Feature - [ ] Docs - [ ] CI / tooling - [ ] Refactor - [ ] Other ## Verification Run in a worktree at this branch's head, rebased onto 8d74fb9. - `uv run --frozen --all-extras pytest tests/test_subagent_probe.py tests/test_subagent_acp.py tests/test_rpc_subagents.py tests/test_rpc_schema_match.py tests/test_subagent_acp_mcp_lifecycle.py tests/test_acp_unprompted.py tests/test_subagent_manager.py tests/test_acp_backend_cwd.py tests/test_subagent_third_party.py -q -p no:randomly` -- 1157 passed. - `npx vitest run src/features/xa/` in ui-web -- 91 passed. - Full ui-web suite, serially, with the localstorage flag this box needs -- 1840 tests, 2 failed, 0 unhandled errors. Both failures are in src/features/workspace/, which this branch does not touch, and the failing ID set is identical to the one recorded before any of this work. They are an artifact of node v26 shadowing happy-dom's localStorage on this machine. - `make lint-python lint-imports lint-types lint-ui lint-tui` -- all pass. lint-bridge fails on a TypeScript deprecation in bridge/tsconfig.json; it fails identically on an untouched checkout and this branch does not touch bridge/. - `make check-commits check-large-files check-source-language` with the three-dot range -- all exit 0. - Seventeen new tests, fourteen python and three in the page's own suite. Every one was watched failing for the missing behaviour before the code that answers it was written, except the three that cover already-written defensive branches, which are verified by the coverage measurement below instead. The test_cancel regression that guards the in-flight lock was proved load-bearing by widening the lock to the whole row and watching it go red. - Review round one, two blocking findings, both reproduced here before being fixed and both re-run afterwards. needs_auth from the advertisement alone: forcing the message heuristic to False left needs_auth true against the stub, and the stub gained a `no_session_other` mode because every existing refusal mode refuses in auth-shaped words, so no test could separate the two halves. The unrecoverable row: a recorded refusal, then `run_test(source="preset")` returning ok, then the same snapshot still reading needs_auth true; it now reads false and ready. Re-measured against the real agents afterwards, Grok Build still reports needs_auth true and CodeBuddy still ready. - The diff-coverage gate read 88.00% on the first head. Of the 120 lines this branch adds to probe.py, none is now uncovered; the four that remain in that file are the base's own. - A second sweep over the review-fix commits found one more defect, fixed in its own commit: a test comment still gave the removed gate as the reason for an assertion that stands for a different one. - End to end against real agents on this machine: the backfill reaches the 9 presets that are installed and skips Copilot, Qwen and Kimi, which are not; verify reports CodeBuddy ready with needs_auth false and Grok Build attention with needs_auth true; and the rows come back reading ready/false, missing/false and attention/true respectively. - [x] Relevant tests pass locally - [x] Relevant lint / type checks pass locally - [ ] User-facing docs or screenshots are updated when needed No docs change: the agents page carries no user-facing documentation page, and the two new strings are catalogue entries. ## Risk User-visible behaviour changes, all on the agents page and its server surface: - A Connect press disables its own button until the call returns. A second press for the same row is dropped client-side rather than sent. - A row whose command is not on the machine, and a row whose agent asked to be signed in, no longer offer Connect. They show a word and are not pressable. The way back for the second is the card's Test. - A gateway start now handshakes the installed shipped presets once, in the background. - subagents.list rows gain needs_auth. It is optional on the wire and defaults false, so a client that does not read it behaves exactly as before. Rollback is the revert of this squash commit: nothing here migrates data, and a snapshot row recorded by the widened backfill is re-derived or ignored by the previous code, which reads the same store and simply never wrote those rows. - [x] Security impact considered - [x] Backward compatibility considered - [x] Rollback path is clear for risky changes ## Related Issues N/A --------- Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
Resolves 64 conflicts against main at cf3c95f. Twenty-five paths were taken from main unchanged: this branch never edited them, it only carried stale copies picked up by the hand-squashed catch-up commit 80ecae6, and main has moved on since. Nineteen delete/modify conflicts stay deleted, the old frontend tree having been replaced. Three resolutions made a judgement call, and are the ones to re-read: - subagent/probe.py: main's #554 records a preset's capability snapshot where this branch's #559 refused to. Main's decision stands, through this branch's machinery, so record_capabilities remains the one writer and _test_record still keeps a previous measurement when a verify fails. The boot backfill re-measures on either side's condition. - rpc/methods/subagents.py: both sides read the same snapshot under different names. One read now serves meta, own and needs_auth. - ui-web check-class-namespace.mjs: the rail ratchet moves 11 to 12 for .wdt, which arrives with main's folder chip beside eleven other unprefixed classes of that domain. Two gaps this merge does not close. The rebuilt settings page offers no timezone control, so the backend half of #506 has no surface; and main's shell/workdir.ts folder picker is left out, since it imports shell modules this branch deleted. The rail half of #542 did merge into features/rail. lib/upload.ts follows raven/rpc/files.py to 100 MB, and the offline fixtures now send session workdir and subagent needs_auth. Verified on this head: ui-web type-check and gen:check clean (190 methods), 40 of 40 convention gates, 1279 backend tests, source-language check clean. The full ui-web suite fails the same files it fails on the unmerged head of this branch, and no others. Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
Both files took their conflict as a union of the two sides' appended tests, which left the joins one blank line short of what ruff format wants. Four blank lines, no test touched. Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
gloryfromca
left a comment
There was a problem hiding this comment.
Blocking: the rebuilt external-agent UI does not preserve main's authentication-refusal behavior.
Two merge-specific defects are marked inline. Together they undo the user-visible contract of #554: a measured credential refusal must reach the page as a disabled Unauthorized state, including when an agent that worked before later logs out.
I covered the full target-to-head file set and main arrival history, the remerge/combined conflict resolutions, the affected callers and persisted snapshot path, backward compatibility and generated RPC/fixture shapes, test changes for weakening, AGENTS.md and the ui-web CONTEXT/CONTRIBUTING architecture rules. I found no weakened tests or additional architecture violations worth raising.
Verification on this head: the focused backend set passed 593 tests with one environment skip (python-pptx unavailable); ui-web type-check and gen:check passed (193 methods); eslint exited 0 with five existing exhaustive-deps warnings; and all 192 ui-web test files passed (2,573 tests). A direct _test_record reproduction also showed a new needs_auth=True refusal persisted as false when a prior successful snapshot exists.
The author-designated follow-ups remain nonblocking from my side: the rebuilt settings page still lacks the timezone control, and the old workdir folder picker was not ported.
gloryfromca
left a comment
There was a problem hiding this comment.
Blocking: the two open authentication-refusal findings still have to be fixed.
The only delta from the reviewed head is four formatter-required blank lines in tests/test_rpc_console.py and tests/test_rpc_subagents.py. Neither affected implementation path changed, so both existing threads remain applicable and unresolved. I rechecked the resulting target-to-head diff, its callers and history, compatibility and generated-contract surfaces, the prior test coverage, AGENTS.md, and the ui-web architecture rules; I found no new findings and no weakened tests.
Verification on 4113b729c184: uv run pytest tests/test_rpc_console.py tests/test_rpc_subagents.py -q passed 228 tests; uv run ruff format --check reports both files formatted; uv run ruff check passed; git diff --check passed.
…sured `_test_record` keeps a previous record's capabilities when a verify fails, so one flaky Test does not cost a row its menu and statefulness. It rebuilt that record from `previous` and copied only the verdict fields, and `needs_auth` was not among them -- it arrived on main with #554, after this helper was written. An agent that worked and has since been logged out therefore stayed `needs_auth: false` through the very Test that detected the refusal, and the page went on offering it a Connect the agent would decline. `needs_auth` travels with the verdict, not with the capabilities: a menu measured earlier is still the best account of what the agent can do, but whether it will open a session without a credential is a fact about this handshake alone. Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
…hake The server reports `needs_auth` on an acp row whose handshake refused to open a session without a credential. The rebuilt external-agent domain never read it: `extAgentRowOf` is a whitelist and did not name the field, so `stageOf` saw an unconfigured row and answered `add`, and the row rendered Connect and sent `subagents.add` to an agent that had already said no. Carries main's #554 contract into the domain that replaced its page: the field on the row, the two stages that name why a write would fail rather than a write to make, and a disabled label in place of the button, in the row and in the sheet, which read one decision so they cannot disagree. The store refuses those stages too. It maps anything that is not `add` or `stale` onto `toggle`, and it is reached from the row, the sheet and the wizard step -- a barred row arriving there would have turned a refused add into a toggle of an entry that does not exist. Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
gloryfromca
left a comment
There was a problem hiding this comment.
No blockers; this can merge as far as I am concerned.
The new delta fixes both prior blockers: the rebuilt web domain now carries the handshake authentication verdict through its source mapper and stage/UI/store decisions, and preserved capability records take needs_auth from the newest measurement while retaining only the older capability data. I replied to and resolved both settled threads.
Review coverage: the repository rules and UI architecture constraints; the delta against the previously reviewed target diff; affected callers and wire compatibility (the optional field remains safe with older servers); relevant history and intended #554 behavior; and the additive regression tests to confirm no tests were weakened.
Verification:
uv run pytest tests/test_subagent_acp.py tests/test_subagent_probe.py tests/test_rpc_subagents.py -q: 384 passednpx vitest run src/features/extAgents/source.test.ts src/features/extAgents/store.test.ts src/features/extAgents/ExtAgentsPage.test.tsx: 81 passednpm run type-check: passednpm run lint: passed with 5 existing warnings in untouched files and 0 errorsuv run ruff check raven/agent/subagent/probe.py tests/test_subagent_acp.py: passed
Summary
Merges current
main(cf3c95fd7) into the ui-web rebuild branch and resolves the64 conflicts it raises. Base is
refactor/ui_web_architecture, notmain.PR #475 has reported
CONFLICTING/DIRTYfor six review rounds. The conflicts grewfrom two files to sixty-four, and five of them had reached
raven/agent/, so theoverlap was no longer confined to the frontend.
Most of it is not a merge at all. Twenty-five paths were taken from
mainunchanged:this branch never edited them, it only carries stale copies picked up by the
hand-squashed catch-up commit
80ecae66b, andmainhas moved on since. Each one wasverified before it was taken, by checking that the branch's copy is byte-identical to a
mainrevision (harness/participants.pyandhook/participant.py, for instance, bothmatch
mainatdbe259f6d). Nineteen delete/modify conflicts stay deleted, the oldfrontend tree having been replaced. Two generated files were regenerated from the merged
sources rather than merged by hand.
Three resolutions made a judgement call, and are the ones worth re-reading:
raven/agent/subagent/probe.pyis a design conflict, not a text one. Main's fix(*): stop the connect button promising what it cannot do #554records a preset's capability snapshot; this branch's feat(*): take the agents page to the agent hub prototype #559 refused to, on the older
rule that a preset is a template. Main's comment answers that rule directly: the page
draws a preset row's verdict from the snapshot, so Test is that row's only way back
and only is one if it writes. Main's decision stands here, through this branch's
machinery:
record_capabilitiesremains the one writer and_test_recordstill keepsa previous measurement when a verify fails. The boot backfill now re-measures on
either side's condition (an unmeasured model menu, or a credential refusal, which
never goes stale on its own).
raven/rpc/methods/subagents.py: both sides read the same snapshot under differentnames. One read now serves
meta,ownandneeds_auth.ui-web/scripts/check-class-namespace.mjs: the rail ratchet moves 11 to 12 for.wdt, which arrives with main's folder chip beside eleven other unprefixed classesof that domain. Prefixing it properly would touch main's component and its test, which
is the author's call rather than this merge's.
Two gaps this merge does not close, both worth a follow-up:
key for one, so the backend half of fix(*): let the web settings timezone control write, and drop the retired one #506 lands here with no surface.
ui-web/src/shell/workdir.tsfolder picker is left out: it importsshell/modules this branch deleted, so taking it verbatim would not compile. The rail half of
feat(*): a conversation starts in a folder of your choosing, and the rail groups by it #542 did merge into
features/railand its tests pass.Two follow-through fixes the merge needed:
ui-web/src/lib/upload.tsgoes from 25 MB to100 MB, following
raven/rpc/files.py, which its own comment names as the source andwhich main raised in #549; and the offline fixtures now send
session.listworkdirand subagent
needs_auth, which the fixture-shape gate requires and which was alreadyfailing on this branch's own head.
Type
Verification
Run on the merge head, with
node_modulescopied into tmpfs because the sharedfilesystem hit 100 percent twice during the work and a write into it fails the run:
npm run type-check(ui-web): clean.npx tsc --noEmit -p tsconfig.json(ui-tui): exit 0.npm run gen:check:generated.ts matches the contract (193 methods).npx vitest run scripts/gates/(ui-web): 40 of 40 files passed.uv-freepytestover the touched backend areas: 1279 passed acrosstest_subagent_probe,test_subagent_acp,test_rpc_subagents,test_rpc_console,test_subagent_manager,test_subagent_dag_runner,test_subagent_activity,test_rpc_tasks,test_rpc_settings,test_rpc_session,test_config_live.scripts/check_source_language.py origin/main...HEAD: exit 0.The full ui-web suite was run on both this head and the unmerged head of this branch, in
two clean clones, and compared file by file: 31 files fail here against 32 there, no file
fails here that did not already fail there, and
scripts/gates/fixture-shape.test.mjsgoes from failing to passing. Those failures are a local Node 26 environment problem; CI
runs Node 22.
Three checks that were not left to inspection: the i18n keys dropped with main's block
were each grepped across the surviving tree before dropping them (zero still referenced,
and the seven referenced-but-absent keys that remain are identical to the set on this
branch's own head); the unioned test files were rebuilt at AST granularity after a naive
text union interleaved two test bodies; and the replay of these resolutions onto the
author's newer heads was proved faithful by showing the result differs from the verified
tree by exactly the files in the author's new commits, no more and no less.
No user-facing docs change: this merge adds no feature. The two gaps named in the
Summary are the only user-visible deltas against main, and both are losses of a control
rather than additions.
Risk
Behaviour that changes for a reader of this branch:
Unauthorizedcan be cured from the page. This is main's behaviour arriving, not anew one.
as it was, the page would have refused at 25 MB what the backend accepts.
Rollback: this is a single merge commit with two parents. Reverting it on the target
branch restores the pre-merge tree exactly; nothing here is squashed into another change.
Related Issues
#475