Skip to content

chore(*): merge main into the ui-web rebuild and resolve its conflicts - #590

Closed
LivXue wants to merge 69 commits into
refactor/ui_web_architecturefrom
chore/merge_main_into_ui_web_rebuild
Closed

LivXue wants to merge 69 commits into
refactor/ui_web_architecturefrom
chore/merge_main_into_ui_web_rebuild

Conversation

@LivXue

@LivXue LivXue commented Sep 21, 2026

Copy link
Copy Markdown
Member

Summary

Merges current main (cf3c95fd7) into the ui-web rebuild branch and resolves the
64 conflicts it raises. Base is refactor/ui_web_architecture, not main.

PR #475 has reported CONFLICTING / DIRTY for six review rounds. The conflicts grew
from two files to sixty-four, and five of them had reached raven/agent/, so the
overlap was no longer confined to the frontend.

Most of it is not a merge at all. Twenty-five paths were taken from main unchanged:
this branch never edited them, it only carries stale copies picked up by the
hand-squashed catch-up commit 80ecae66b, and main has moved on since. Each one was
verified before it was taken, by checking that the branch's copy is byte-identical to a
main revision (harness/participants.py and hook/participant.py, for instance, both
match main at dbe259f6d). Nineteen delete/modify conflicts stay deleted, the old
frontend tree having been replaced. Two generated files were regenerated from the merged
sources rather than merged by hand.

Three resolutions made a judgement call, and are the ones worth re-reading:

  • raven/agent/subagent/probe.py is a design conflict, not a text one. Main's fix(*): stop the connect button promising what it cannot do #554
    records a preset's capability snapshot; this branch's feat(*): take the agents page to the agent hub prototype #559 refused to, on the older
    rule that a preset is a template. Main's comment answers that rule directly: the page
    draws a preset row's verdict from the snapshot, so Test is that row's only way back
    and only is one if it writes. Main's decision stands here, through this branch's
    machinery: record_capabilities remains the one writer and _test_record still keeps
    a previous measurement when a verify fails. The boot backfill now re-measures on
    either side's condition (an unmeasured model menu, or a credential refusal, which
    never goes stale on its own).
  • raven/rpc/methods/subagents.py: both sides read the same snapshot under different
    names. One read now serves meta, own and needs_auth.
  • ui-web/scripts/check-class-namespace.mjs: the rail ratchet moves 11 to 12 for
    .wdt, which arrives with main's folder chip beside eleven other unprefixed classes
    of that domain. Prefixing it properly would touch main's component and its test, which
    is the author's call rather than this merge's.

Two gaps this merge does not close, both worth a follow-up:

Two follow-through fixes the merge needed: ui-web/src/lib/upload.ts goes from 25 MB to
100 MB, following raven/rpc/files.py, which its own comment names as the source and
which main raised in #549; and the offline fixtures now send session.list workdir
and subagent needs_auth, which the fixture-shape gate requires and which was already
failing on this branch's own head.

Type

  • Fix
  • Feature
  • Docs
  • CI / tooling
  • Refactor
  • Other

Verification

Run on the merge head, with node_modules copied into tmpfs because the shared
filesystem hit 100 percent twice during the work and a write into it fails the run:

  • npm run type-check (ui-web): clean.
  • npx tsc --noEmit -p tsconfig.json (ui-tui): exit 0.
  • npm run gen:check: generated.ts matches the contract (193 methods).
  • npx vitest run scripts/gates/ (ui-web): 40 of 40 files passed.
  • uv-free pytest over the touched backend areas: 1279 passed across
    test_subagent_probe, test_subagent_acp, test_rpc_subagents, test_rpc_console,
    test_subagent_manager, test_subagent_dag_runner, test_subagent_activity,
    test_rpc_tasks, test_rpc_settings, test_rpc_session, test_config_live.
  • scripts/check_source_language.py origin/main...HEAD: exit 0.
  • Zero unresolved paths, zero conflict markers in the tree.

The full ui-web suite was run on both this head and the unmerged head of this branch, in
two clean clones, and compared file by file: 31 files fail here against 32 there, no file
fails here that did not already fail there, and scripts/gates/fixture-shape.test.mjs
goes from failing to passing. Those failures are a local Node 26 environment problem; CI
runs Node 22.

Three checks that were not left to inspection: the i18n keys dropped with main's block
were each grepped across the surviving tree before dropping them (zero still referenced,
and the seven referenced-but-absent keys that remain are identical to the set on this
branch's own head); the unioned test files were rebuilt at AST granularity after a naive
text union interleaved two test bodies; and the replay of these resolutions onto the
author's newer heads was proved faithful by showing the result differs from the verified
tree by exactly the files in the author's new commits, no more and no less.

  • Relevant tests pass locally
  • Relevant lint / type checks pass locally
  • User-facing docs or screenshots are updated when needed

No user-facing docs change: this merge adds no feature. The two gaps named in the
Summary are the only user-visible deltas against main, and both are losses of a control
rather than additions.

Risk

Behaviour that changes for a reader of this branch:

  • An acp preset row now gets its capability snapshot recorded by Test, so a row showing
    Unauthorized can be cured from the page. This is main's behaviour arriving, not a
    new one.
  • The attachment ceiling the page enforces rises to 100 MB, matching the gateway. Left
    as it was, the page would have refused at 25 MB what the backend accepts.
  • Everything else is main's own behaviour, already shipped there.

Rollback: this is a single merge commit with two parents. Reverting it on the target
branch restores the pre-merge tree exactly; nothing here is squashed into another change.

  • Security impact considered
  • Backward compatibility considered
  • Rollback path is clear for risky changes

Related Issues

#475

xfng-sd and others added 30 commits September 17, 2026 11:56
…ads (#463)

## Summary

`main` is red on three jobs (`python lint`, `unit 2/4`, `unit 3/4`), and
has been
since #451 merged. All of it is fallout from one arrival: `image_search`
came with a
second switch beside its credential.

`tools.web.search.images` decides whether the tool is registered at all
(`raven/agent/loop/wiring.py`), so a lane that never places a picture
keeps the tool
face it had. The capability table was not told. `is_configured` answers
on the vendor
key alone, so with a Serper key and the switch left off the table ticked
a capability
the agent does not hold, and the two tests that drive a real `AgentLoop`
caught exactly
the disagreement that module exists to prevent.

The switch belongs on `is_disabled`, not `is_configured`. That
function's own docstring
already names the reason: a switched-off tool usually has its credential
set, and
calling it unconfigured sends the deployer to set a key that is already
there.
`tools.disabledTools` and this flag are the same kind of answer, so
`is_offered` needs
no change and the doctor row reads "off" rather than "missing a key".

The test configs that mean "everything supplied" now turn the switch on,
the way
`raven/core/runtime.py` passes it from the config, and a new case pins
the combination
that was wrong: key present, pictures off, nothing offered, table saying
off.

Two more pieces of the same arrival:

- `ImageSearchTool` copied `WebSearchTool`'s key resolution without its
two attribute
annotations, so the untyped lambda `wiring.py` hands it left the
`api_key` property
returning `~AlwaysFalsy | str` against a declared `str`. The type
checker pinned in
  #416 reports it; the annotations are the fix.
- `WebSearchConfig` gained the `images` field, so the writer's validated
subtree now
carries it and two `test_config_update_tools` expectations that pin the
written
  section by value went stale.

Not changed here: `get_web_search` still reports `{provider, api_key,
max_results}` and
does not surface `images`, though `set_web_search` can write it. That is
a gap in the
read side of the config API, not a CI failure, and it wants its own
change.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Run on Windows 11 in a clean worktree cut from `main` (`5a29ea22`):

- `uv run pytest tests/test_tool_capabilities.py
tests/test_config_update_tools.py -q`:
71 passed. All four tests failing on `main` now pass, plus the new case.
- `uv run --frozen --python 3.12 --all-extras ty check raven evolver
agents plugins-dist scripts`:
the `web.py:516` diagnostic is gone. The diagnostics that remain locally
are
unresolved imports for optional packages absent from this machine's
environment
(`boxlite`, `nio`); CI reported exactly one diagnostic before this
change, the one
  fixed here.
- `uv run --frozen --python 3.12 --extra dev ruff check raven tests`:
all checks passed.
- `ruff format --check` on the four changed files: already formatted.
- `uv run pytest tests/test_cli_doctor_commands.py
tests/test_rpc_console.py -q`, which
consume `is_disabled`: the failing-test set is byte-identical before and
after this
change (8 pre-existing Windows-only failures around file modes and an
absent browser
  package, none of them in the set CI reports).

## Risk

`is_disabled` gains a second reason to answer true, so a deployment with
a search key
and `tools.web.search.images` off now reads as "switched off" in `raven
doctor` and the
console capability report, where it previously read as offered. That is
the state the
loop was already in; only the description changes. Nothing about
registration, gating,
or the tool's runtime behaviour moves.

The other two changes are an annotation and two test expectations, with
no runtime
effect.

Rollback is reverting this commit; it touches nothing that other work
builds on.

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
… busy runner (#466)

## Summary

Since the idle ceiling became a gate (#428), a unit shard has failed on
`main` and on most pull requests with every test passing: the ceiling
reads wall clock minus CPU and calls the difference waiting, and a
thread that is runnable and not on a CPU spends it the same way. With
four workers, the controller and coverage on a four-vCPU runner that
queueing reached three to four seconds inside one long test, so the
shard died on whichever test was executing when contention peaked, a
different one each run. Warming the catalogue (#452) did not stop it.
#449 halves the workers per shard, which pays for the same measurement
with every shard's wall clock.

This changes the measurement instead. Idle now subtracts the time the
test's threads spent queued for a CPU, read from the scheduler's own
account in `/proc/self/task/<tid>/schedstat`:

- every thread of the process is read, not only the one running the
test, because the rpc model handlers do their catalogue reads on
`asyncio.to_thread`;
- a thread writes its own last reading on the way out, through a wrapper
on `Thread._bootstrap_inner`, because the per-test event loop is closed
at the end of the test and takes its executor threads with it, and
`/proc` forgets a thread the moment it ends;
- CPU stays on `os.times`, which keeps the CPU of ended threads and of
reaped children; the scheduler's per-thread run time would have been
lost the same way the queue time was.

A sleeping thread is not on a run queue, so a test that waits on purpose
is still charged the whole wait. Without `schedstat` (macOS, Windows)
the arithmetic is the old one; the strict gate only runs on Linux, and a
Linux-only test fails loudly if the runner's kernel stops keeping the
statistics rather than letting the gate degrade quietly. `-n 4`, the 3 s
ceiling and the four shards are unchanged. If this lands, #449 is not
needed.

## Type

- [ ] Fix
- [ ] Feature
- [ ] Docs
- [x] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Unit tests for the clocks (12 on macOS, 13 on Linux, the last one reads
the live kernel):

```
uv run --frozen --all-extras pytest -q tests/test_conftest_idle_ceiling.py
  macOS: 12 passed, 1 skipped
  Linux container: 13 passed
```

Full suite on macOS, where the fallback arithmetic runs:

```
uv run --frozen --all-extras pytest -q -p no:cacheprovider
  23137 passed, 109 skipped in 291.30s
```

Contrast runs in a Linux container (Docker, linux/arm64, kernel 6.12),
the checkout and venv on container-local storage, `main` at 2bdc4bd
against this branch. Four workers pinned to two CPUs stands in for the
four-vCPU runner; the ceiling is lowered to 0.05 s so the effect shows
on one file:

```
docker run --cpuset-cpus=0-1 ... pytest tests/test_rpc_model.py --idle-ceiling 0.05 -n 4
  main:   70 unmarked tests over the ceiling, worst 1.67s idle of 2.89s
  branch:  2 unmarked tests over the ceiling, worst 0.12s idle of 0.50s (0.19s queued)

docker run --cpuset-cpus=0-3 ... pytest tests/test_rpc_model.py --idle-ceiling 0.05 -n 16
  main:   81 over, worst 6.11s idle of 8.16s
  branch:  1 over, 0.13s idle of 1.11s (0.73s queued)

an unmarked test that does time.sleep(0.5), on the branch:
  0.53s idle of 1.70s (0.00s queued)   # an honest wait is still charged in full
```

The CI shard command, same container, four workers pinned to two CPUs:

```
pytest -q --shard 3/4 --durations=25 --idle-ceiling-strict -n 4 --cov=raven --cov=raven_everos --cov-branch --cov-report=
  main:   5131 passed, 10 skipped; "1 unmarked test(s) waited more than 3s (failing the session)"
          3.11s idle of 6.04s tests/test_ppt_engine_template_tool.py::test_binding_a_template_by_path; exit 1
  branch: 5378 passed, 11 skipped; no idle report; exit 0
  branch, shards 1/4, 2/4, 4/4: no idle report in any (three failures, all `git ls-files` exit 128 because
          the copied worktree's gitdir is not inside the container)
```

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed (none; the
option help text and the summary line say what is now subtracted)

## Risk

- Behaviour change is confined to what the ceiling charges a test: less
on a busy runner, the same for a real wait. Nothing under `raven/`
changes.
- `Thread._bootstrap_inner` is private CPython API (present from 3.8
through 3.13); the wrapper is installed once, guarded against a second
import of the module, and pinned by a test. If a future interpreter
drops it the module fails to import at collection, which is loud.
- Rollback: revert the commit; the gate goes back to wall clock minus
CPU.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

#428 made the ceiling a gate; #452 warmed the catalogue; #449 halves the
workers instead.

Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
…458)

## Summary

`os.kill(pid, 0)` is the POSIX existence probe. On Windows, CPython
implements
`os.kill` as `TerminateProcess`, so signal 0 is not a probe there at
all, and it fails
in two directions depending on what the caller may do to the target:

- a process the caller can terminate, which is any child of it, is
killed and the call
returns cleanly, so the probe reports "alive" about a process it has
just destroyed;
- a process it cannot, such as the detached `raven web` supervisor,
raises `OSError`
with `ERROR_INVALID_PARAMETER`, so the probe reports "dead" about a
process that is
  running.

The second is why `raven web --stop` did nothing on Windows.
`_read_web_state` and
`_read_serve_pid` both answered `None`, so `_stop_resident` signalled
neither the
supervisor nor the gateway, and printed "nothing to stop" while the
engine went on
serving the page.

`raven/utils/pid.py` holds the platform question now. Windows asks
`OpenProcess` plus
a zero-timeout `WaitForSingleObject`, the pair
`raven/updates/upgrade.py` already uses
to watch a parent exit: a handle that fails to open for any reason but
`ERROR_ACCESS_DENIED` is a dead process, and one that opens and then
times out is a
live one. The wait rather than `GetExitCodeProcess`, because a process
whose real exit
code is 259 cannot be told apart from `STILL_ACTIVE`.

`_pid_alive` keeps its name in `serve_commands`, since `_gateway_page`
imports it and
the stop tests patch it; only the answer moves.

Four other call sites carry the same probe and are left for their own
change. The
sharpest is `raven/agent/tools/background_exec.py`, which reaches it on
the Windows
branch specifically, against the agent's own background tasks.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Run on Windows 11, in a clean worktree cut from `main`:

- `uv run pytest tests/test_utils_pid.py
tests/test_cli_serve_commands.py -q`:
  77 passed, 3 failed.
- the same 3 failures reproduce on `main` without this change
(`test_the_state_file_stays_owner_only`,
`test_sigterm_removes_the_state_file`,
`test_sigterm_during_the_prologue_still_removes_the_state_file`): they
are
pre-existing Windows gaps in the supervisor tests, unrelated to this PR,
and this
branch adds no new failure. Baseline on `main` is 73 passed, 3 failed;
the 4 extra
  passes are the new `tests/test_utils_pid.py`.

The new test drives real subprocesses. Survival is asserted by a wait
that must time
out, not by a poll taken straight after the probe: termination is
asynchronous, and
the immediate read still says "running" for a process already on its way
down. Forcing
the POSIX branch on every platform fails two of the four cases.

- [x] Relevant tests pass locally
- [ ] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

## Risk

POSIX behaviour is unchanged: the helper keeps calling `os.kill(pid, 0)`
there. On
Windows, `raven web --stop` starts working and stops silently
terminating child
processes it only meant to look at, which is the point of the change.

Rollback is reverting to the inline `os.kill(pid, 0)` in
`raven/cli/serve_commands.py`
and dropping `raven/utils/pid.py`.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
…es (#457)

## Summary

A Windows checkout of this repo with the default `core.autocrlf=true`
rewrites every
text file to CRLF, and two byte-exact gates in `ui-web` then fail on a
tree that is
byte-identical to the one CI passes on Linux:

- the boot-order and open-conversation node tests slice their source on
`"\n}\n"`,
which never matches a CRLF file, so the sliced function comes out empty
and 9 tests
  fail with `ReferenceError`;
- `gen-rpc-client --check` compares its own LF output against the file
on disk and
  reports permanent drift in `src/rpc/generated.ts`.

`.gitattributes` pins the checkout filter so the working tree matches
what the
repository already stores. The repository stores LF today, so this adds
no
renormalization diff: `git add --renormalize .` over the whole tree
stages nothing
beyond the two files in this PR. Only the checkout filter changes.

`ui-web/build.py` is the other half of the same defect. It assembled the
page under
Python's default newline translation, so `dist/index.html` came out CRLF
on Windows
even from an LF source tree, and `check-page.mjs` could no longer
extract the script
payloads (its `<script>\n` anchor needs the LF to be the next byte).
That artifact
ships inside the wheel, so it must not depend on which platform
assembled it.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Run on Windows 11 with `core.autocrlf=true`, in a clean worktree cut
from `main`:

- `git add --renormalize .` after the change stages nothing (confirms no
  renormalization diff);
- `file ui-web/src/rpc/generated.ts` reports CRLF on `main` and LF after
a forced
  re-checkout on this branch;
- `npm run gen:check` (ui-web): `generated.ts matches the contract (178
methods)`;
- `npx vitest run` (ui-web): 108 files, 1836 tests passed;
- `npm run --prefix ui-web build && python ui-web/build.py`:
`dist/index.html` is LF;
- `node ui-web/scripts/check-page.mjs`: `OK (1,482,975 bytes, 3 scripts
parse)`;
- `node ui-web/scripts/count-shared-globals.mjs`, `check-css.mjs`,
  `check-class-namespace.mjs`: all OK.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

## Risk

No runtime behaviour changes for users on Linux or macOS, where the
checkout already
produced LF. On Windows the working tree stops being rewritten to CRLF,
so a
contributor with existing local checkouts may see one large
renormalization touch on
the next checkout; the committed bytes do not move. The built
`dist/index.html` is now
LF on every platform, which is what the Linux-built wheel already
shipped.

Rollback is deleting `.gitattributes` and dropping the `newline=""`
argument in
`ui-web/build.py`.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
…here (#464)

## What

A skill installed through the market had no way to be removed from the
page, because `ext.list` decided "was this installed from the hub" by
looking for one marker file, and the installer most skills arrive by
writes a different one. Closes #456.

Two installers put hub skills on disk:

| | market module (`raven/skill_hub/hub.py`) | context engine +
`use_skill` (`raven/agent/tools/skill_hub.py`,
`raven/context_engine/segments/skills.py`) |
| --- | --- | --- |
| lands at | `<skills>/<name>/` |
`<skills>/hub/<slug>@<version>/<name>/` |
| stamps | `.skillhub.json` | `.install-meta.json` |

`ext.list` tested only for the first, so the second -- the ordinary path
-- reported `hub: false`. The page gates its removal control on that
flag, so it appeared for the exception and never for the rule. On the
machine this was found on, three of four installed skills were the
second kind; the one that showed a remove button was the one that came
through the market module.

## How

Backend only; nothing under `ui-web/` or `ui-tui/` changes. The page
already renders the control when `hub` is true, and already skips the
market-detail fetch when `hub_id` is empty, so both halves light up on
their own once the flag is right.

- `raven/skill_hub/audit.py` names the stamp it writes (`INSTALL_META`)
so the two readers test for the file the installer wrote rather than a
spelling of their own.
- `raven/rpc/methods/console.py` reports `hub: true` when either stamp
is present. The bundle stamp carries a slug, not a market id, so that
skill has an empty `hub_id` -- a state the page already handles.
- `raven/skill_hub/hub.py` lets `skillhub.remove` find a bundle by the
skill name inside it, resolving the wrapper folder the same way
`SkillHubClient._bundle_root` did at install time, and removes the
bundle whole: the bundle is the install unit, and the CLI's `skill
remove` deletes the same folder. Only a stamped directory is eligible --
a folder placed under `hub/` by hand is left alone, which is the rule
`MARKER` already enforced on the other layout.

The second change is not optional given the first: reporting `hub: true`
without it would put a remove button on a skill the handler then
refuses, which is worse than no button.

## Verification

-
`test_ext_list_calls_a_skill_hub_installed_whichever_installer_stamped_it`:
three skills, one per stamp and one with neither, asserting `(hub,
hub_id)` for each -- `(True, "uuid-market")`, `(True, "")`, `(False,
"")`. The marker names are pinned as literals in the test so a drift
between writer and reader is a failure, not a pass.
- `test_remove_reaches_a_bundle_the_other_installer_cached`: a stamped
bundle is removed whole by its skill name; an unstamped folder under
`hub/` is refused and left in place, and so is the neighbouring bundle.
- Mutation-checked: dropping the second stamp from `ext.list` fails
exactly the first test; dropping the bundle lookup from `remove` fails
exactly the second. 139 passed in the two suites on the unmutated tree.
- Dry-run against the real workspace: all four installed skills now
report `hub: true`; the three bundles resolve to their `<slug>@v0`
folders by skill name; the market-module skill and an unknown name
resolve to none.
- Ruff check and format clean.

## Not done here, on purpose

Removal does not stop the context engine re-installing the skill on its
next catalogue hit; the CLI documents the same limit and points at
`skill block`. Whether the page should offer that beside removal is a UI
question and out of scope while `ui-web/` is frozen. Also left as they
were: the two installers keeping two layouts and two stamps -- this
makes the readers agree with both writers rather than deciding which
writer should win.

Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>

---------

Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
Co-authored-by: KT <74288668+0xKT@users.noreply.github.com>
… group (#455)

## Summary

A conversation started in the TUI was listed by the served page and then
could not be opened: clicking it produced an empty transcript.

`SessionManager.session_path` looks for a transcript in this process's
own group, and falls back to searching the other groups when it is not
there. That fallback was gated on `project_slug`, so only a surface that
groups by launch directory ever ran it. The gateway groups by channel
and never did -- yet `list_sessions` scans every group, so the page
offered a conversation filed under a launch directory's slug and then
looked for it under `sessions/tui/`. `session.resume` missed the file,
fell through to its fresh-mint branch, and handed the reader an empty
transcript under a new id.

The write verbs resolve through the same function, so the miss was not
only a failed read. A rename, a pin or a turn from the page filed a
second transcript under the channel group, splitting one conversation
across two files:

```
sessions/-srv-alpha/abc.jsonl -> 2 messages
sessions/tui/abc.jsonl        -> 1 messages   (written by the gateway)
```

Ungating the search is necessary but not sufficient, and the second half
is the part worth reviewing. The glob matches on the chat_id, and a
chat_id names no channel: `cli:<id>` and `tui:<id>` are two
conversations, so an ungated search let a `tui:` save adopt the `cli:`
file and append to it. A candidate is now taken only once its own
metadata key says it is this session. A transcript carrying no key
predates that field and is still adopted on the stem alone, which is the
case the search was originally added for.

Matching on the stored key rather than on the directory name is also
what makes this correct on Windows. `key_from_path` tells a slug from a
channel by a leading `-`, which holds only for POSIX paths; a Windows
launch directory slugs to `C--Windows-system32`, with no leading
separator to read.

Root cause was confirmed against a control rather than by inspection:
two `tui:` sessions, same handler, same manager, differing only in which
group holds the file. The one under `sessions/tui/` resumed with its 26
messages; the one under a slug minted a fresh empty session.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

```
uv run pytest tests/test_session_project_grouping.py -q      2 failed, 32 passed
uv run pytest tests/test_cli_session_commands.py -q          43 passed
uv run ruff check raven/session/manager.py tests/test_session_project_grouping.py
                                                             All checks passed!
uv run ruff format --check <same two files>                  2 files already formatted
uv run ty check raven/session/manager.py                     All checks passed!
```

The 2 failures in `test_session_project_grouping.py` are present on
unmodified `main` in this Windows checkout and are not touched by this
change: `test_project_dir_is_not_rewritten_by_a_later_run_elsewhere` and
`test_slug_collides_for_a_separator_and_an_underscore` both build
fixtures from POSIX paths such as `/srv/alpha`, which `Path` renders as
`\srv\alpha` on Windows.

Regression scope was measured rather than asserted. Running `pytest
tests/ -k "session or transcript or workdir or resume"` (about 1085
tests) with and without the change gives **372 failures before and 372
after, with identical failure sets** -- no test changes state in either
direction. The same comparison over the session-adjacent files alone
gives 156 before and 156 after.

Five tests were added to `tests/test_session_project_grouping.py`,
extending the existing file rather than adding a new one. Two of them
fail without the fix and pass with it:

- `test_the_gateway_opens_a_session_filed_under_a_project_slug` -- the
reported bug;
- `test_the_gateway_appends_to_the_transcript_it_opened` -- the
split-file half;
- `test_two_channels_sharing_a_chat_id_keep_their_own_transcripts` --
the guard against over-correcting, and the case that caught the first
attempt at this fix;
- `test_a_keyless_transcript_is_still_adopted_across_groups` -- pre-key
transcripts stay reachable;
- `test_the_gateway_still_files_its_own_sessions_by_channel` -- reaching
across groups is a fallback for an existing transcript, not a change of
where this process writes.

The first attempt, which ungated the search without the key check, broke
`test_cross_channel_ambiguous_same_id_two_channels` in
`tests/test_cli_session_commands.py`. That failure is what produced the
key check.

Fix verified against the real workspace that produced the report: the
affected session now opens with both of its messages, resolving to its
actual file under the launch directory's slug.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

No docs change: this corrects the behaviour the surrounding docstrings
already described, and coins no domain term.

## Risk

User-visible change: a surface can now open a transcript filed under
another group. On the gateway this closes a gap rather than widening
reach -- `list_sessions` already scans every group, so those sessions
were already listed and offered; only opening them was broken. For a
slug-grouped surface the search is unchanged apart from the new key
check, which makes it strictly stricter: a candidate that would
previously have been adopted on its stem alone is now refused when its
stored key names a different session.

Transcripts written before the metadata `key` field are unaffected --
they answer no key and are still adopted on the stem.

`session_path` now reads the first line of each candidate when the
direct path is absent. The read is one `readline` on a file the caller
is deciding whether to open, and the candidate set for an exact chat_id
is normally empty or a single file.

Rollback is a revert of the single commit; the change is confined to one
function and one new module-private helper, and writes nothing to disk
that a revert would strand.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
Co-authored-by: KT <74288668+0xKT@users.noreply.github.com>
…t moved on (#476)

## Summary

The installers stop telling you what to run and run it. `install.sh`
ends with
`raven web --stop && raven web --foreground`, so a finished install is
the page
open in a browser rather than a hint to go and type something; first-run
setup
is walked through by the page's own onboarding. The capability summary,
the
upgrade-versus-first-run wording and the config.json probe that fed it
go with
it. `install.ps1` mirrors all of this, with one deliberate divergence: a
non-zero page exit only warns there, because under `irm | iex` that is
the
reader's own interactive PowerShell and Ctrl-C, the ordinary way to end
a
foreground page, returns non-zero.

Two gaps that only bite a source install are closed on the way:

- `build_web_assets` compares each artifact against its sources instead
of only
testing that it exists. An editable install relinks Python and nothing
else,
so `ui-tui/dist/entry.js` and `ui-web/dist/index.html` kept serving
whatever
the tree held at first install while every `.py` beside them moved on.
The
page watches `i18n/` too, since `ui-web/build.py` inlines the catalogue.
`node_modules`, `dist` and `.modern` are pruned: `npm ci` rewrites the
first
  on every build, and the other two hold the artifact being judged.
- A node packaged without npm now fetches a private runtime that carries
one.
Debian and Ubuntu split the two packages, so a system node >= 22
satisfied
`ensure_node` and the build then found no npm with no private runtime
ever
provisioned. The download body is split out as `provision_private_node`
and
called from the build probe as well, subshelled so a failure there is a
skipped build rather than a failed install. It stays lazy rather than
moving
into `ensure_node`, so a wheel install that never runs npm does not pay
for a
  50 MB download it has no use for.

Both READMEs are reordered and the clone install is documented under
Quick
Start, where it was missing: `./install.sh` and a piped run do different
things
from inside a checkout, because local mode requires `$0` to be a real
file, and
`RAVEN_LOCAL_SRC` is the opt-out. The Chinese README also picks up the
EverMind
ecosystem table the English one had already moved to.

One test changed rather than being deleted.
`test_readme_quickstart_matches_the_installer_hint`
required an `Onboard and run` heading in both READMEs; the installers no
longer
close with a first-run hint and neither README carries that section, so
the pin
had no subject left. It is repointed at what the two must still agree
on, which
is the clone install.

`install.ps1` has not been parsed or executed. There is no pwsh on the
machine
this was written on and no PowerShell step in CI, so every Windows
change here
is reviewed by eye and pinned by text tripwires only. A real Windows run
before
release would be worth it.

## Type

- [ ] Fix
- [x] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

`uv run python -m pytest tests/ --ignore=tests/integration` on the
merged tree:
23161 passed, 109 skipped.

`uv run pytest tests/test_install_script.py`: 27 passed, including a
functional
test that extracts the real `is_stale` and `resolve_node_dir` from
`install.sh`
and runs them under `sh`. `is_stale` is checked across missing artifact,
fresh
artifact, touched page source, touched catalogue, and build output
touched;
`resolve_node_dir` across system node with npm, node without npm, node
without
npm and a failed fetch, and no node at all. Both were mutation-checked:
removing
the fallback line from `install.sh` turns the new tests red.

`sh -n install.sh`, `make check-source-language`, `make
check-large-files`,
`ruff check`, `ruff format --check`: all clean.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

## Risk

User-visible: a finished install now holds the terminal on a running
page
instead of returning to the prompt. That also means the installer's exit
code is
the gateway's, so on a clone where the page build only warned,
`install.sh` now
exits 1 with "No page is built" where it used to exit 0 with a working
TUI. A
re-run over a checkout that has moved on will rebuild the frontend,
which costs
an `npm ci` it used to skip.

Rollback is per commit: the launch, the rebuild-on-change and the npm
fallback
are independent edits to `install.sh` and `install.ps1` and can be
reverted
separately. No runtime code outside the installers changed.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

---------

Co-authored-by: admin <admin@SH-HuKai.local>
Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
## Summary

Turn "bug fix to regression case" from a manual chore into a gated
pipeline, and give trajectory regressions their own CI gate.

- `raven trajectory replay --json`: stable machine-readable report
(schema_version 1) with defensive JSON coercion; on a halted replay the
full JSON is still emitted before exit code 2.
- Case metadata contract: `case.yaml` requires non-blank
issue/owner/why/re_record; residual-scan findings pass only with an
explicit per-token human sign-off (`reviewed_residuals`: full-token
sha256 plus a verbatim reason, nothing exempted automatically -- a
machine-produced redaction report is not approval).
- `raven trajectory regression validate [CASE_DIR|--all]`: static commit
gate -- whole-directory discovery so a broken case cannot hide, both
schemas, cassette completeness down to the replay contract (valid JSON
is not a usable payload: missing payloads via the extracted
replayability check, corrupt inner fields via `validate_recording`),
rejection of files the residual scanner would silently skip, and a 256
KiB / 1 MiB size budget; every parse boundary collects problems instead
of raising, so one malformed case cannot abort the --all sweep.
- `raven trajectory regression init <source> --name <case>`: scaffold
from a bundle directory, attempt id, trajectory report tarball, or bug
report package (two-layer extraction; both passes accept only regular
files/directories with root-contained names). Minimize lands in staging;
residual findings are reviewed interactively (the full token is shown
locally, files receive digest and reason); publication copies onto the
destination filesystem under a `.init-` staging name, re-checks the case
name, then renames on one device -- failures clean up only what the run
created.
- Vacuous-pass guard (`MIN_COMMITTED_CASES`, raised by hand) and a
`trajectory` CI job: replay the committed cases, run `validate --all`
even after a pytest failure so both gates report, upload per-case
failure reports (replay --json schema plus check failures, mode from
each case's expect.yaml) as a 14-day artifact.
- Docs: the Trajectory Regression Case entry in CONTEXT.md extended in
place; `tests/trajectories/README.md` documents the per-case layout and
workflow.

Also: `config_secrets_loaded` (a field name of redaction.json itself, a
guaranteed self-referential false positive in every cassette) joins the
residual scanner's benign literals, and both sample cases gain real
`case.yaml` metadata.

## Type

- [ ] Fix
- [x] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

- `uv run pytest tests/ -q` (after rebasing onto latest main): 23149
passed; the 37 failures and 20 errors are environment-bound and
reproduced identically on a pristine origin/main worktree without these
changes.
- `uv run pytest tests/test_trajectory_regressions.py
tests/test_cli_trajectory_regression_commands.py
tests/test_cli_trajectory_commands.py tests/test_trajectory_replay.py
tests/test_trajectory_cassette.py tests/test_trajectory_redact.py -q`:
282 passed (about 90 new tests across the five steps).
- `uv run ruff check .`, `uv run ruff format --check .`, `make
lint-imports`, `make lint-deps`, `make check-large-files`, `make
check-source-language`: pass. `make lint-types` reports one diagnostic
in `raven/agent/tools/web.py`, a file this branch does not touch,
pre-existing on main.
- CI workflow YAML parses (`yaml.safe_load`); the first real run of the
new `trajectory` job happens on this PR (see this PR's checks).
- CLI assertion width-independence re-verified with `COLUMNS=60`.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

## Risk

- All additive: the default `replay` output path is unchanged;
`regression init/validate` are new subcommands; the one scanner change
adds a fixed benign literal (any variation still flags).
- The new `trajectory` CI job becomes a required-passing gate for PRs;
to lift it in an emergency, revert the CI commit (`ci: add trajectory
regression gate job and case-count guard`) -- the remaining commits are
independent.
- Tar extraction accepts only regular files and directories with
root-contained relative names on both layers, with `filter="data"` kept
as defense in depth.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

---------

Co-authored-by: 江国庆 <guoqingjiang@deepglint.com>
Co-authored-by: Claude (claude-fable-5) <noreply@anthropic.com>
## Summary

Present Raven's orchestration capabilities and four specialized agents
with benchmark figures, and keep the English and Chinese READMEs
aligned.

- Add the Multi-Agent Orchestration Benchmark at the end of the
introduction. Compare Hermes, Claude Code, and Raven with shared 0-1
scales for Node F1, Edge F1, Partial Order Accuracy, and Exact Match
Rate.
- Move Raven Agents before Benchmarks and add concise Research, Code,
Design, and Oncall descriptions with their figures.
- Replace the third-party agent table with a centered card image and
explain how Raven connects to and orchestrates these agents. Host image
files as GitHub attachments.
- Standardize primary heading markers and align section order, captions,
model settings paths, and the EverMind ecosystem catalog across both
languages.

## Type

- [ ] Fix
- [ ] Feature
- [x] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Commands run after rebasing onto the latest main:

```text
uv run --no-sync pytest tests/test_readme_scope_canon.py tests/test_living_docs.py -q -o addopts=''
# 4 passed
uv run --no-sync pre-commit run --files README.md README.zh-CN.md
# All applicable checks passed
make check-large-files
# Passed
make check-source-language
# Passed
env PYTHONPATH=. uv run --no-sync python scripts/check_commit_messages.py origin/main..HEAD
# Passed
git diff --check
# Passed
```

Compared all 27 headings, 12 code blocks, image URLs and sizing,
commands, paths, numeric values, and architecture graph structure across
both languages. Verified all 24 orchestration scores against the
supplied results and confirmed that the uploaded chart matches the local
PNG byte for byte.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

## Risk

This changes documentation structure, wording, and linked figures.
Runtime behavior is unchanged. Figures depend on GitHub attachment
availability. Roll back by reverting the changes to README.md and
README.zh-CN.md.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

#462, #465, #474, #477, #479

---------

Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
…the roles (#468)

## Summary

The agent loop already names four strategy roles (Memory, Planning,
Capability, Action), but two judgements a turn makes still lived in the
shell around them. This moves both onto the roles, in two dependent
commits, with no behaviour change intended.

**1. The charter reads through Memory and Action.** A dispatched
sub-agent's playbook (`_meta["raven.playbook"]`) was read from four
places inside the loop and the tool registry. `prompt` and `stopWhen`
are now read by `DefaultMemory.assemble()`, which fills `task_brief` and
`task_done_when` on the `TurnContext` the identity segment renders.
`checks` and `code` are read by `DefaultAction.judge()`, which
`ToolRegistry.execute` asks once per proposed call, the way it already
asks the permission gate. `tools` stays in `withheld_names()` and the
registry still decides what a refusal does, so a replaced role withholds
nothing it could not already withhold.

**2. Mid-turn window shrinking seats on Memory.** The five recoveries
the loop ran inline go through one method,
`MemoryModule.shrink(messages, pressure, state, model)`, with
`WindowPressure`, `WindowState` and `ShrinkResult` as the carriers:

- proactive compaction, once the last billed reading crosses the trigger
- the standing image window
- a provider's overflow error
- a picture refused inside a tool result
- pictures refused for size

The retry mechanism stays in the shell. The six `continue` sites and the
iteration rollback did not move, so a hook sees a retried call exactly
as it saw the first. `WindowState` is a per-turn object the shell owns
and hands to each call, because the role outlives the turn;
`image_window` has no default, since 0 is a real value meaning "no
picture stays".

The pure pieces both sides need move to a new package,
`raven/agent/window/`, which the import contracts treat as cargo that
may not import the loop shell. `REASONING_EFFORT_LADDER` moves to
`contracts/llm_provider.py` because the loop's empty-response descent
and the window's head summary both read it and neither package may
import the other.

Contract papers: `CONTRACTS_VERSION` 26 to 28, `CONTRACTS_LINE_CEILING`
3,040 to 3,200, each bump with its own review paragraph in
`tests/test_kernel_budget.py`.

A follow-up PR carries the third step, which gives a sub-agent nine
verbs in place of the six hook phases. It is kept separate because it
touches every bundled plugin.

## Type

- [ ] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [x] Refactor
- [ ] Other

## Verification

Gates, run at each of the two commits:

```
uv run ruff check --no-cache raven evolver agents plugins-dist tests scripts
uv run ruff format --check --no-cache raven evolver agents plugins-dist tests scripts
uv run lint-imports
uv run pytest -q
npx commitlint --from origin/main --to HEAD --config commitlint.config.cjs
uv run python scripts/check_commit_messages.py origin/main..HEAD
uv run python scripts/check_source_language.py origin/main..HEAD
uv run python scripts/check_large_files.py origin/main..HEAD
```

Import contracts: 10 kept, 0 broken at both commits. The suite is green
at both; the failures this machine shows are the same ones it shows on
an unmodified `main` (missing browser and optional packages), and
neither commit adds one. The loop's own e2e files, which the default
selection leaves out, were run too (charter, playbook, truncation,
direct chat, fallback chain, ACP stdio): no failure that `main` does not
also show.

Behaviour preservation. Each commit was run against a worktree of its
own parent under the same scripted provider, outputs normalised for
nonces and temp paths, then byte-compared. Four scenarios, eight
comparisons, all identical:

- a whole turn with no playbook: host turn, spawned sub-agent, tool
calls and refusals
- a playbook-on worker lane in process, including two charter refusals
- six window-shrink scenarios: proactive compaction, the standing image
window, overflow, a refused picture, oversized pictures, the effort
ladder
- the five-plugin hook seam, every phase's decision and metadata per
iteration

Mutation check on the second commit: breaking each of the five moves
inside `DefaultMemory.shrink` turns exactly its own scenario red, so the
comparison is not vacuous.

Real model, Sonnet 4.5 through OpenRouter, at both commits:

- playbook on, in-process lane: both workers receive the charter, write
their files and finish
- playbook on, ACP fork lane: the charter prompt crosses the process
boundary, and a charter check blocks a `write_file` outside the fence

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

## Risk

No user-visible behaviour change is intended, and the comparisons above
are byte-identical on every path they cover.

For anyone replacing a harness role: `MemoryModule` gains `shrink`, so a
replacement written against the old protocol needs it.
`DefaultMemory.bind()` names the missing member at assembly rather than
failing inside a turn.

Rollback is a revert of the merge commit. Nothing persisted on disk
changes shape.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

---------

Co-authored-by: yao pengfei <yaopengfei@shanda.com>
Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
Co-authored-by: KT <74288668+0xKT@users.noreply.github.com>
## Summary

Synchronize the Chinese README with the updated English overview. The
change aligns the Raven introduction and built-in agent section,
including the Host Agent positioning, Raven Evolver description,
built-in agent wording, and orchestration readiness statement.

The existing benchmark charts, links, captions, and the remaining
bilingual sections are preserved.

## Type

- [ ] Fix
- [ ] Feature
- [x] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

Commands run:

- `uv run --no-sync pytest tests/test_readme_scope_canon.py
tests/test_cli_onboard_commands.py -k readme_quickstart -q` (1 passed)
- `uv run --no-sync pytest tests/test_docker_runtime.py -k
container_exposes_no_engine_selector -q` (1 passed)
- `UV_NO_SYNC=1 make check-large-files check-source-language
PYTHON_VERSION=.venv/bin/python COMMIT_RANGE=HEAD` (passed)
- `git diff --check` (passed)

## Risk

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

This is a documentation-only change. Reverting the commit restores the
previous README wording.

## Related Issues

N/A

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
## Summary

Five upgrade-path and reporting defects in the embedding pin that #414
landed. Three were named as bounded follow-ups in that PR's final
review; two came out of driving the acceptance environment afterwards,
and both of those were invisible to the unit suite because every fixture
already held the new shape.

The two found by acceptance:

- **A config still holding the retired endpoint shape does not load at
all.** The block used to carry the address and the key;
`EmbeddingConfig` forbids extras, so `raven doctor` ends in a traceback
rather than in a degraded feature, and so does every other command that
reads the extension blocks. The migration takes those keys now: the
address is adopted onto whichever configured provider answers there, and
when nothing answers the whole block goes rather than half of it, with
the dropped address named in the log. A model left with no provider is a
pin every reader resolves to nothing while every screen reads it as
configured.
- **A pin could not name a provider this config holds.** The provider
half was checked against the registry alone, and Raven carries no spec
for every vendor LiteLLM can reach. A pin the wizard had stored and
every reader resolved came back from the settings page as a provider
that does not exist, leaving the endpoint uneditable on the page that
exists to edit it. A section the operator wrote counts as proof the
vendor exists; a name nothing holds is still refused as a typo.

The three from review:

- **Clearing the pin was a no-op.** The picker's inherit option sends
both halves empty and the writer dropped empty values before deciding
anything, so the call returned applied while the file kept the old pair.
- **Doctor reported a move it had not made.** The endpoint only moves
when a configured provider answers at the retired address, and the
return value was discarded. The unmade move stays in the remaining list
now, with a line saying why.
- **The canonical SkillForge glossary entry still said
`skillForge.everos`** and attributed the local extraction pipeline to an
embedded copy of the memory backend.

## Type

- [x] Fix

## Verification

- `uv run --all-extras pytest -q` -- 23058 passed, 110 skipped, 4
failed; the four are the same pre-existing failures main carries
(`test_agent_loop_token_budget`, `test_agents_code_launcher`,
`test_agents_oncall_launcher`, `test_ppt_engine_image_search`), matched
by test id against a run on the merge base.
- `uv run ruff check .` and `uv run ruff format --check .` -- clean.
- `make check-commits`, `make check-source-language`, `make
check-large-files` -- clean.
- Every fix was reproduced before and after against a real config or a
real gateway, not only through its test: the unloadable config came from
the acceptance home itself and `raven doctor` was run on it before and
after; the provider refusal was reproduced over a real WebSocket to a
running gateway; clearing and the doctor report were each run end to
end.
- Each new test was checked by removing its fix and confirming it fails.
- Wider acceptance on the running product, with the memory plugin and
without it: a gateway with no memory plugin installed built a knowledge
base whose width (4096) was measured from the live endpoint and answered
a query sharing almost no keywords with the source (score 0.728); the
real TUI recalled a fact stored in an earlier turn, with the recall
payload in the audit artifact as evidence rather than the model's
answer.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

## Risk

Two changes widen what is accepted rather than narrowing it: a provider
section counts alongside the registry, and a migration edits a block it
previously left alone. The migration only runs for a config carrying the
retired keys and is stamped, so it runs once; a config already in the
current shape is untouched. Rolling back any single commit is safe --
they do not depend on each other.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ith (#483)

## Summary

An EverOS agent case carries three fields: `task_intent`, `approach` and
`key_insight`. Recall built `Memory.text` from the first and the last
only, so the method a case was solved with never reached the prompt. The
reader saw what was attempted and what was concluded, with the steps
between them missing.

The other two consumers of the same row already keep all three fields:
the sub-agent memory record joins intent, approach and insight, and the
memory page maps `approach` to a card's body. Only the recall path that
feeds the model dropped it.

This joins the three fields in field order. Two constraints shaped how:

- Intent stays on the first line. The skill router takes this text's
first line as a hit's display name and as its line in the search
listing, so moving intent off line one would rename every recalled case.
- Empty fields drop out, so an EverOS that answers without an `approach`
still reads as prose instead of carrying a blank paragraph.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

```
uv run --frozen pytest tests/test_everos_backend.py tests/test_skill_router_everos_source.py -q
182 passed in 17.69s

make lint-python lint-types
ruff check: All checks passed!
ruff format: 1960 files already formatted
ty check: All checks passed!

make check-source-language
clean, no findings
```

The case test asserts the exact joined text rather than substring
membership, and it was run against the pre-change conversion first to
confirm it fails there (AssertionError on the missing approach). A
second test covers a row that arrives without an `approach`, which is
the shape that would otherwise produce a blank paragraph.

No user-facing doc states the recall text shape. The CONTEXT.md passage
that lists all three case fields describes the sub-agent memory record,
which is a different code path and is unchanged.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

## Risk

Each recalled case now carries one more paragraph, so the agent track's
share of a prompt grows by roughly the length of one `approach` per case
hit. Nothing downstream parses this text: the skill router reads the
first line for a display name and passes the remainder through
untouched, and that first line does not move.

Rollback is a revert of the single commit.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
## Summary

Synchronize the Chinese README with the updated English README. The
change aligns the Raven introduction, benchmark captions, built-in agent
copy, third-party agent positioning, and the removal of the obsolete
benchmark summary section.

The change also adds the updated wording for data analysis, PowerPoint
slide generation, and unattended AI4S benchmark captions while
preserving all chart links and remaining bilingual sections.

## Type

- [ ] Fix
- [ ] Feature
- [x] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

Commands run:

- `uv run --no-sync pytest tests/test_readme_scope_canon.py
tests/test_cli_onboard_commands.py -k readme_quickstart -q` (1 passed)
- `uv run --no-sync pytest tests/test_docker_runtime.py -k
container_exposes_no_engine_selector -q` (1 passed)
- `UV_NO_SYNC=1 make check-large-files check-source-language
PYTHON_VERSION=.venv/bin/python COMMIT_RANGE=HEAD` (passed)
- `git diff --check` (passed)

## Risk

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

This is a documentation-only change. Reverting the commit restores the
previous README wording.

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
…ch (#485)

## Summary

Enabling an entrance from the page (weixin, for one) builds and starts
the adapter on the spot through `ChannelManager.start_one`. The gateway
command hands every channel its spine dispatch in one loop at launch,
and that loop has already run by then, so a channel started this way has
no `intake._submit`: it logs `no spine dispatch wired` once per message
and drops every one, while the page draws it connected and the adapter's
own login has succeeded.

The outlet half of the same hazard was fixed earlier:
`channels.on_started` registers an outlet so a hot-started channel can
be replied to. The intake half stayed a launch-time loop, and
`start_one` fires `on_started` and nothing else, so nothing after launch
ever calls `set_submit` for a channel born later.

The fix composes the intake onto the hook the outlet already uses. A
module-level `_wire_channel_intake(channels, dispatch)` sets the submit
on each channel the manager holds and wraps `channels.on_started` so a
channel started later gets the outlet it already got plus the intake;
the command calls it at the point where `_inbound_dispatch` is born.
Channels present at launch are wired exactly as before. The helper is
module-level rather than a closure in the command body so that it can be
executed by tests, which the command body cannot be. Nothing under
`ui-web/` or `ui-tui/` changes.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

- `.venv/bin/python -m pytest tests/test_cli_gateway_commands.py -q` on
this branch rebased onto current `main`: 50 passed. Four new cases drive
`_wire_channel_intake` through fake channels and a fake manager: every
channel present at launch gets the dispatch; a channel the manager
starts later gets it through `on_started` while the outlet hook that was
already there still runs; a manager with no outlet hook still wires the
intake; and the command body is pinned by source to call the helper with
`_inbound_dispatch` and to have lost the launch-only loop.
- Mutation-checked both ways: dropping `wire(ch)` from the composed hook
fails exactly the two cases that start a channel late; dropping the
outlet call fails exactly the outlet assertion of the hot-started case.
- `python scripts/coverage_gate.py diff --base-ref upstream/main
--threshold 90` after the file ran under `--cov=raven --cov-branch`:
91.67% (11/12 executable changed lines). The one uncovered line is the
call inside the `gateway` command body, which no unit test executes, the
same as every other line of that body.
- `.venv/bin/ruff check` and `.venv/bin/ruff format --check` on
`raven/cli/gateway_commands.py` and
`tests/test_cli_gateway_commands.py`: clean.
- The drop was reproduced on a live gateway before the fix (2026-09-14):
enable, disable and re-enable weixin through the page's own RPC puts the
adapter on the `start_one` path, and every inbound message was logged as
`no spine dispatch wired` and dropped.
- No user-facing docs change; the behaviour was always the documented
one.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

## Risk

- Security: none. No new input path; the dispatch a hot-started channel
now receives is the same one launch-time channels already had.
- Backward compatibility: channels present at launch are wired by the
same call as before; the only change is that a channel started later is
wired too.
- Rollback: revert the squash commit; the hook falls back to the
outlet-only lambda.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
…es (#469)

## Summary

Follows #468 (merged as `869ad78a`), which moved the charter reads and
the window's mid-turn shrinking onto the harness roles. This branch is
rebased onto `main` and carries only its own three commits.

A sub-agent's product logic was an `AgentHook` subclass: six phase
methods over a fourteen-field context it could write, with the turn's
own state kept in the context's free-form `metadata` dict. Surveyed
before this change, the five bundled plugins had 71 phase methods, 42 of
them doing real work, reading fourteen context fields and using six
decision fields. Every one of them fits a small set of verbs. This PR
makes those verbs the contract, in three commits.

**1. A sub-agent writes a conduct instead of six hook phases.** New
factory-loop-tier paper `contracts/agent_conduct.py`: `AgentConduct`
with nine optional verbs (`intake`, `select_tools`, `advise`,
`system_addendum`, `review`, `salvage`, `outbound`, `archive`,
`observe`), answered against a read-only `StepView` of fourteen fields,
with `Intake` and `Verdict` (Accept / Resample / End) as the answers and
a `ConductFactory` because the host builds one conduct per turn.

It sits beside `loop_hooks` rather than replacing it. The six phases
remain the loop's *timing* contract: when the loop asks, in what order,
what a refusal does. This is the *judgement* contract: given one step,
what does this agent say about it. `ConductHook` seats one in the other,
so a plugin that still ships a plain `AgentHook` keeps working and the
composite's ordering, merging and rollback are untouched.

All five bundled plugins are ported: design-engine, ppt-engine,
code-flow, oncall-flow and research-flow. Research keeps its six-phase
gate chain: the gates now subclass a plugin-private `Gate` and read a
`GateCtx` the conduct builds from its `StepView` and its own
turn-private facts, so six thousand lines of gate bodies did not move
and its wrappers still see the cross-gate state they read.

The turn's conduct and the system-message addendum it splices live in
the turn's own `metadata` dict, not on the hook: one hook instance
serves every turn a process runs, and a system turn can overlap a user
one on the same chain.

**2. The conducts seat on the roles.** `MemoryModule.intake`,
`PlanningModule.advise`, `ActionModule.review` and
`ActionModule.salvage` each take this turn's conducts, so a replaced
role decides what a plugin's judgement does rather than the seat
deciding it. The composition rules live in
`raven/agent/harness/conducts.py`: the first non-Accept verdict wins,
advice joins in order, the first salvage wins, an intake threads the
text through each conduct. `run_turn` binds the harness for the turn the
way it binds the model, and the seat falls back to the same rules when
nothing is bound, which is how a plugin test drives a hook directly.

Contract papers: the conduct paper is factory-loop tier, so the
contract-tier digest and `CONTRACTS_VERSION` do not move with it.
`CONTRACTS_LINE_CEILING` 3,200 to 3,420, each bump with its own review
paragraph in `tests/test_kernel_budget.py`.

**3. What the review found.** Eleven threads, all resolved in the third
commit. The ones that changed behaviour rather than prose:
`CodeFlowHook` kept its `rolls_back_iterations = False` (inheriting
`True` from the new base had it hold every reply behind the draft gate
instead of streaming); `StepView` now publishes deep-frozen nested
values, since freezing only the outer tuple left a conduct able to write
through to `ctx.messages`; a turn's conducts are collected at the role
boundary, so a replaced role sees one batch of the turn's judgements
rather than one call per seat; an addendum is located by its own text
rather than a saved offset, which two conducts appending system text
used to shift out from under each other; `Planning` and `Action` gained
binders, so a replacement role missing a new member is named at assembly
rather than swallowed by the composite's `except Exception`.
`CONTEXT.md` carries the Agent Conduct term and the Agent Hook entry no
longer describes what this PR replaces.

## Type

- [ ] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [x] Refactor
- [ ] Other

## Verification

Gates. CI runs these at the branch tip (17 pass, 1 skipped -- the
skipped row is the disabled Claude Mention workflow). Locally, at each
of the three commits:

```
uv run ruff check --no-cache raven evolver agents plugins-dist tests scripts
uv run ruff format --check --no-cache raven evolver agents plugins-dist tests scripts
uv run lint-imports
uv run pytest -q
npx commitlint --from origin/main --to HEAD --config commitlint.config.cjs
uv run python scripts/check_commit_messages.py origin/main..HEAD
uv run python scripts/check_source_language.py origin/main..HEAD
uv run python scripts/check_large_files.py origin/main..HEAD
```

Import contracts: 10 kept, 0 broken at each of the three commits. The
suite is green at each, and no commit adds a failure that `main` does
not already show on this machine (this checkout has the heavy extras
uninstalled, so the local run carries failures `main` carries too -- CI
is the authority, and it is green). The loop's own e2e files, which the
default selection leaves out, were run too (charter, playbook,
truncation, direct chat, fallback chain, ACP stdio): no new failure.

Behaviour preservation. The oracle for this change is the hook seam
itself: all five plugins driven through every phase of a multi-iteration
turn under the same scripted provider, recording each phase's decision
and the observers it files. Outputs normalised for nonces and temp
paths, then byte-compared.

Re-run after the review fixes at the branch tip `81e6f423` (all three
commits), against `main` at `c65c2059`:

| scenario | bytes |
| --- | --- |
| five-plugin hook seam | 23,840 == 23,840 |
| a whole turn with no playbook (host turn, spawned sub-agent, tool
calls and refusals) | 12,821 == 12,821 |
| six window-shrink scenarios | 133,371 == 133,371 |

These cover the tip, not each commit separately. The seam comparison is
the one that matters most for the third commit: its first fix changed
`rolls_back_iterations` for code-flow, which decides whether replies
stream or are held, and the seam oracle records the decision at every
phase, so an unintended change there shows as a diff rather than as
silence.

A unit test proves the seat is a real decision point rather than a
pass-through: with a lenient Action role bound, a conduct's Resample is
not applied, and Planning's advice and Memory's intake are what the loop
receives.

Real model, Sonnet 4.5 through OpenRouter, at the branch tip:

- playbook on, in-process lane: two workers, each receives the charter
(3 checks), writes its file, finishes; 26 recorded model calls
- playbook on, ACP fork lane: the charter prompt crosses the process
boundary (the fork answers ZEBRA where the task text says PONG, so the
brief reached the forked model through Memory), and a charter check
blocks `write_file` to `notes/hi.txt` against a `pathPrefix: out/` rule
- Raven-Research, a fork running the full gate chain: `plain-first:
escalated_tool`, then `sufficiency-gate: released on the pages at
iteration 4 after 2 searches / 2 pages`

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

## Risk

No user-visible behaviour change is intended, and the hook-seam
comparison is byte-identical on every phase it covers.

For plugin authors: the six-phase `AgentHook` is unchanged and a plugin
shipping one keeps working. `ConductHook` is an adapter that seats a
conduct in that same chain.

For anyone replacing a harness role: the role protocols gain `intake`,
`advise`, `review` and `salvage`, so a replacement written against the
old protocol needs them.

Rollback is a revert of the merge commit. Nothing persisted on disk
changes shape.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

---------

Co-authored-by: yao pengfei <yaopengfei@shanda.com>
Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
…res (#410)

## Summary

Absorbs the first of two remaining commits from the research fork before
it stopped
iterating. This one lands instruments only: nothing here changes what a
shipped run does,
and the research flow label does not move because no model-visible text
changed.

`raven/agent/loop/dead_end.py` (new) asks whether a finished turn ended
on the model's own
prose. The test is structural, with a textual sub-label for the tail
shape. Three decisions
are worth review:

- It lives under `raven/agent/loop/` rather than in the research plugin,
because the loop
is what would re-run a dead turn, and a host module cannot import a
plugin without
  breaking `make lint-imports`.
- The ask labeller is injected (`ask_kind`) rather than imported,
because only the product
that writes the harness asks knows how they are worded. With no labeller
the failure mode
is still recognised structurally and reported as `harness_ask_unknown`,
so the boolean
  never depends on wording.
- The fork marked its injected turns with a key of its own. This loop
already had an
equivalent in `_HOOK_INJECTED_KEY`, so the predicate reads that rather
than adding a
  second marker for the same fact.

The conditional rerun that reads the predicate arrived on this branch
afterwards, in
#418, along with the turn's wall clock. Both are supplied as data a hook
writes onto the
turn's metadata rather than as config the loop reads, so an agent whose
hooks leave no
budget runs exactly as it did before either existed: the rerun is
unreachable for it and
the clock is never consulted. The rerun starts from the original
question rather than
salvaging the failed attempt, because salvage was measured on 25 real
duds to produce
content 23 times and the right answer none of them, turning a detectable
zero into a
confident wrong answer.

Three defects in that work were found by review and fixed on this
branch, each with a
reproduction that is red on the commit before it. The loop numbered
episodes from an
iteration count that restarts, so a rerun emitted indices [0, 1, 2, 0]
and the TUI, which
keys episode rows and their fold state by that index, had the rerun's
first step collide
with the first attempt's. Each attempt took its own clock reading, so a
first attempt
ending just under the limit handed the rerun a fresh full budget and a
ten-second turn ran
to eighteen. And the research hook opens its per-turn ledger at
iteration 1, which a rerun
reaches again, so the second open replaced the path the close reads and
left the first
attempt's file unreachable on disk.

Reading that back afterwards turned up three more, none of them raised
in review. The
turn's end record carried the attempt's iteration count beside the
turn's elapsed time
with nothing saying which was which, so it gained ``attempt``. The hook
contract described
that record as "status and iteration count" while it carried five
fields, which hid
``stopped_by`` from anyone writing a terminal gate against the contract.
And the rerun's
progress line named research in a loop that serves every agent, and
called itself the
first attempt when the budget allows more than one.

A verbatim body sink (`RAVEN_VERBATIM_SINK`) records search and fetch
results at three
chokepoints. It is a second file rather than more columns on the ledger
because a 250 KB
row would leave no complete line in the 64 KB tail window
`_resolve_session_seq` reads,
which would silently reset the re-run split. Off unless the environment
names a file, and a
write failure is logged rather than raised.

`DigestOutput` lets a digest return a write-only ledger annotation
beside its text. The
annotation is stripped from the tool result and merged into the fetch
row, so a digest can
never label its own output where the model can read the label. Nothing
emits one yet; this
is the pipe a later digest sidecar uses.

`is_hard_tool_failure` was blind to the whole reader path. `web_fetch`
never returns a bare
string, it returns a JSON envelope, so a spent quota, a rejected URL and
a failed
validation all read as success and reset the streak `web_search` had
accumulated.
Transient markers are still tested first, so a 429 inside an envelope
stays retryable.

Forced finalize now books its own salvage spend. The phase fires on
every run that produced
no answer and burns up to `max_tokens` doing it, and it was counted
nowhere.

The Chinese refusal opener lives in `raven/i18n/zh_lexicon.py`, which is
the exemption zone
this repo keeps Chinese forms in, and the predicate imports it from
there rather than
carrying a literal outside the zone.

Two defects found in the first review round are fixed on this head, both
reproduced before
being changed.

The verbatim sink borrowed the ledger's `_resolve_session_seq`, which
recovers the re-run
split by reading the last 64 KB of the file and parsing the last
complete line. That works
for the ledger because its rows are small; this sink stores the large
ones, so the first
record wider than the window leaves nothing parseable and the sequence
resets to 1, making
two runs indistinguishable. Runs are now told apart by a tag minted in
memory once per
process and never recovered from disk, which removes the read a large
row defeats. The
argument for keeping this file separate from the ledger was already in
its docstring; it
applies to the sequence too, and now says so.

`is_hard_tool_failure` started counting JSON envelopes while
`failure_class` still read
them as undifferentiated text. The streak key is `(tool, failure_class)`
and the break
threshold is two, so a blocked URL followed by a reader's HTTP refusal
fired the
stop-repeating nudge at a model that had changed both its cause and its
approach. Both
predicates now share one envelope parse, and an envelope is classified
by its own `error`
string rather than by a textual ladder that cannot see into one.
Coarseness is for model
prose, which varies without meaning anything; a tool writes its error
from a small fixed
vocabulary, so equal strings are the same failure and different ones are
different. The
parts that vary with the page travel in the envelope's other keys, so
they cannot split a
streak.

The last commit removes an `isinstance` guard in the envelope reader
that the diff-coverage
report flagged as the only uncovered changed line. JSON has one shape
that opens with a
brace, so the branch was unreachable; removed rather than covered.

## Type

- [ ] Fix
- [x] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Run at this head (95c975b), on `origin/main` with nothing to rebase
onto.

```
uv run --frozen --python 3.12 --all-extras pytest -q tests/
  22132 passed, 43 skipped, 109 warnings in 388.67s
```

Integration and e2e are deselected by default (`-m "not integration and
not e2e"`), so the
run above does not include them. This change adds no integration
surface.

```
make coverage
make coverage-ratchet
  line   current=88.66% baseline=87.30% delta=+1.36pp
  branch current=81.59% baseline=79.43% delta=+2.15pp
  Coverage ratchet passed with 0.05pp tolerance.
make coverage-baseline-check
  Coverage baseline update is monotonic.
COVERAGE_BASE_REF=origin/main make coverage-diff
  Diff coverage: 100.00% (126/126 executable changed lines)
```

Diff coverage reported 99.21% on the commit before last, with one
uncovered changed line:
the sentence the rerun says to the reader, which no test drove because
none passed an
`on_progress` into a retrying turn. The last commit pins it rather than
waiving it.

```
make lint-python
  All checks passed! / 1946 files already formatted
make lint-imports
  Contracts: 10 kept, 0 broken.
make lint-deps
  Success! No dependency issues found.
make check-source-language
  clean
make check-large-files
  clean
make check-commits
  clean
```

The kernel budget gate caught one thing worth naming: documenting the
turn-end record in
the hook contract put `raven/contracts` at 2808 lines against a ceiling
of 2791. The
per-field detail moved to the write site in `turn_path`, which is not
budgeted, and the
paper now names the fields and says where the rest is. `raven/contracts`
is at 2788.

Each of the three review fixes was proved against the commit it fixes
rather than
reasoned about: a detached worktree at that commit, the new tests copied
in, red there and
green here. They reproduce the reviewer's own numbers -- the episode
list `[0, 1, 2, 0]`,
the eighteen seconds, and the stranded ledger file. The strengthened
comparison in
`test_the_retry_starts_from_the_question_not_from_the_wreckage` is a
tightened guarantee
rather than a regression test and passes either way, which is why it is
called out here.

`agents/raven-research/.env` was moved aside for every suite run: it is
gitignored and
absent in CI, and two launcher tests that delete a key from the
environment otherwise find
one in the file.

This machine has no search-provider key, so no research brief can
actually be run here.
The retrieval-path changes are verified by unit tests only, which is why
the three
chokepoints are pinned by tests that assert both the presence and the
absence of a row
rather than by an observed run. The same limit applies to the rerun: the
clock and the
dead-end trigger are driven by a deterministic clock in tests, never by
a real turn.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

No user-facing surface changed, so there is nothing to document. The one
operator-visible
knob, `RAVEN_VERBATIM_SINK`, is an instrument that is off unless set and
is documented in
the module that reads it.

## Risk

The `is_hard_tool_failure` change is the only unconditional one, and it
reaches every agent
rather than just research: any tool returning a JSON envelope with a
non-empty top-level
`error` now counts toward the tool-failure streak that triggers the
change-approach nudge.
That is the intended correction, and the direction is conservative
because transient
markers are still tested first. Reverting it is a single branch.

Everything else is inert by default. The body sink writes nothing unless
the environment
names a file. `DigestOutput` has no emitter in this change, and a digest
returning a plain
string produces the same envelope it always did, which
`test_a_bare_string_digest_still_works` pins.

The dead-end predicate now has a caller, and it reaches only an agent
whose hooks write a
budget onto the turn's metadata. No agent on this branch does:
`agents/raven-research`
ships `wallClockSeconds` and inherits the rerun from its plugin's class
default, and every
other agent leaves the key unwritten and is bounded exactly as before.
The turn's overrun
is one budget plus at most one iteration, because the clock is read
between iterations and
never mid-generation: cancelling a call in flight would discard a
finished generation and
leave no answer at all. Rolling the rerun back is the budget key;
rolling the clock back is
the same.

Rollback is the whole commit; nothing here is depended on by anything
already on main.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
Co-authored-by: KT <74288668+0xKT@users.noreply.github.com>
…es (#472)

## Summary

After a conversation switches model from the page, the agent keeps
introducing itself as the model it left. The switch itself works -- the
session record, the request and the usage ledger all carry the new id --
but the system prompt's `You are running on model: ...` line was read
from `agents.defaults.model`, so the agent was told it still runs on the
configured default and repeated that back. From the page it is
indistinguishable from a switch that never happened.

The fix reads the id from the turn's active `ModelBinding` instead. The
loop already opens `use_binding(binding)` around every turn
(`raven/agent/loop/main.py`, around `_run_turn`), and that binding is
the session's own `/model` pick when it has one (`binding_for_session`
in `raven/agent/loop/wiring.py`), else the default.
`render._resolved_model_id` now takes its model id from
`active_binding()` when one is set and from config outside a turn, and
runs either through the same storage-to-wire conversion
(`providers.wire`) it always ran the default through. One source file
changes; nothing is threaded through the assembly path.

This is the approach the review suggested over the earlier revision of
this PR, which carried the id down as a new contract field, and it is
better on every point that revision was faulted for:

- **Right spelling.** The earlier revision forwarded
`session.metadata["model"]`, the storage spelling, and rendered it
as-is; on Codex, Azure and every gateway install that is not the id the
request carries. Here the bound id goes through `wire_model` exactly as
the default does, so `openai-codex/gpt-5.1-codex` renders as
`gpt-5.1-codex` and a gateway install keeps its `openrouter/` prefix, in
both branches. The conversion stays on the one path that owns it, and a
lazily built provider's identity `wire_model_id` is not relied on.
- **The estimation path agrees by construction.**
`ContextBuilder._get_identity` (`raven/agent/context/builder.py`) sizes
the system prompt for the token budget and calls the same renderer with
no arguments; it is invoked from `_assemble_context_messages`, inside
the binding, so it now names the same model the assembled prompt does. A
field handed down the assembly path reached the segment and not this
caller.
- **No contract change.** `AssemblyContext` and `TurnContext` are
untouched, so there is no field position to argue about, no producer
(the curator's `TurnContext` rebuild) that can forget to copy it, and
`CONTRACTS_VERSION` stays where `main` has it.
- **`stable = True` stays true.** A binding is per session, constant
within it; the line changes once per switch, which is what a switch is.
A router's per-message pick is not reflected in the line (the router is
opt-in and off by default); that is the one case a contract field would
have covered, and it is also the case that made the segment's stability
claim false.
- **No empty-string hole.** `ModelBinding` refuses an empty model id at
construction, so the line cannot be deleted by a blank pick.

A fallback hop further down a model chain still reads a prompt naming
the primary: context is assembled once, before the call. Not changed
here and not a regression -- the sibling `can_see_images` field
documents the same caveat.

Nothing under `ui-web/` or `ui-tui/` changes.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Five new tests in
`tests/test_segments.py::TestIdentityNamesTheBoundModel`:

- `test_a_switched_conversation_is_told_the_model_it_is_bound_to` --
inside `use_binding` on a gateway config, `_resolved_model_id()` names
the bound model with the gateway prefix and never the configured
default.
- `test_the_bound_id_goes_through_the_same_storage_to_wire_conversion`
-- a codex-form id bound to the turn renders as the wire spelling the
client sends, not the stored one; this is the assertion the earlier
revision could not fail.
- `test_outside_a_turn_the_configured_default_still_answers` -- no
binding, config default, as before.
- `test_the_segment_names_the_bound_model` -- the whole segment, through
`IdentitySegmentBuilder.build`.
-
`test_the_estimation_prompt_and_the_turn_prompt_agree_on_a_switched_model`
-- `ContextBuilder._get_identity()` equals the segment's text inside a
binding and names the switched model.

Commands and results:

- Mutation-checked: replacing the binding read with the config read
fails exactly the four tests that bind a model (4 failed, 46 passed in
`tests/test_segments.py`); the outside-a-turn test and the three
pre-existing wire-form tests stay green.
- `.venv/bin/python -m pytest` over every test file that names
`identity_text`, `_resolved_model_id`, `IdentitySegmentBuilder` or
`ContextBuilder(`, plus `tests/test_contracts_two_tier_ledger.py`,
`tests/test_kernel_budget.py` and `tests/test_context_stable_prefix.py`
(11 files): 375 passed.
- `python scripts/coverage_gate.py diff --base-ref upstream/main
--threshold 90` after those suites ran under `--cov=raven --cov-branch`:
100% (3/3 executable changed lines).
- `.venv/bin/ruff check`, `.venv/bin/ruff format --check` and
`.venv/bin/ty check` on `raven/context_engine/segments/render.py`:
clean.
- The contracts gates pass unchanged: this revision touches nothing
under `raven/contracts/`.
- No user-facing docs change: the prompt line already documents itself
as naming the model the turn runs on.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

## Risk

- Security: none. The binding's model id is already in every request and
in the usage ledger; nothing new is read or sent.
- Backward compatibility: outside a turn (no binding) the renderer
behaves exactly as before. Inside a turn the line names the bound model
instead of the configured default; for a conversation that never
switched they are the same id, so the rendered prompt is byte-identical
there. No wire, config or contract change.
- User-visible change: after a `/model` switch the agent names the
switched model in its identity line from the next turn on. The identity
prefix changes once per switch, so a prompt-cache prefix keyed on it is
invalidated once per switch, not per turn.
- Rollback: revert the squash commit; one file, nothing else to
reconcile.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

Fixes #470

Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
## Summary

Update the hosted AI4AI (Nanochat 50M Pretraining) chart reference in
both bilingual READMEs.

- Replace the old image attachment URL in README.md and README.zh-CN.md.
- Use the updated chart with the Bits Per Byte (BPB) label and the Lower
Is Better annotation above the BPB values.
- Keep the chart hosted as a GitHub attachment; no image binary is added
to the repository.

## Type

- [ ] Fix
- [ ] Feature
- [x] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

Commands:

- `git diff --check`
- `UV_NO_SYNC=1 make check-large-files check-source-language
PYTHON_VERSION=.venv/bin/python COMMIT_RANGE=HEAD`
- `UV_NO_SYNC=1 uv run --no-sync pytest tests/test_readme_scope_canon.py
tests/test_cli_onboard_commands.py -k readme_quickstart -q`
- Downloaded attachment SHA-256 matches the local PNG.

## Risk

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

This is a documentation-only URL update. Reverting the commit restores
the previous chart.

## Related Issues

#465

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
…ber like the fleet (#493)

## What

Two defects in the shipped `agents/raven-code` profile, both inherited
from the retired vendored twin's comparison hygiene and both silent at
runtime.

**The skill route was advertised but dead.** Trunk's skills are
pull-discovery: a name+description menu rides the user envelope, and its
footer tells the model, verbatim, *"Read one with read_skill(id); search
differently with find_skill"* (`raven/context_engine/scent.py`).
raven-code kept `read_skill`, `use_skill` and `find_skill` in its
disable rows, so the menu taught a route that answered "unknown tool" --
while the router still ran, retrieval still happened, and every body it
fetched was dropped unread, with nothing logged to say so. All three
tools register off the skill registry alone
(`raven/agent/loop/wiring.py`); a Hub endpoint only adds their `hub/`
branch, and `hub` itself stays disabled.

**Memory was double-off.** design, oncall, ppt and research all ship
`memory.backend: "everos"`; raven-code alone shipped backend `null`
**plus** the `everos-memory` plugin opt-out -- the one shipped agent
that could not remember. The identity was already `raven-code` in all
three places (`memory.userId/agentId`, the plugin slice,
`subagent.json`); only the two off-switches go.

The redundant `skillForge` block goes too: its two knobs
(`rewriteEnabled`/`llmGateEnabled`) are only built under push discovery
(`factory.py`), which no shipped agent runs, and the other four agents
carry no such block.

**Review follow-up (second commit):** enabling everos-memory also offers
the plugin's `understand_media` tool through the entry-point lane, which
the hermetic face fixture cannot see -- an unledgered 14th tool in
production. Same remedy as the playbook rows that board past the fixture
the same way: the disable row is the pin, and `PLUGIN_LANE_WITHHELD`
puts the lane on the ledger. raven-research holds the same line; design,
oncall and ppt serve the tool deliberately.

## Fleet posture after this change

| | everos | read/use/find_skill | skillForge block |
|---|---|---|---|
| design / oncall / ppt / research | on | design serves all three | none
|
| **raven-code (before)** | **double-off** | **all three disabled** |
explicit false/false |
| **raven-code (after)** | on | all three served | none |

## Tests

The tool face moves on its ledger, not by luck
(`tests/test_agents_code_launcher.py`): `SKILL_LANE_TOOLS` joins the
face arithmetic, `find_skill` leaves `TRUNK_NEW_WITHHELD` (five stay
withheld), `PLUGIN_LANE_WITHHELD` pins the plugin lane, and the everos
posture test is rewritten from *stays-factory-off* to *ships-on*. The
face itself is hermetically rebuilt from the render: 13 tools = the
previous 10 + the three skill tools, verified by name and schema.

- `tests/test_agents_code_launcher.py`: 70 passed
- reviewer's set (`test_agents_code_launcher` +
`test_agent_loop_skill_tools` + `test_context_scent` +
`test_everos_plugin_discovery`): 110 passed
- ruff check + format: clean

---------

Co-authored-by: litong <238663200+TongLi31@users.noreply.github.com>
## Summary

Add `.github/CODEOWNERS` with a single entry mapping `/raven/agent/` to
`@LivXue`, so a
pull request that changes a file under that directory automatically gets
a review request
from the directory's owner.

This only requests a review. The branch ruleset keeps
`require_code_owner_review` off, so
no approval requirement changes and no contributor's merge path is
affected. Making it a
merge gate would be a separate ruleset change, deliberately not made
here.

Two limits are inherent to the mechanism rather than gaps to close
later:

- GitHub never requests a review from the author of a pull request, so
one opened by the
  owner does not assign them. The API answers that case with HTTP 422.
- A draft pull request gets no automatic request. It is sent when the
draft is marked
  ready for review.

CODEOWNERS is read from the base branch, so this governs pull requests
opened after it
lands and does not apply retroactively. Ten of the sixteen pull requests
open at the time
of writing touch `raven/agent/`. Those were assigned manually, apart
from the one the
owner authored, which cannot take the assignment for the reason given
above.

## Type

- [ ] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [x] Other

## Verification

Gates run in a clean worktree branched from `origin/main`, each exit
code read rather than
inferred from empty output:

```
make check-large-files      exit 0
make check-source-language  exit 0
make check-commits          exit 0
```

Behaviour checked against GitHub rather than by reading the file. The
CODEOWNERS validator
on the pushed branch returns no errors, which is what proves both that
the pattern parses
and that the named owner resolves to an account holding write access:

```
gh api "repos/EverMind-AI/Raven/codeowners/errors?ref=..."
{"errors":[]}
```

Not run, and why: `make lint` and `make test-python` analyse Python,
TypeScript and
JavaScript, and this diff adds one ten-line text file containing no
code, so neither
measures anything about the change. No test in the repository exercises
CODEOWNERS
resolution, which is why the validator above is cited instead of a test.
Separately
confirmed that nothing reads the directory this file lands in as a set:
of the three tests
mentioning `.github/`, one validates backticked documentation paths
carrying a known file
extension, one names a single workflow file, and one globs
`.github/workflows` only.

- [ ] Relevant tests pass locally
- [ ] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

## Risk

Nothing changes for any contributor: an added review request does not
gate a merge, and
CODEOWNERS confers no permission of its own. The owner handle named here
is already public
on every commit and pull request in this repository, so the file
discloses nothing new.

Rollback is deleting the file, or the one entry in it. Nothing depends
on it.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
## Summary

Replace the first image in README.md and README.zh-CN.md with the
updated Raven banner, "One raven, a whole flock of specialists." Both
versions use the same GitHub-hosted attachment.

## Type

- [ ] Fix
- [ ] Feature
- [x] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

- `uv run --frozen --extra dev pytest tests/test_readme_scope_canon.py
-x -q`: 2 passed.
- `uv run --frozen --extra dev pre-commit run --files README.md
README.zh-CN.md`: all applicable hooks passed.
- `git diff --check origin/main...HEAD`: passed.
- `make check-commits check-large-files check-source-language`: passed.

GitHub confirmed the attachment upload. Download timeouts prevented a
complete comparison with the supplied PNG. Browser rendering at desktop
and mobile widths has not been checked.

## Risk

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

This changes the README banner only. The image is hosted as a GitHub
attachment. Reverting the two URL edits restores the previous banner.

## Related Issues

Fixes #126

Co-authored-by: Codex <noreply@openai.com>
#507)

## Summary

Stands up a bilingual MkDocs site under `docs-site/`, published to
GitHub Pages
by a workflow that builds on every pull request and deploys only from
`main`.
Fourteen topics ship in English and Chinese: the quick start,
self-hosting,
Docker, the WebUI, the sandbox manual, the runtime architecture, the
command
reference, the repository layout, the developer workflow, the
proactivity design
and reference pair, the self-evolution map and the tracing API.

This pull request is additive. Both READMEs keep their prose and gain
one link
to the site in the header row; trimming them is deliberately a second
pull
request, so the site can be confirmed live before anything depends on
it.

Key decisions:

- **`docs-site/`, not `docs/`.** `docs/README.md` defines `docs/` as
design
notes and dated records, and `docs/plans/` is an archive explicitly not
held
  to today's layout. A user manual has the opposite lifecycle.
- **Full bilingual parity**, via the `mkdocs-static-i18n` suffix
structure
(`page.md` beside `page.zh.md`). Parity decays immediately without a
guard, so
  a test fails when a page has no twin.
- **Anchors are pinned on the Chinese side.** A CJK heading slugifies to
  nothing, so Markdown numbers it `_1`, `_2`, and inserting one heading
renumbers every anchor below it, silently breaking links readers already
hold.
Each Chinese heading therefore states the anchor its English twin gets
for
  free: 133 heading anchors, identical across the two languages, none
  positional.
- **A left rail, not a top bar.** The header carries the brand, the
search box
and the navigation tree down the left; the table of contents on the
right is
  one continuous polyline clipped to the headings currently on screen.
- **No template overrides.** Material customises through
`theme.custom_dir` with
Jinja partials, those partials are `.html`, and
`scripts/check_large_files.py`
  holds `.html` in `BLOCKED_ASSET_EXTENSIONS`. The skin is therefore one
stylesheet remapping Material's own custom properties, plus one script
for the
table-of-contents rail. The brand mark is a `data:` URI for the same
reason.
- **An owner-signed `docs-site/` exemption zone** in AGENTS.md section
1.3. The
i18n config carries the Chinese navigation labels, and `mkdocs.yml` is
not
`*.md`, so the blanket suffix exemption does not reach it. A contract
test
holds the AGENTS.md list and `EXEMPT_PREFIXES` equal, so both moved
together.
- **The site never links out to a file it does not host.** Documents
that stay
in the repository are named as plain paths rather than linked, and a
guard
fails a page that links to an external webpage or a repository-relative
path.

Two defects found and closed here rather than after publication. The
Chinese
site rendered an English sidebar: navigation labels live in
`mkdocs.yml`, not in
any page, so comparing the two languages' Markdown finds nothing and
only the
rendered output shows it. And a control could be present, sized,
animated and
answer `querySelector` while being unclickable, because a popup Material
sized
for a top bar landed outside the viewport once the header became a rail.

## Type

- [ ] Fix
- [x] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

All commands run at the pushed head, rebased onto `origin/main`.

```
make check-large-files                                         exit 0
make check-source-language                                     exit 0
make docs-build                                                exit 0, 0 errors
  docs-site/site/index.html and docs-site/site/zh/index.html   both present
  mkdocs_static_i18n: Translated 17 navigation elements to zh
uv run pytest tests/test_docs_site.py tests/test_source_language_check.py \
  tests/test_readme_scope_canon.py tests/test_docker_runtime.py \
  tests/test_living_docs.py                                    44 passed
uv run pytest tests/integration/test_docs_site_controls_e2e.py 13 passed
pre-commit run --files <changed files>                         all hooks pass
git diff --check                                               clean
```

Anchor parity was measured on the built HTML rather than assumed: 14
page pairs,
133 heading anchors compared at h2 through h6, every one identical
between the
two languages and none positional.

The controls are hit-tested rather than queried.
`document.elementFromPoint` at
the exact point a reader clicks must return the control itself, which is
the
only check that catches an off-viewport popup; the language switcher,
the search
panel, the repository link and the table-of-contents rail each pass it.

The brand lockup is measured, not eyeballed. A failing test was written
first
and reported the wordmark sitting 4.00px below the mark's centre line;
it is
0.00px after the fix, and a 4x screenshot puts the residual optical
offset at
0.25px.

Every guard this branch adds was confirmed to fail first. The
brand-class guard
was mutated from both sides in turn: renaming `.em-eyebrow` in the
stylesheet
turned it red, renaming `class="em-standfirst"` in `index.md` turned it
red, and
it is green again with both restored. The anchor guard was shown to
catch both
failure modes, a missing anchor and a wrong one.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

## Risk

No behaviour change to any shipped code path. Nothing imports
`docs-site/`, the
new Makefile targets are additive, and the only change outside
`docs-site/` and
its tests is one link added to each README's header row.

Rollback is deleting the branch: the site does not exist until Pages is
enabled
with `Source: GitHub Actions`, and every other change is inert without
it.

Disclosed residue, checked and deliberately not changed:

- The README header now links to `https://evermind-ai.github.io/Raven/`,
which
404s until Pages is enabled. It is the same URL the Documentation
section
  already carries.
- The `deploy` job is skipped on pull requests by design, so a green
check here
does not prove deployment works. The first real evidence is the `deploy`
job's
  own conclusion on the `main` run after merge.
- GitHub Pages is not enabled on this repository. Measured: `GET
/repos/.../pages`
returns 404. Under `Deploy from a branch` the uploaded artifact is
silently
ignored while the workflow still reports success, so the source must be
set to
  `GitHub Actions`.
- `tests/integration/` is excluded by `norecursedirs`, so the browser
suite is a
  local guard and never runs in CI.
- The self-hosting page's Makefile target, docker path and port claims
are
unguarded. They were equally unguarded in the README they came from, and
the
  owner ruled against adding a guard for them.
- The self-evolution SOP and the benchmark contract stay in the
repository and
are named as plain paths from the pages that reference them. The SOP is
a
translation of an upstream project's internal document and carries
redaction
placeholders, which is a publication decision rather than a formatting
one.
- `docs/plans/` still shows the pre-restyle `mkdocs.yml`. CONTEXT.md
defines that
directory as an archive of the tree as it stood, with an explicit rule
against
  updating an old plan to match a later change.
- The Chinese pages drop the brand's italic serif accent word. Cormorant
Garamond
carries no CJK glyphs, so it would fall back to another face without any
sign.
- The code font is Material's default Roboto Mono. The brand names four
  typefaces and no monospace, so none was invented for it.

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
… states (#503)

## Summary

The overview says the four built-in agents are "Powered by the **Raven
Evolver** engine". The architecture section of the same file says
`evolver/`
is a separate tool that consumes Raven as a library and that the runtime
does
not import Evolver, and `pyproject.toml` enforces that with an
import-linter
contract literally named "the runtime does not import the evolver". A
reader
gets two incompatible answers to whether Evolver is a runtime engine or
an
external development tool, and the contract settles which one is true.

This was gloryfromca's blocking review finding on #486. The thread was
resolved when that PR merged, but the text was never changed, so the
contradiction is live on `main` today in both language mirrors.

Both mirrors now put Evolver where it actually runs: refining the shared
harness from outside, against benchmarks, developing the agents rather
than
running inside them. That is the reading the review offered for the case
where
the intended claim was that Evolver was used to develop the agents. The
overview also takes the canonical spelling of harness self-evolution
from the
architecture section, as the review asked.

Scope note: only the Evolver relationship changes. The SOTA sentence in
the
same paragraph is untouched - it was not part of the finding, and what
it
should say is a product call rather than a consistency fix.

## Type

- [ ] Fix
- [ ] Feature
- [x] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Prose only: two lines, one per language mirror, no code and no new
files.

- `git diff --stat` -> `README.md | 2 +-`, `README.zh-CN.md | 2 +-`
- The two gates that apply here are satisfied by construction rather
than by a
command: `check-source-language` exempts `*.md` by suffix (AGENTS.md
1.3),
  and `check-large-files` sees no added or resized file.
- Read back against the two places the claim has to agree with: the
Evolver
  row of the agent table, and the harness self-evolution bullet in the
  architecture section. Both mirrors now match both.

- [ ] Relevant tests pass locally
- [ ] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

## Risk

Documentation wording only. No runtime behaviour, no interface, no
build.
Rollback is a plain revert. The claim removed was the one the
import-linter
contract forbids being true, so nothing downstream depends on it.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A. Lands the finding from #486 that was resolved without a change.

Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
)

## Summary

`failure_class` keys the loop's tool-failure streak on an envelope's
`error`
string alone. That works only while the string comes from the small
fixed
vocabulary its producers are meant to keep, and it is a convention the
module
cannot enforce - the docstring added in #410 names exactly this failure
mode:
one cause counting as many.

Every media tool broke it. `_format_http_error` interpolated up to 400
characters of vendor body (which carries a request id) into `error`; the
unreadable-reference handlers interpolated the path; the four catch-alls
put
`str(e)` there, which spells the host. So a dead endpoint produced a
class per
call, `(tool, failure_class)` never repeated, the streak never reached
`_LOOP_BREAK_THRESHOLD`, and the stop-repeating nudge was unreachable
for
`image_generate`, `text_to_speech` and `video_generate` alike.

The URL gate in front of `web_fetch` had it too, in both copies of the
tool:
most reasons `validate_url_target` composes name the address they
refused, so
a model walking an internal range got one class per address. #410 fixed
the
handlers behind that gate and left the gate itself.

`error` now carries only what a reader could enumerate. The variable
part moves
to `detail`, which keeps it in front of the model and inside
`is_hard_tool_failure`'s transient-marker scan. Where the split gave the
single
call a further key, the batch reducer forwards that too: `hint`, which
names the
proxy a 403 can be routed through, and `url`, which names the reference
a
guarded fetch refused. A batched picture is told what the single call is
told,
which is the property the split has to preserve. Distinct causes stay
distinct,
which is the failure mode this trades against: two HTTP statuses, two
exception
types and the four vendor refusals each keep their own class, and the
tests
assert that direction too.

Found by sweeping for the shape rather than by reading the code: every
`"error":` whose value is not a literal is a candidate, and that grep is
the
cheapest guard against this recurring.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

- `uv run pytest tests/ -q -p no:randomly` -> 23443 passed, 109 skipped
- `make lint-python` -> ruff check and format clean
- `make lint-types` -> all checks passed
- `make lint-imports` -> 10 contracts kept, 0 broken
- `make check-source-language`, `make check-large-files` -> clean
- Revert-to-red, per fix, restoring the source from a backup afterwards:
reverting `media_gen.py` alone turns the four new class tests red and
leaves
`test_two_different_statuses_stay_two_classes` green; reverting either
copy
of `web.py` turns only that copy's gate test red and leaves the two #410
  streak tests green. For the batch fold, each key is pinned on its own:
reverting it to `detail` alone turns both new batch tests red, restoring
`hint` alone leaves the reference test red, restoring `url` alone leaves
the
  403 test red.
- After rebase onto 45203a4: `uv run pytest
tests/test_media_gen_tool.py
  tests/test_security_web_ssrf.py -q -p no:randomly` -> 100 passed;
  `make lint-python`, `make lint-types`, `make check-source-language`,
  `make check-large-files` -> clean.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

## Risk

Tool output shape changes: a caller reading the reason out of `error`
now
finds it in `detail`. The readers in this repo are the model itself and
the
tests updated here; `tests/test_security_web_ssrf.py` is the one
behavioural
assertion that moved, and it still pins that a private address is
refused for
being private. Rollback is a plain revert of this commit - no state, no
migration, no config.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A. Finishes the envelope work started in #410, which fixed the
reader's
handlers and left the gate in front of them and every media tool.

---------

Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
…its (#490)

## Summary

`raven gateway` exited with SIGSEGV (-11, or 139 in a shell) on every
SIGTERM,
so a systemd or docker stop recorded a crash instead of the clean stop
it was.
`raven tui` and `raven agent` had the same defect.

Root cause: `LocalSkillCatalog` auto-starts `SkillFileWatcher`, a daemon
thread
that parks inside the Rust `watch()` of watchfiles. Daemon status does
not make
process exit safe while the thread sits in native code -- CPython runs
`Py_FinalizeEx` under that call. Every Python shutdown step succeeds
first, so
the crash lands after the graceful chain is complete and the real exit
code is
masked. With faulthandler armed all three faults report `<no Python
frame>`,
i.e. no Python code is on the stack.

Every owner now retires it:

- the gateway's shutdown chain, for the generation still bound at exit;
- the TUI's teardown;
- the one-shot agent's teardown;
- `RavenRuntime.dispose`, for a generation retired at a swap.
`build_runtime`
mints a fresh `AgentLoop` per generation, so each swap starts a new
watcher
  and abandons the previous one; without this a reloaded gateway still
segfaults on stop. The generation-organ roster grows the same entry, so
the
  two halves of that contract stay in step;
- the gateway's shutdown again, for a candidate staged by a reload that
the
serving loop never consumed. That generation is not the bound `agent`,
so
  nothing else reaches it;
- the gateway's shutdown once more, for a candidate `take()` handed out
whose
binding a cancelled unbind never finished. `_serve_generations` holds
that
candidate in a local across `await _unbind_generation()`, so a shutdown
cancelling the await unwinds the coroutine and drops the local, leaving
the
generation staged nowhere and bound nowhere. `SwapCoordinator` now keeps
it:
`take()` records the candidate it hands out and `release()` clears it,
which
spans exactly handover to the moment the caller's own binding points at
the
new loop. The serving loop needed no new line, because it already calls
both
  ends of that span.

The replay runner already stopped the same watcher for the same reason;
these
were the other owners. The stop sits in the generation contract and in
each
command's teardown rather than in `AgentLoop.stop`, whose other caller
is the
generation swap -- that path keeps the process running and must keep
skill
auto-refresh with it.

The gateway's sweep now lives in a module-level
`_retire_generation_watchers`
rather than inline in the serve closure. `run()` is a 550-line closure
with no
import seam, so the inline form could only be pinned by reading its
source text
-- which never executes it, and so earned no coverage on the lines that
matter.
The helper drains all three seats and keeps going when one stop raises,
since a
sweep that stops early still leaves a watcher parked in native code.

Two comments describing a removed guard go with it: one in
`agent_commands`
pointing at an exit chokepoint in `raven.cli.commands.run`, and the
test-side
scaffolding that patched `os._exit` for an exit the command no longer
makes.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Process-level, faulthandler armed, throwaway home and a dummy key
nothing calls:

```
                                      before   after
raven gateway   SIGTERM                 -11      0     (-11 on 5 of 5 runs)
raven gateway   SIGHUP then SIGTERM     -11      0     (2 reloads after)
raven tui       PTY, then signal        -11      0
raven agent     ordinary one-shot exit  139      1     (1 is the turn's own
                                                        auth failure, which
                                                        139 had been masking)
```

Controlled pair isolating the cause, on the loop factory the TUI uses --
build
the loop, let the interpreter finalize:

```
watcher left running   exit 139
stop_file_watcher()    exit 0
```

Suites, from the branch tree:

```
uv run pytest tests/integration/test_gateway_sigterm_smoke.py -q -n 0 -m integration
  1 passed          (this is the guard test that was failing)

pytest tests/test_cli_agent_loop_parity.py tests/test_cli_gateway_commands.py \
  tests/test_cli_gateway_health.py tests/test_cli_tui_commands.py \
  tests/test_cli_agent_commands.py tests/test_gateway_spine.py \
  tests/test_generation_swap.py tests/test_plugin_services_lifecycle.py \
  tests/test_rpc_control.py tests/test_rpc_system.py \
  tests/test_skill_watcher_cleanup.py tests/test_trajectory_replay.py \
  tests/test_utils_asyncio_runner.py -q
  307 passed, exit 0

ruff check / ruff format --check over the changed files   clean
ty check over the changed files                           no diagnostics on the changed lines
scripts/check_large_files.py origin/main..HEAD            exit 0
scripts/check_source_language.py origin/main..HEAD        exit 0
commitlint --from origin/main                             exit 0
```

Second round, after review found the in-transition window; rebased onto
`origin/main` at 6d8c363:

```
pytest tests/test_cli_gateway_commands.py tests/test_core_runtime_swap.py \
  tests/test_generation_swap.py tests/test_cli_tui_commands.py \
  tests/test_cli_agent_commands.py -q
  171 passed, exit 0

every test file importing raven.core.runtime          131 passed, exit 0
real gateway: boot, SIGTERM, exit code                0 on 3 of 3, no fault text
ruff check / ruff format --check                      clean
make check-source-language / make check-large-files   exit 0
coverage_gate.py diff --threshold 90                  95.83% (23/24), passed
```

Revert check on that round: removing the `_in_transition` record, the
in-transition seat in the sweep, or the sweep's `try/except` each fails
its own
test and nothing else. The cancellation window itself is reproduced
against real
`SkillFileWatcher` threads, counted by thread name, and is red without
the fix.

Diff coverage on changed lines was 50% (4 of 8) before that round,
because the
gateway-side tests read the chain's source instead of running it. The
one line
still uncovered is the call site inside `run()`, which cannot execute
without
booting a real gateway; the wiring to the seam is pinned by source
inspection.

Each of the four regression tests was confirmed red before its fix. The
TUI and
agent stops were additionally re-confirmed afterwards by removing both
again and
watching those two tests go red, then restoring them.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

## Risk

User-visible change: these three commands now exit with their real
status
instead of a segfault. Anything reading the exit code sees a different
value --
0 where it used to see 139, and for a failed one-shot the turn's own
non-zero
code rather than 139.

Rollback: revert the commits. The watcher is a cache-invalidation
helper, so
removing the stops restores the previous behaviour exactly, crash
included.

Checked and deliberately not changed:

- The stops are not wrapped in try/except. `ContextBuilder.__init__`
assigns
`skills` unconditionally, `stop_file_watcher` is documented safe when no
watcher was started, and the watcher's own contract is that neither
`start`
nor `stop` raises. This matches the existing production caller in the
replay
  runner, which is also unwrapped.
- `raven acp` reaches finalization the same way but calls `os._exit(0)`
after a
real session, so the fault is masked there rather than fixed. Left as
is;
  making it consistent is a separate change.
- No central gate was reinstated. A future command that builds a loop
and exits
normally would reintroduce this, and the source-inspection tests here
only
  guard the retirement seats that exist today.
- The watcher starts in `ContextBuilder.__init__`, which is
construction-time
work, and `RavenRuntime.discard` is contract-bound to stay call-free
because
a call there would mean an organ was started before FREEZE. So a staged
or
in-transition candidate is retired by the gateway's sweep rather than by
  `discard`, and the docstring records why. Moving the watcher out of
  construction would let `discard` own it; that is a separate change.
- The shutdown-before-consumption window was not reproduced by timing:
three
runs sending SIGTERM 0.2s, 1.0s and 2.0s after a reload all completed
the
  swap first and exited 0. The staged-candidate retirement is driven by
  forcing the state in a test rather than by racing for it.

Known gap in the evidence: the two `raven tui` runs were ended with
SIGTERM
after Ctrl+C went unanswered, because the TUI never finished loading
against a
throwaway home. The ordinary-exit case is evidenced by the agent
one-shot,
which exits by itself with no signal involved.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
…ked again (#467)

## Summary

Two regressions in the hand-built deck lanes (medium and high on
Raven-Design), and the stream failure that made every run of them a coin
toss.

**The deck lane draws its own pages again.** PR #451 pointed the
hand-built lane at the engine's helper modules (`ppt_layout`,
`ppt_theme`, `ppt_shapes`), which were written for the max lane's script
backend. Their defaults became the deck: `heading()` put a white surface
band bled across every page, `card()` became the way anything was
placed, text frames were top-anchored with the lower third empty, and a
build script named three colours. Measured on the same request, the
accent colour appeared 8 times across 20 pages. The skill, route note
and assets reference go back to the pre-#451 text: the lane draws with
python-pptx and takes only the icon set, `add_formula` and `math_runs`
from the engine. On the reverted lane the accent appears 54 times, the
deck carries 12 real pictures and the layouts differ page to page.

**Colour is named beside the type sizes.** The skill listed sizes for
every role and said nothing about colour, so a lane with the right
palette still drew in one colour. One paragraph now says what colour is
for; the layouts note no longer reads as one accent per page, which a
lane took as a ceiling. Where `image_search` is not offered, the
fallback no longer tells the lane to give the pictures up: `web_search`
for the page, `web_fetch` it, and download where the fetch backend keeps
image links.

**A stall in a silent think is asked again.** Across two measured runs
the lane's turn failed 28 times on `Upstream idle timeout exceeded`,
every time in the one round that writes the whole build script into a
tool argument. `stream_llm_call` refuses to retry after output a watcher
has seen, and counted reasoning deltas and half-built tool-call
fragments as that output, though neither is ever shown as the reply. So
the configured retry ladder was never reached. `rendered()` now asks
what a watcher could have seen; every retry path resets the buffers
first (a slot left standing merged with the next attempt into a call the
model never made); a stall handed back as an error response carries the
usage it spent. The in-process builtin backend, whose comment claimed a
retry ladder and passed none, now receives the deployment's
`llmErrorRetryDelays` and `llmRetryAfterOutput` through the
SubagentManager. No packaged product takes that path; all five run as
ACP agents on the loop's own call site.

**Two bundled templates edited in place**, by the maintainer's word.
`beige_geometric_general_report` loses its page 20, a dark photograph
cut by a circle under pale rounded cards; 23 pages become 22 and the two
reference pages the lane borrows from it (17, 19) keep their numbers.
`black_circuit_tech_launch` loses the 27 percent circuit texture on its
master and the photograph on its Section Header layout, so its body and
section pages sit on the master's plain `#000000`; cover and closing are
unchanged. `templates.manifest.json` is repinned; `fetch_templates.py
--verify` holds all ten pins. The registry archive was not re-cut.

No configuration default changes. The only behaviour every lane sees: a
stream that fails before any reply text reached a watcher is retried on
the configured ladder instead of failing the turn.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

```
uv run --frozen pytest tests/test_agent_loop_llm_error_patience.py tests/test_agent_loop_stream.py \
  tests/test_subagent_manager.py tests/test_subagent_routing_backend.py -q
265 passed in 6.85s   (after rebase onto main)

uv run --frozen pytest tests/test_litellm_provider_stream.py tests/test_litellm_provider_stream_end.py \
  tests/test_provider_stream_fallback.py tests/test_providers_stream_idle_budget.py -q
102 passed

uv run --frozen ruff check raven/agent/loop/main.py raven/agent/subagent/backends/raven_loop.py \
  raven/agent/subagent/manager.py raven/providers/streaming.py
All checks passed!

grep -nP "[^\x00-\x7F]" agents/raven-design/route-raven-ppt.md   -> no matches (the note is pinned ASCII)
```

Both new tests were checked against the unfixed code:
`test_a_stall_during_a_silent_think_is_asked_again` fails with the old
`emitted` test, and
`test_what_the_failed_attempt_left_behind_is_dropped` fails with the
buffer resets removed.

The failure itself was reproduced outside the runtime: the same
37-message history, six pictures and tool table, sent through litellm
1.101.0 to the same vendor, succeeds on a short answer every time and
fails mid-stream one time in three when asked to write the full build
script; healthy long streams show a maximum inter-chunk gap of 1.9s over
234s, so the cut is a dropped connection, not a slow model.

Decks built on the changed lane, same request as the #451 acceptance run
(glm-5.3-flash via OpenRouter): medium 21 pages with 12 pictures and 7
text colours, 48 gate findings (4 `excessive_whitespace`, down from 18);
high 13 pages with 13 pictures and 9 text colours, 37 findings. The high
run's page count came from the host agent's forwarded brief, which
stated 10-14 pages the user never asked for; that is a separate issue.

Template edits: `tests/test_ppt_engine_templates.py
tests/test_ppt_engine_default_templates.py
tests/test_ppt_engine_template_bands.py
tests/test_ppt_engine_template_theme.py
tests/test_ppt_engine_measure_geometry.py` -> 152 passed, 1 skipped;
`make check-large-files` -> clean; both templates rendered and read page
by page.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

## Risk

- [ ] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

User-visible: on every lane, a streamed call that fails during reasoning
or mid tool-call, with no reply text yet shown, waits out the configured
ladder (default 15/30/60s) and asks again instead of ending the turn. A
call that already streamed reply text behaves as before. Reasoning
tokens of the abandoned attempt are spent again and are now recorded in
usage. Rollback: revert the `providers` commit alone to restore the old
failure; the skill text and the builtin-backend plumbing are independent
of it.

## Related Issues

N/A
## Summary

Space on an attempt row in the `raven trajectory` browser previously
showed two truncated per-turn preview lines. This PR replaces it with an
interactive full-conversation viewer. It re-lands the content of PR
#382,
which was auto-closed by the 2026-09-12 history linearization without
being merged; the branch is rebased onto the linearized main, with one
extra commit expressing the CJK display-width test fixtures as unicode
escapes to satisfy the source-language gate that main gained since.

- Data layer `raven/trajectory/conversation.py`: rebuilds an attempt's
  span snapshot into a causally ordered stream of labeled records (User
  input, LLM input/output, Tool input/output, skill, memory, and
  subagent events). Event-time semantics (inputs at span start, outputs
  at span end) plus nesting depth from parentSpanId keep the causal
  order. Full content comes from artifact files (paths restricted to
  the trace store, streamed 512 KiB cap per file); span previews are
  only a visible fallback. ERROR spans and expected-but-unreadable
  payloads always yield a placeholder record. LLM inputs omit only the
  item-by-item verified common message prefix against the previous call
  (omissions and rewrites are announced); tool_calls of unexecuted
  tools keep their full arguments.
- Rendering: kind-colored labels aligned to one column, hanging-indent
  bodies with word-aware CJK-safe wrapping, degraded/error/meta note
  lines, single-pass turn separators (interleaved traces are never
  regrouped), and a stacked layout on narrow terminals that never drops
  characters.
- Interactive viewer (prompt_toolkit full-screen, entered for every
  preview on a terminal): scrolling, a muted bottom bar with an
  'h for help' hint, filter/collapsed state and the scroll range (laid
  out by display width, never overflowing any terminal), a scrollable
  help page reachable in full on small screens, a global collapse (s:
  bodies over five display lines fold to five plus an ellipsis; error
  and degradation notes never fold), and a label filter (%: real input
  line, case-insensitive, empty clears; filtered-out turn groups lose
  their separators). Resizes re-wrap and re-judge folding live; the
  content caret is hidden (only the filter input shows one). q/Esc
  returns to the list; Ctrl+C cancels the whole browser from any viewer
  state. A rebuild failure falls back to the legacy per-turn preview;
  without a terminal the preview degrades to direct printing.
- Every dynamic field passes the sanitization gates (control
  characters, rich markup, and ANSI never reach the terminal). Adds the
  Conversation Record term to CONTEXT.md.

## Type

- [ ] Fix
- [x] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

- `uv run pytest tests/test_trajectory_conversation.py
  tests/test_cli_preview_viewer.py tests/test_cli_trajectory_browse.py
  tests/test_trajectory_store.py -q` -> 298 passed (viewer assertions
  run against really rendered screens)
- `uv run pytest -q` (full, on the rebased branch) -> 23463 passed,
  125 skipped; 33 failed + 20 errors reproduce with identical node ids
  on a clean origin/main worktree - environment baseline, not
  introduced here (passed-count delta of 96 equals the tests this
  branch adds)
- `make check-source-language`, `make check-large-files`,
  `make lint-types`, `make lint-python`, `make lint-imports` -> all pass
- Manual: rendered 6 real local attempts at widths 72/120/30 with zero
  over-width lines (largest attempt 7060 lines in 0.35s); drove the
  real TUI in a pty at normal and 44x12 sizes (viewer, help paging,
  collapse, label filtering, Ctrl+C cancel)

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

## Risk

- Read/display side only; no persisted format changes. The untrusted
  input surface (artifact content, typed filter input) is contained by
  the store-prefix path check, the 512 KiB streamed cap, and per-field
  sanitization.
- User-visible change: the Space preview is now a full-screen viewer;
  the legacy rendering remains as the automatic fallback when the
  rebuild fails, and non-terminal environments print directly.
- Rollback: revert this PR, or restore the previous `_preview_screen`
  call site to get the old preview back.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-fable-5) <noreply@anthropic.com>
…e shell (#516)

## Summary

The documentation site built one search index for both languages, so a
reader on
the English pages was offered Chinese ones: results they cannot read,
behind
links that leave the language they chose. A build hook now splits the
merged
index in two and points the Chinese pages at their own copy, which the
theme
also resolves every result link against.

Where that split runs is the whole fix. The merge lands in a plugin's
own
`on_post_build`, so the split has to be the last thing that event does.
Running
it on `on_shutdown` looks equivalent and is not: `mkdocs serve` never
shuts
down, so the preview went on serving the merged index however often it
rebuilt.
The new guard makes the same assertion against `mkdocs build` and
against
`mkdocs serve`, because passing one says nothing about the other.

The rest settles the shell against the reference layout:

- the rail's foot is its own box the navigation cannot show through, and
the tab
  icon is the EverMind mark;
- the footer keeps to the content column instead of covering the fixed
columns
  once the reader reaches the page foot;
- the contents column reaches the screen foot with its scrollbar hidden
behind a
  fade, and follows the reader by centring the first lit entry;
- search opens as a centred dialog over the page. The form itself
travels into
that dialog, so what stays in the rail is drawn from the field's own
metrics
and its shape does not change on a click. Neither state animates: the
theme
transitions the two shapes into each other, and leaving search played
that out
  in the corner of the page after the reader had moved on.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

- `uv run pytest tests/integration/test_docs_site_search_e2e.py` -> 2
passed.
Without this change the preview case fails: the Chinese index is HTTP
404 and
  the English index carries 146 Chinese pages.
- `uv run pytest tests/integration/test_docs_site_controls_e2e.py` -> 25
passed
- `uv run pytest tests/test_docs_site.py
tests/test_source_language_check.py` ->
  31 passed
- `make check-commits`, `make check-source-language`, `make
check-large-files` ->
  exit 0; `ruff check` on the added files -> no findings
- `make docs-build` -> the root index holds 149 entries and no Chinese
page; the
  Chinese index holds 146, none of them addressed from the site root

Driven in a real browser at 1560x820 in both languages: a word that
appears in
both trees returns 10 results and 0 from the other language, and the
corner of
the page is pixel-identical to its never-opened state at every frame
after
leaving search. Every page of the built site was checked at the page
foot (28
pages, 0 faults) for a footer covering a column, for a column whose last
entry
cannot be reached, and for the rail's foot icons being clickable.

No page content changed, so nothing else needed updating with it.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

## Risk

User-visible: search no longer returns pages from the other language,
and the
search control, the footer and the two fixed columns change shape. The
field
left in the rail while the dialog is open is a drawn replica, and every
rule
behind it sits above the theme's desktop breakpoint, so narrow viewports
keep
the theme's own full-screen search untouched.

Rollback: revert the commit. The hook is one file, and dropping it from
the
`hooks:` list in `docs-site/mkdocs.yml` restores the previous merged
index on
its own.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
)

## Summary

An EverOS role points a section at a provider: the settings card writes
the
model and asks `lend_provider_credentials` for that provider's key and
address to go beside it. The address was copied only when there was one
to
copy -- the reasoning being that a section with no lender address may be
holding one the reader typed by hand.

A section being pointed at a new lender is holding the OLD lender's
address,
though, and leaving the field out of the answer keeps it. Moving the
memory
role from an OpenRouter model to a DeepSeek one left DeepSeek's key at
OpenRouter's address, which the far end refuses. DeepSeek is not a
corner
case here: every vendor LiteLLM routes carries no `default_api_base`, so
the
lend answers with no address for all of them.

The lender answers about the address either way now, empty included. An
empty
value reads where it lands the way an unset one does: the EverOS writer
sets
the field, and the section stops carrying somebody else's endpoint.

Found while running the settings page's acceptance case for the EverOS
roles
against a real gateway. The lend rule is from #414, not from the
settings
work that surfaced it.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

- `uv run pytest tests/test_config_update_providers.py
tests/test_rpc_settings.py
tests/test_cli_onboard_commands.py tests/test_rpc_console.py -q` -- 641
passed
  on the branch this was written on; 191 on the two closest suites here.
- The new case fails without the change: the lender answers `{"api_key":
...}`
  where the section needs `{"api_key": ..., "base_url": ""}`.
- `uv run ruff check` and `ruff format --check` over both files --
clean.
- Real gateway, temp `RAVEN_HOME`, EverOS plugin present:
`settings.everosSet`
  with `borrow_from: "deepseek"` writes the model and DeepSeek's key and
leaves no `base_url` behind; the same call with `borrow_from:
"openrouter"`
writes OpenRouter's address back. Before the change the first of those
left
  `base_url = "https://openrouter.ai/api/v1"` under DeepSeek's key.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

## Risk

- Two callers, both of which point a section at a provider: the EverOS
card
(`settings.everosSet`) and the CLI onboarding reuse step. For a lender
that
  has an address, nothing changes. For one that has none, the borrowing
  section now stops carrying the previous lender's -- which is the fix.
- A reader who typed an address by hand into an EverOS section and then
borrowed from an addressless provider loses that typed value. That is
the
same event as the bug, read the other way; the section is being pointed
  somewhere else, and the card shows the resolved address.
- Rollback: revert the commit.

## Related Issues

N/A

---------

Co-authored-by: gloryfromca <23442919+gloryfromca@users.noreply.github.com>
Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
LivXue and others added 20 commits September 20, 2026 19:16
…ction (#557)

## Summary

Open the Showcase with a Raven-Code run that picked, built and judged an
FPS
game over 42 rounds, and give README.zh-CN.md the Showcase section it
never had.

The entry sits in its own table above the existing ones: the left cell
carries
the run's task graph and a 30s clip of the result, and the right cell is
left
empty for a second sample. It is a separate table rather than a row in
the deck
grid, because a playable game is not a deck and an empty cell inside
that grid
would read as a broken row.

The clip is a GitHub attachment referenced by a bare URL on its own
line. That
is the only form that renders a player: a video element is removed by
the
markdown sanitizer, which was confirmed by rendering both forms through
the
markdown API before writing either one. The same rendering also shows
that the
player header displays the name the file carried at upload, so the clip
was
re-uploaded under its intended name rather than left under a working
name.

README.zh-CN.md previously had eight sections against the English nine,
and the
missing one was Showcase. Rather than hand-copying it, the section was
derived
from the English one and its entry titles translated, so the two files
stay
structurally aligned; alt text is left in English, matching every other
alt
attribute in that file. Both files now render the same structure: one
player,
seventeen images, fourteen cells, one empty cell.

The section's opening sentence is reworded in both files while the
section is
being touched: it now says Raven generates the orchestration rather than
chooses it, and the Chinese copy names the panel as the multi-agent
orchestration graph.

Checked and deliberately not changed:

- The task-graph alt text describes three nodes running "across 42 game
rounds".
The panel actually shows round 42 of a capped loop, with those three
nodes
being that one round's pick, do and judge steps, so the wording is
loose. The
author chose to keep it; it is recorded here rather than silently
altered.
- The first upload of the clip is still attached to its hosting issue
under the
earlier filename. Nothing references it, and attachments are not
removable
  once posted, so it stays as upload history.

## Type

- [ ] Fix
- [ ] Feature
- [x] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Rendering, through GitHub's own markdown API rather than by eye, on the
exact
section text of each file:

    gh api -X POST /markdown --input -   # mode gfm, repo context

Both files returned one video element, seventeen images, fourteen cells
and one
empty cell, with the player header reading the intended filename and the
first
entry title reading as intended in each language.

Byte equality of every uploaded asset against its local source, fetched
back
over HTTP:

    curl -sL <asset-url> | sha256sum

All matched. A HEAD request is not usable here: the attachment host
answers HEAD
with 403, which reads as a failed upload.

Tests that actually read these two files, found by grepping for them
rather than
assumed absent:

uv run --frozen --all-extras pytest tests/test_cli_onboard_commands.py
tests/test_docker_runtime.py -q
    344 passed

Repository gates:

    make check-commits check-large-files check-source-language

All passed. The two file gates were also run against a bare revision,
which
compares the working tree, because a commit-range form reports nothing
while an
edit is still uncommitted.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

No test asserts anything about Showcase content, asset URLs or section
parity
between the two READMEs. The 344 above cover one heading lookup and one
whole-file read; nothing would fail if an entry were malformed. The
rendering
check above is what stands in for that, and it is manual.

## Risk

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

Documentation only; nothing executable changes. The assets are public
attachments on a repository that is already public and carry no
credential or
private detail. The clip is 53 MB, so the Showcase section is heavier to
load
than before, and the player appears only in the rendered file view: the
diff and
compare views show the bare URL as text. Rollback is a revert of these
commits;
the assets stay hosted either way.

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
## Summary

- Align English and Chinese documentation for Docker/Compose, sandbox
execution, and the in-tree tracing API with verified runtime behavior.
- Separate proactivity usage guidance from developer design and
implementation details, with configuration, troubleshooting, cost, and
safety boundaries.
- Update navigation and remove redundant section separators. Keep page
URLs unchanged.
- Preserve published proactivity section and title anchors in both
languages. Use native aliases for same-page content and visible
migration links for moved sections, with regression tests.
- Clarify Sentinel delivery fallback and menu-selection requirements,
tracing-on/off argument handling, headless CLI menu output, and viewer
descriptor lookup. Add focused regression coverage without changing
runtime behavior.
- Clarify the initial frozen baseline and improve Evolver wording in the
bilingual site and repository mapping. Keep its structure, navigation,
implementation terms, and technical claims unchanged; broader Evolver
documentation work is excluded.
- No runtime changes.

## Type

- [ ] Fix
- [ ] Feature
- [x] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Passed on the rebased revision:

```bash
uv run --frozen --no-sync pytest tests/test_docs_site.py tests/test_tracing_api.py tests/test_sandbox_unit.py tests/test_sentinel_runner.py tests/test_sentinel_planner.py tests/test_sentinel_fast_path.py tests/test_nudge_policy.py tests/test_proactive_spawn.py tests/test_core_sentinel_stack.py tests/test_core_cron_stack_ledger.py tests/test_cli_sentinel_commands.py tests/test_decision_router.py -q -n 0 -x
# 395 passed

uv run --frozen --no-sync --with 'mkdocs-material>=9.5' --with 'mkdocs-static-i18n>=1.0' --with 'mkdocs-git-revision-date-localized-plugin>=1.6.0' --with 'jieba>=0.42.1' mkdocs build -f docs-site/mkdocs.yml --strict --site-dir /tmp/raven-pr551-push.9ZfOEi/site
# Strict English and Chinese builds passed

UV_NO_SYNC=1 make lint-python PYTHON_LINT_TARGETS='tests/test_docs_site.py tests/test_tracing_api.py tests/test_decision_router.py tests/test_cli_sentinel_commands.py'
UV_NO_SYNC=1 make check-large-files check-source-language COMMIT_RANGE=origin/main..HEAD
PYTHONPATH=. uv run --frozen --no-sync python scripts/check_commit_messages.py origin/main..HEAD
npm exec --yes --package=@commitlint/cli@21.1.0 -- commitlint --from origin/main --to HEAD --config commitlint.config.cjs
git diff --check origin/main...HEAD
# All passed
```

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

Documentation checks also covered bilingual examples and the proactivity
JSON configuration schema. A built-HTML audit verified all 110
historical proactivity section fragments and four title fragments,
including same-language migration destinations, and found zero broken
internal fragment links across 29 HTML pages (672 links). Replaying
pre-fix content made all four page compatibility cases fail as expected.
The new descriptor-directory test failed on all three documents before
their stale paths were corrected. Added coverage exercises tracing
on/off, bare-number versus explicit menu selection with incomplete
classifier configuration, and discovery-menu dispatch through the real
CLI stack's headless sink. Runtime type checks and the full repository
test suite were not run; no runtime code changes. No browser click test,
live LLM, real-VM, or container end-to-end run was performed.

## Risk

Documentation-only, with regression tests. Proactivity navigation
changes, but page URLs and historical heading fragments are preserved.
An old link to content moved between pages lands on a visible migration
entry and requires one click to reach the new section; no JavaScript
redirect is needed. Existing queued-nudge policy and external-agent
isolation limitations are described, not fixed. Roll back by reverting
the documentation and associated test changes; no data migration is
involved.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

---------

Co-authored-by: zhao.wang <270284818+userName20260323@users.noreply.github.com>
Co-authored-by: Codex <noreply@openai.com>
## Summary

A 26 MB deck dropped on the composer was refused. The upload ceiling
stood at 25 MB, a number it never chose: it was the file viewer's,
adopted when `fs.upload` landed beside it. The viewer picked 25 MB
because no renderer in the page does anything useful past that size,
which is not a reason that applies to an upload.

An upload answers a different question. It only has to land a path in
the workspace -- a non-image attachment is named to the model and never
read into the message -- so the bytes exist to reach the disk, not a
renderer. This raises `MAX_UPLOAD_BYTES` to 100 MB and leaves
`MAX_VIEW_BYTES` at 25 MB. A file between the two can therefore be
attached and not previewed; the constant's comment now states that
instead of implying the two move together.

`frame_ceiling_for_upload` derives the WebSocket frame ceiling from this
constant, so the base64 expansion that once capped attachments at
roughly 3 MB cannot come back from raising the number. The page's
mirrored copy moves with it.

The ceiling that remains is memory, not policy: one upload at this size
costs the gateway the frame, the JSON string parsed out of it and the
decoded bytes at once, and the parse holds the event loop while it runs.
A much larger limit wants the streamed HTTP upload path, not a bigger
number here. That is left for its own change.

This also sweeps four comments that stated 25 MB as the current ceiling;
they name the constant now. The three describing the old 3 MB frame bug
keep their number, because they report what was true then.

Deliberately out of scope: the composer and the subagent page measure an
attachment after base64-encoding it rather than before, so an oversized
file is read whole and thrown away while the tab blocks.
`uploadRefusalBySize` already exists for the pre-encode check and the
knowledge page uses it; wiring the other two callers crosses the
composer source contract, so it belongs in its own change.

## Type

- [ ] Fix
- [x] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Server side, end to end at the size that was refused (26.2 MB), through
the real aiohttp WebSocket and the real method:

- a frame carrying that file as base64 comes back answered, not
disconnected
- `fs_upload` accepts it and the bytes land in `uploads/` at the right
size
- `MAX_UPLOAD_BYTES + 1` is still refused, with `100 MB` in the message

Those three ran as a throwaway probe and are not added to the suite. The
committed guard is
`test_the_frame_ceiling_can_carry_the_largest_upload_the_method_accepts`,
which reads the constant rather than a literal and so keeps the two
limits coupled at any value.

Gates run locally, after rebasing onto github/main at 82a51d8:

uv run pytest tests/test_rpc_transport.py tests/test_rpc_files.py
tests/test_rpc_console.py -q
      172 passed
make lint-python All checks passed / 1625 files already formatted
    make lint-imports    clean
    make lint-deps       Success! No dependency issues found.
    make lint-types      All checks passed!
npm run gen:check --prefix ui-web generated.ts matches the contract (178
methods)
    npm run type-check --prefix ui-web   clean
    npm test --prefix ui-web             108 files, 1837 tests passed
    node ui-web/scripts/check-page.mjs             OK
    node ui-web/scripts/count-shared-globals.mjs   OK
    node ui-web/scripts/check-css.mjs              OK
    node ui-web/scripts/check-class-namespace.mjs  OK

Page side, in a real browser against a `raven serve` built from this
branch, on an isolated RAVEN_HOME:

- a 26.1 MB .pptx dropped on the composer attaches: the chip goes solid
at "26.1 MB" with no failure note, and the file lands in
`<workspace>/uploads/` at exactly 27,400,000 bytes with its non-ASCII
name intact
- a 105 MB file is refused, and the note names the new ceiling ("105.0
MB", "100.0 MB"), so the limit was raised rather than removed and the
message follows the constant

The drop was a synthesized DragEvent: it exercises the page's own drop
handler and the whole upload path, but not the operating system's part
of the drag.

No user-facing docs name this ceiling. The one doc mention of 25 MB is
the viewer's render cap in
`docs/specs/2026-08-26-desk-deliverables-shelf.md`, which this change
does not touch.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

## Risk

Behaviour change: an attachment up to 100 MB is now accepted where 25 MB
was the wall. Per upload the gateway holds the frame, the parsed JSON
string and the decoded bytes at once -- roughly 370 MB peak for a
maximal one -- and the JSON parse blocks the event loop for a second or
more. Single-user desktop use absorbs that; a shared gateway taking
concurrent uploads would feel it.

Security: the route is unchanged. `fs.upload` is reachable only over an
authorized socket, so the larger allocation is available to the
signed-in local reader and to nobody new. What grew is how much memory
that reader can ask for in one call.

A file between 25 MB and 100 MB can be attached and not opened in the
viewer, which still refuses past `MAX_VIEW_BYTES`.

Rollback: revert the two constants. Nothing persists, no schema or
contract moved, and `frame_ceiling_for_upload` recomputes from whatever
the constant says.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

Co-authored-by: arelchan <204152633+arelchan@users.noreply.github.com>
Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
)

## Summary

Rebuilds the Showcase layout in both READMEs, and replaces one image.

**Why the old layout did not hold.** The section had grown to nine cases
across three tables of mixed shape. A heading that wrapped to two lines
pushed its column's images down. Paired task graphs of unequal height
left
the artifacts below them starting at different heights. And GitHub's
stylesheet stripes every second row while a case spanned three, so the
pattern drifted until some headings sat on white and others on the grey.

**What it is now.** Each case spans three rows - heading, task graph,
artifact - so cells in a row share a top edge and a long heading cannot
cascade into the images. Each pair gets its own table, which restarts
the
stripe and puts the usual gap between pairs. The game run keeps the
opening
slot and now takes the full width instead of leaving half a row empty.
Image widths are set per pair so both sides render to one height, and
the
narrower image is centred. The divider sentence that separated the old
tables is gone, since the headings already say it.

**A new comparison board.** The Frameworks artifact is replaced with a
redrawn plate: each framework carries its vendor mark, and the
orchestration style became a three-way tag - explicit graph or canvas,
declared task flow, dynamic at runtime - where the old plate split the
six
in two. Its alt text said "versus" and named two categories; it names
three
now. The render is 2000x1547 against the old 2000x1332, so its width is
set
to 86 percent to stay level with the sweep chart beside it.

## Type

- [ ] Fix
- [ ] Feature
- [x] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

- `make check-large-files` -> exit 0
- `make check-source-language` -> exit 0
- The replacement board was fetched anonymously from its attachment URL:
HTTP 200, 2000x1547, and its md5 matched the local source file byte for
  byte.
- Rendered the section through GitHub's own markdown endpoint under the
published markdown stylesheet and read back the computed row
backgrounds.
  All five heading rows now report one background; before the split they
  reported two.
- Measured every image box in that render. All four pairs match on top
edge,
  and within 2px on height:

  | pair | left | right |
  | --- | --- | --- |
  | Song / Greece | 379x129 and 379x718 | 379x129 and 379x718 |
  | Pop / Abstract | 379x127 and 379x713 | 371x127 and 379x713 |
  | Frameworks / Sweep | 379x126 and 326x252 | 379x126 and 379x252 |
  | Poster / Explainer | 314x131 and 379x348 | 379x131 and 367x347 |

- Confirmed against the repository-rendered README that the game run's
bare
  attachment URL is served as a video player, and kept it on its own
  paragraph inside the cell so it still is.

- [ ] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

## Risk

Documentation only, in both READMEs. One of the seventeen attachments is
replaced, with its alt text rewritten to match the new plate's three-way
legend; the rest are untouched, and no file is committed to the
repository.
Rollback is reverting the commit.

- [ ] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

Follows #547 and #557.

Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
…rounds directory (#553)

A campaign asked the shared rounds directory how much had been spent,
and the
directory answered for every job in it, including jobs its sibling
campaigns
had submitted. Rounds that keep one directory therefore bill each other,
and
the bill is cumulative: round three opens already charged for rounds one
and
two.

Measured on a run of 2026-09-11. Five campaigns declared the same rounds
directory. Round two had a budget of 130 and had actually used 34.6, but
the
eighteen jobs the earlier rounds had left in the directory brought the
reading
to 125.6, so the remaining 4.4 would not cover one more trial at 5.9 and
the
round stopped. 142 GPU-minutes went unspent and that round's multi-seed
conclusion was lost. A sibling run that gave each round its own
directory was
not affected, which is why this reads as a directory-hygiene problem
until you
look at the meter.

Two changes. Spend is now counted from the campaign's own ledger, so a
job a
sibling submitted is not billed here; the filter sits above the existing
timeline arithmetic, so same-device overlap is still charged once and
two
cards in parallel are still charged twice. And declare refuses a rounds
directory a live sibling campaign already keeps, naming the sibling, so
the
two rounds do not land in one directory in the first place.

Type: fix

Verification:
- All oncall tests: 775 passed, 1 failed. The one failure
  (test_the_products_tool_face_equals_the_forks_config_intent, a missing
optional a2a_client dependency on this machine) also fails on 0464aa1,
so
  this adds six passing tests and no new failures.
- The six new ones pin: a sibling's job is not billed; the campaign's
own two
trials overlapping on one device are still billed 20 minutes once; a job
this executor just submitted is billed before the ledger has caught up
(no
refund window); billing_only leaves a foreign backend untouched; declare
  refuses a directory a live sibling keeps; a concluded sibling, and the
  campaign itself, do not count as a clash.
- The first test also pins the measured number: the same fixture
directory
reads 125.6 unfiltered, which is 34.6 of its own plus 91.0 of its
siblings.

Risk: low for the meter, moderate for declare. The meter can only ever
bill
less than before, and a run that already gives each round its own
directory
reads the same. Declare now refuses something it used to accept, so a
caller
that deliberately shares a directory between live campaigns will see an
error
naming the sibling; concluded siblings and re-declaring the same
campaign are
both still fine.

---------

Co-authored-by: xiaotian.luo <xiaotian.luo@thetahealth.ai>
Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
…rail groups by it (#542)

## Summary

A conversation can now be started in a folder of the person's choosing,
from the page, the way a code editor's folder switcher works: on the
new-task screen the composer carries a folder chip that opens No folder
/ Recent / Open folder..., and the folder picked rides on
`session.create` as its `workdir`. The engine already pinned a session's
turns to that directory (shell, filesystem, deliver and media tools
follow the binding); the page had no way to set it.

The choice is made once. After the first message the chip only reports
the conversation's folder and cannot be pressed, which is exactly the
engine's semantics: the pin is taken at creation and every later turn
reads it.

The rail splits the rest of the list on the same fact: conversations
pinned to a folder sit under their own collapsible heading, each row
wearing the folder's name (full path on hover); conversations on the
policy default keep the old heading, renamed to say what it is once the
other group exists. Pinned and cron rows are untouched.

Two small backend additions carry it:

- `session.list` items gain an optional `workdir`, the stored override
or absent, which is what the rail groups on.
- `fs.dirs` lists the subdirectories of one absolute directory (home
when omitted), directories only, dotfiles omitted, capped at 500, and
marks each entry -- and the listed directory itself -- with whether
`session.create` would accept it (`validate_override`: not the agent
home, not its ancestors, not its memory/skills/sessions trees). The
picker greys those out instead of offering a folder the create would
refuse.

"Recent" is derived from the rail's own rows, newest activity first, so
nothing new is stored and every browser sees the same list. The offline
demo page installs neither hook; the chip works and "Open folder..."
says it needs a gateway.

## Type

- [ ] Fix
- [x] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

- `uv run pytest tests/test_rpc_console.py tests/test_rpc_session.py
tests/test_rpc_contract_shapes.py tests/test_rpc_schema_match.py
tests/test_rpc_registration.py tests/test_ui_language_repaint.py`: 661
passed. New: six `fs.dirs` tests (subdirectories only, dotfiles omitted,
sort, `ok` on entries and on the listing, home default, root has no
parent, relative/file/missing/unreadable refused, a child vanishing
mid-listing), one `session.list` test (the override or null), and the
`fs.dirs` round trip against its declared result model.
- `scripts/coverage_gate.py diff --base-ref upstream/main`: 97.83% of
changed executable lines (45/46).
- `ruff check` and `ruff format --check`: clean.
- ui-web: `npm run gen:check` (179 methods, in sync), `npm run
type-check`, `npx vitest run`: 109 files, 1849 tests passed. New:
`shell/workdir.test.ts` (draft/locked states, recents deduped, stage and
unstage, the browse walk with greyed entries and a disabled use button,
the offline and refused-path notes), three rail tests (the two groups,
the tag on either separator, folding), two `open-conversation` tests
(the staged folder rides the create and lands on the row; none staged
creates with none).
- `python ui-web/build.py`, `check-page.mjs`, `check-css.mjs`,
`check-class-namespace.mjs`: OK.
- ui-tui: `npm run gen:rpc`, `npm run gen:i18n`, `npm run lint:rpc`,
`npm run lint:i18n`, `npm run type-check`: in sync and clean.
- On a real gateway, in the served page: opened the picker on the
new-task screen, browsed home into a demo folder, picked it; the chip
named the folder; promoted the draft; the chip locked with the
fixed-once-started note and no longer opened; `fs.list` for the new
session rooted at the demo folder; a turn asking the agent to run `pwd`
answered with that folder; `session.list` carried it as `workdir`; the
rail showed the row under the folder group with its tag, and the other
rows under the no-folder heading.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

## Risk

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

`fs.dirs` lists directory names on the gateway's machine to an
authenticated page. The page already lists and reads files under a
session's working directory, and the agent's shell tool reaches the
whole machine, so this widens what the page can see by directory names
only; dotfiles stay hidden and the agent's own data is marked unusable,
matching `session.create`'s refusal. A page that never presses the chip
behaves exactly as before: `session.create` is called with no `workdir`,
and a `session.list` consumer that ignores the new optional field sees
the old shape. Reverting the commit restores the previous page and
contract; sessions already pinned keep working, since the pin was the
engine's feature already.

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
## Summary

Two commits. The first fills the Launch WebUI section of both READMEs,
which
carried three `Screenshot placeholder` lines and no images, with two
figures
each: the new task page and the subagents page. The second lands the
Raven mark
those figures show.

On the figures:

- Two, not three. The third placeholder asked for a task graph, which
cannot be
captured from stored data: the web DAG panel is driven by live `dag.*`
events
and its disk restore is gated on a localStorage note that only a live
event
writes, so showing one would have required a real `run_subagent_dag`
turn.
That slot is dropped rather than filled with something its caption does
not
  match.
- They come from an isolated instance with its own `RAVEN_HOME` and an
empty
session store, so no conversation titles, memory entries or workspace
file
names appear. The A2A block was removed from its config, and the config
carries no channels and no cron, so the throwaway instance had no
outward
side effects. Its home, which held a copy of real provider keys, was
deleted
  after the run.
- Each README gets its own language, captured with `language` set to
`en` and
to `zh`. Following the convention already in these files, `alt` text
stays
  English in both while the caption follows the prose language.

On the mark:

- The new file is a dark bird on an antique-gold halo. Against the dark
theme's
surface that ink measures 1.07:1, and no tone filter can lift it because
it is
not one colour to invert, so it carries an opaque plate rather than the
`prefers-color-scheme` rule the previous file used. Ink against plate
measures
  18.29:1 on both themes.
- The plate is a card, so the mark fills its frame the way
`miromind.svg`
already does. Keyed by file rather than by `data-agent`, which carries
the
  preset and raven's own rows have none.
- The halo stroke goes from 1.6 to 3.2. Against a 64 unit canvas the old
width
lands under one device pixel at the two smaller slots the mark renders
in,
0.45 at 18px and 0.65 at 26px, so the ring washed out wherever the
roster
  draws it.

Reviewer note: the figures were captured before the stroke went to 3.2,
so the
halo in them is thinner than what this branch now renders. Everything
else in
them is current. Re-capturing would mean a second upload for a
difference of
half a device pixel, so they are left as they are.

Assets are hosted as a comment on issue #505 rather than committed, per
the
repository asset rule.

## Type

- [ ] Fix
- [ ] Feature
- [x] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

```
.venv/bin/python -m pytest tests/test_ui_agent_marks.py -q -p no:randomly
    -> 7 passed

.venv/bin/python -m pytest tests/test_scope_canon.py tests/test_docs_site.py \
    tests/test_docker_runtime.py -q -p no:randomly
    -> 27 passed

npx --prefix ui-web vitest run --root ui-web \
    scripts/agent-mark-css.test.mjs src/shell/agent-mark.test.tsx
    -> 2 files, 18 passed

make check-large-files        -> exit 0
make check-source-language    -> exit 0
make check-commits            -> exit 0
```

The gates were run over a committed range rather than an empty diff, and
the
mark suite was re-run afterwards because `make check-commits` can
rebuild the
virtualenv under it.

The mark change was rendered and measured rather than eyeballed: at
1440px the
mark is 28px, and the two candidate files differ only in a sub-pixel
halo and a
plate that is near the tile colour in the light theme, so the two are
not
visually separable at full size. The renders were compared by pixel
signature
against standalone renders of each file.

Every uploaded asset was fetched back with a GET and compared by sha256
against
the local render; all four are byte identical. A HEAD request is not a
valid
check here, because user-attachments URLs answer 403 to HEAD and 200 to
GET.

Both README sections were rendered through the GitHub markdown API to
confirm
the figures resolve, the alt text survives and each file references the
assets
for its own language.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

## Risk

Two user-visible changes. Both READMEs render two figures in Launch
WebUI where
three placeholder lines used to be. The Raven mark in the subagents
roster and
in every other slot it draws in is a different image, larger in its
frame, and
no longer answers the theme from inside its own file.

Rollback is a revert of either commit independently; they touch disjoint
files.
The hosted assets can stay either way.

Backward compatibility is not applicable: no interface, config key or
stored
value changes.

- [x] Security impact considered
- [ ] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

#505

---------

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
Restore the missing icon mappings for the overview cards in both locales.
## Summary

The four Launch WebUI figures were flat rectangles butted against the
page.
They now carry a 24px corner radius and a soft drop shadow, so each
reads as a
window floating on the page rather than a slab of pixels pasted into it.

- The capture is untouched. The same 1440x900 content sits inside a
1588x1048
frame; the extra 74px on each side is the room the blur needs, and
nothing in
  the interface was re-shot.
- The canvas is RGBA. Both the rounded corners and the shadow are
transparent,
so GitHub composites them correctly on the light theme and on the dark
one.
On the dark theme a black shadow against `#0d1117` is nearly invisible
by
construction, and only the rounding reads; a tinted shadow would fix
that and
  would look dirty in the light theme, so the shadow stays neutral.
- `width` stays at 90 percent. Because the padding is inside the image,
the
interface itself now renders about nine percent smaller than before.
That is
  the space the shadow needs to read as depth, so it is kept rather than
  compensated for.
- Both language pairs are replaced: English in `README.md`, Chinese in
  `README.zh-CN.md`.

The corner mask is drawn at 4x and resampled down, because PIL's
`rounded_rectangle` is aliased at 1x and leaves a visible stair-step on
a
1440px edge.

Shadow parameters, for anyone reproducing them: radius 24, Gaussian blur
34,
offset 16 down, 30 percent black.

Assets are hosted as a comment on issue #505 rather than committed, per
the
repository asset rule.

## Type

- [ ] Fix
- [ ] Feature
- [x] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

```
.venv/bin/python -m pytest tests/test_scope_canon.py tests/test_docs_site.py \
    tests/test_docker_runtime.py -q -p no:randomly
    -> 27 passed

make check-large-files        -> exit 0
make check-source-language    -> exit 0
make check-commits            -> exit 0
```

The gates were run over a committed range, not an empty diff.

Each uploaded asset was fetched back with a GET and compared by sha256
against
the local render; all four are byte identical. A HEAD request is not a
valid
check here, because user-attachments URLs answer 403 to HEAD and 200 to
GET.

Both README sections were rendered through the GitHub markdown API,
which
confirms each file now references its own language's new asset and that
the alt
text survived. The rounding and the shadow were inspected at 1:1 on a
white and
on a `#0d1117` backdrop before upload, to confirm the corner is clean
and the
edge carries no light fringe on the dark theme.

The diff is four URLs, each appearing twice as the link target and the
image
source. Alt text, captions, width and layout are unchanged, and no stale
asset
id survives in either file.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

## Risk

Presentation only. Both READMEs render the same four captures with
rounded
corners, a shadow, and slightly smaller interface content. No code, no
build
output, no behaviour change.

Rollback is a revert of the single commit; the previous assets are still
live
on issue #505, so the old figures would come back intact.

Backward compatibility is not applicable: no interface, config key or
stored
value changes.

- [x] Security impact considered
- [ ] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

#505

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
…#583)

## Summary

Refreshes the Raven-Design visual design figure in both READMEs with the
resupplied
Opus-5 evaluation. Raven-Design and Claude Code moved on all three
benchmarks in that
figure; the GPT5.6-Luna rows and the Hermes column carry over unchanged.

| Benchmark (Opus-5) | Raven-Design | Claude Code | Hermes |
| --- | --- | --- | --- |
| ArtifactsBench Dashboard | 64.4 -> 66.3 | 63.6 -> 62.8 | 64.4
(unchanged) |
| ArtifactsBench SVG | 83.5 -> 86.2 | 82.7 -> 83.4 | 83.0 (unchanged) |
| GDPVal | 89.4 -> 91.9 | 88.7 -> 89.2 | 89.1 (unchanged) |

Two consequences of the update are worth naming, because they change
what the figure
shows rather than only the bar lengths. Raven-Design now leads Hermes on
all three
benchmarks, where ArtifactsBench Dashboard was previously a tie at 64.4.
And Claude Code
now leads Hermes on ArtifactsBench SVG and GDPVal while trailing it on
ArtifactsBench
Dashboard, having previously trailed on all three.

The chart was re-rendered with the same renderer, the same palette and
the same layout as
the figure it replaces, so only the six Opus-5 bars and their value
labels differ. The
asset is hosted as a comment attachment on #477, per the repository rule
that image files
are not committed. Both READMEs reference the same asset, so the diff is
one URL swap per
file; the alt text, the anchor, the width and the caption are untouched.

## Type

- [ ] Fix
- [ ] Feature
- [x] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Chart data, before rendering:

- Built the eighteen expected plotted values by hand from the supplied
numbers, then
asserted the renderer's readout equal to them. Diffed against the
previous value set:
exactly six cells differ, all of them Opus-5 Raven-Design or Claude
Code.
- The renderer's own guards passed: `18 values verified; no overlapping
labels`, plus its
assertions that no text escapes the 2000px canvas and that the output is
2000px wide on
  a white background.

Uploaded asset, after rendering:

- `curl -sL <asset-url> | sha256sum` against the local render:
identical, 141368 bytes,
`e175bc24e7d6208ab26d2aaaab213431585398c0de9705f058dfd4092023924f`.
Checked with GET,
  not HEAD, because user-attachments answers HEAD with 403.
- `POST /markdown` on the changed line returns both `href` and `src`
pointing at the new
  asset.

Repository state:

- Asserted the retired asset id occurred exactly twice per README
(anchor plus image)
before replacing, and zero times after. `git grep` over the whole
tracked tree finds no
  remaining reference to it.
- `git grep` for the six retired score values across every markdown,
text and YAML file:
  no hits, so no prose quoted the numbers the figure carries.
- `uv run --frozen --all-extras pytest tests/test_living_docs.py
tests/test_docs_site.py`:
18 passed. `tests/test_cli_onboard_commands.py -k readme`: 1 passed.
These are the tests
  that read the repository-root READMEs.
- `make check-large-files`, `make check-source-language`,
  `scripts/check_commit_messages.py origin/main..HEAD`: all exit 0.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

## Risk

User-visible change is the Raven-Design visual design figure in the
English and Chinese
READMEs. No code, no configuration and no dependency is touched, so
there is no runtime
behaviour change and nothing to be backward incompatible with.

Rollback is a one-line revert per README: the retired asset
`061b5818-595b-4392-b9b6-53a1effb8347` stays hosted on #477 and keeps
serving, because
replacing a README reference does not delete the attachment it pointed
at.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

#477

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
)

## Summary

`raven skill block` could leave a config file that Raven then refuses to
load.

`set_skill_blocked` wrote `data.setdefault("skillForge", {})`
unconditionally. Dict keys
match byte for byte, but the schema treats `skillForge` and
`skill_forge` as one field
(`alias_generator=to_camel` with `populate_by_name`), so a config file
spelling the block
`skill_forge` gained a second block instead of having its own one
patched. `RavenConfig`
forbids extra inputs, so `load_raven_config` binds the alias and rejects
the leftover field
name: `raven agent`, `raven gateway` and the TUI RPC server all then
refuse to start.

`raven status` keeps working throughout, because the base loader pops
extension keys before
validating. That asymmetry is what makes the defect quiet: the command
most people reach for
to check a config is the one that cannot see it.

The same misread has two softer faces, both closed by the same change.
`raven skill unblock`
became a silent no-op on such a config, reporting success while the
skill stayed blocked;
and blocking an already-listed skill reported it as newly blocked. Both
read the blocklist
out of the empty block the call had just created.

The fix probes the spelling the file already carries and reuses it. That
is not a new shape:
`loader.py` already does exactly this for the `skillRouter` and
`everosSkillLight`
migrations, and `set_sentinel_nudge_quota` does it for
`sentinel.nudge_policy`. Among the
top-level blocks a config writer creates, `skillForge` is the only one
whose camelCase and
snake_case spellings differ at all.

The second commit closes what the first one opened. Making the written
key conditional made
two lines that report the write false on a snake_case file: the log
line, and the result
line of `raven skill block` and `raven skill unblock`, both of which
named
`skillForge.blocklist` unconditionally. Measured on a `skill_forge`
config, the write landed
on `skill_forge.blocklist` while both messages still said
`skillForge.blocklist`, sending an
operator to grep their own config for a key it does not carry. They now
report the setting
rather than the file key, which is what `set_sentinel_nudge_quota`
already does: it respects
the same two spellings and logs `sentinel nudge quota patched` with no
key path at all. The
`--help` text and the module docstring keep naming
`skillForge.blocklist`, because they
describe the command rather than the write.

Checked and deliberately not fixed, with reasons:

- A config already corrupted by the old code does not heal. Running
`raven skill block`
again now selects the camelCase block and leaves the snake_case one in
place, so such a
file still needs a hand edit. Merging the two blocks automatically is
lossy and ambiguous,
and belongs in a migration argued on its own terms rather than inside a
fix whose job is
  to stop producing the state.
- `update_cron_config` writes `to_camel(key)` unconditionally, which is
the same shape one
level down, on field keys rather than block keys. Its failure is milder
and different:
`CronConfig` does not forbid extras, so a file carrying both
`default_timezone` and
`defaultTimezone` loads with the camelCase one winning. Measured
directly. Nothing breaks,
but a later hand edit of the snake_case key is silently dead.
Pre-existing, a different
  failure mode, and in another config surface.
- The three operator-blocklist refusals in
`raven/agent/tools/skill_hub.py` name
`skillForge.blocklist` in text the model reads. These were already
inaccurate for a
hand-written `skill_forge` config before this branch, since the loader
has always accepted
both spellings on the read side. Not introduced here, and fixing them
widens the diff into
  the agent tool surface.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Run on Linux with Python 3.12.13, on a branch rebased onto the current
`main` (which now
carries the A2A work). The rebase conflicted only in
`tests/test_config_update.py`, where
both sides appended a section at the end of the file; both sections were
kept. The rebased
commit's diffstat is byte for byte the pre-rebase one, +70/-1 over two
files, so the
resolution neither absorbed nor dropped anything.

- `uv run --frozen --all-extras pytest tests/test_config_update.py` ->
47 passed
- `uv run --frozen --all-extras pytest tests/ -k "config or skill_forge
or skill_block"` ->
  1709 passed, 33 skipped, 0 failed
- `make lint-python` -> ruff check clean, 2008 files already formatted
- `make lint-imports` -> 10 contracts kept, 0 broken
- `make lint-deps` -> no dependency issues
- `make lint-types` -> all checks passed
- `make check-commits`, `make check-large-files`, `make
check-source-language` -> exit 0

Neither commit's tests are theatre, and each was checked on its own.
Reverting only the
first commit's production hunk turns 4 of its 5 tests red: duplicate
key,
`load_raven_config` ValidationError, unblock no-op, and blocklist read
from the wrong block.
The fifth covers camelCase behaviour that predates the fix and correctly
stays green.
Restoring the message wording the second commit replaced turns its new
CLI regression red on
the `skillForge` assertion, and restoring the fix turns it green.

End to end with an isolated `RAVEN_HOME`: a config carrying the
duplicate block makes
`raven agent` exit with `ValidationError: skill_forge Extra inputs are
not permitted`. The
same starting config driven through the fixed `raven skill block` gets
past config
validation and reaches the expected missing-API-key error, with the
user's existing
`auto_install` setting preserved beside the new blocklist.

Malformed blocks were checked too. `skillForge: null` beside a
`skill_forge` table used to
raise `AttributeError` and now works. `skillForge: null` alone, and
`skillForge` set to a
string, still raise `AttributeError`, unchanged from before.

File mode was measured rather than assumed, because the new A2A
onboarding narrows
`config.json` to 0600 and this branch writes to the same file: after
`initialize_a2a_server`, a `raven skill block` leaves the file at 0600,
since
`atomic_update` preserves the mode it finds.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

`make lint` also runs `lint-ui`, which fails in a fresh worktree with
`ERR_MODULE_NOT_FOUND`
because `ui-web/node_modules` is not installed there. This change
touches no JS or TS file,
so the Python lint targets above are the gates that cover it. No
user-facing docs were
needed; the helper's docstring states the casing contract.

## Risk

Behaviour changes only for a config file that spells the block
`skill_forge`. For a file
using the camelCase spelling, which is what onboarding writes, the bytes
written are
identical to before.

Two changes are visible to a snake_case user, and both are the point of
the fix: the
blocklist is written into their existing section rather than into a new
one, and
`raven skill unblock` now actually removes an entry instead of reporting
success and doing
nothing.

One change is visible to every user regardless of spelling: the result
line of
`raven skill block` and `raven skill unblock` now reads `Skill blocklist
= [...]` where it
read `skillForge.blocklist = [...]`, and the corresponding log line
drops the key path the
same way. This is wording only; no value or exit code changes, and the
`--help` text still
names `skillForge.blocklist`.

Rollback is a revert of the two commits. No migration runs, so nothing
has to be undone on
disk. Reverting only the second leaves the first in place with the
messages inaccurate on a
snake_case config, so revert both or neither.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

No permissions, credentials or file modes are changed; the 0600
measurement above is the
check behind that sentence.

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
…527)

## Summary

`restrict_to_workspace` is an operator's promise that the agent cannot
reach outside
its workspace. It kept that promise only for paths spelled literally.

`ExecTool._check_workspace_restriction` extracted candidate paths from
the command as
typed, then called `os.path.expandvars` on what it had extracted.
Extraction wants a
`/` sitting on a word boundary, and `$HOME/` puts a letter there, so a
variable
spelling produced no candidate at all and the expansion never saw the
text that hid
the path. The expansion sat downstream of the extraction it needed to
feed.

An unset name is the sharpest spelling: the shell drops it, so
`$NOPE/etc/shadow` is
`/etc/shadow`, and there is no literal form to fall back on. Nothing
else was looking
either, because `cat` is read-only and `default_tier` answers `allow`
for a read-only
command without asking; for a read the fence was the only thing in the
way.

Measured with the fence on, before this branch and after it:

| command | before | after |
|---|---|---|
| `cat /etc/shadow` | refused | refused |
| `cat $HOME/.ssh/id_rsa` | ran | refused |
| `cat "$HOME/.ssh/id_rsa"` | ran | refused |
| `cat $NOPE/etc/shadow` | ran | refused |
| `cat ${NOPE:-/etc/shadow}` | ran | refused |
| `cat ${UNSET:-$HOME/secret}` | ran | refused |
| `cat ${PWD%/*}/outside.txt` | ran | refused |
| `cat ${PWD%$PWD}/etc/passwd` | ran | refused |
| `cd /; cat etc/shadow` | ran | refused |
| `cd -- /; cat etc/passwd` | ran | refused |
| `sh -c 'cat $HOME/secret'` | ran | refused |
| `env sh -c 'cat $HOME/secret'` | ran | refused |
| `command cd /; cat etc/shadow` | ran | refused |
| `cat $PWD/notes.txt` | ran | ran |
| `echo 'keys go in $HOME/.ssh'` | ran | ran |
| `cd subdir; cd ..; ls` | ran | refused |
| `cd missing; cd ..; ls` | ran | refused |
| `cd subdir \|\| cd ..; ls` | ran | refused |
| `cd subdir \| cat; cd ..; ls` | ran | refused |
| `cd subdir & cd ..; ls` | ran | refused |
| `pushd ..; ls` | ran | refused |
| `test -d subdir && cd subdir && echo r; cd ..; ls` | ran | refused |
| `cd subdir && cd .. && ls` | ran | ran |
| `cd subdir && cd ..; cat notes.txt` | ran | ran |

The fix expands first, with the shell's own rules, then scans. Each rule
is taken
from the shell rather than invented:

- an unknown name expands to nothing, as the shell does;
- the environment that decides is the child's allowlisted baseline, not
`os.environ`,
  so a name the command will never be given reads as unset;
- single quotes expand nothing, so `echo 'keys go in $HOME/.ssh'` stays
a sentence;
- an escaped `$` yields the bare character, which this level does not
expand but a
  nested shell handed it does;
- a brace body goes back through the same pass, so a parameter inside a
fallback word
  or a trim pattern is read on the same terms as one anywhere else;
- `%NAME%` is read only where cmd.exe is the shell, and there without
regard to case,
  because that is what cmd does and `%VAR%` means nothing to `sh`;
- `PWD` resolves to the directory the command runs in, because a shell
sets it from
  there whatever this process inherited.

A nested shell's payload is scanned in the shell that runs it,
recursively, through
the policy's own embedded-shell helper and depth bound. Both readers of
a segment
unwrap `env`, `sudo`, `command` and leading assignments first, through
the policy's
own helper, so the fence and the deny list agree on what a segment runs
rather than
keeping two lists.

Stepping out is the same escape as reaching out, so a `cd` out of the
workspace is
refused as well. Only the destinations the scan cannot see needed this:
`/` has
nothing after it, `..` is not absolute, and a `cd` with no argument
names `$HOME` by
saying nothing. `cd /etc` was already refused, because `/etc` is a path
like any
other. Each `cd` starts from where the last one landed, but only where a
separator
proves it got there: `&&` runs its right side because the left one
returned zero.
After anything else the move may not have happened, so both readings are
held and a
later step that leaves from either is refused.

### The posture

Bypasses of one family kept appearing: everything the fence did not
model was
allowed, so every gap was silent and the next one would be too. The
contract is now
stated on `_check_workspace_restriction`:

**A parameter expansion is resolved faithfully or the command is
refused.** The set
of spellings is finite, so the residue is closed. Substitution, case
folding,
indirection, offsets and a brace this cannot parse each refuse with the
construct
named. A refusal is visible and can be argued with; the failure it
replaces could
not be.

**Command substitution is a declared limit, not a gap.** It holds an
arbitrary
program, so no textual guard can resolve it, and refusing it would
refuse
`echo "built at $(date)"` along with everything else. Paths written
literally inside
one are still scanned. Two constructs are exempt from the refusal for a
reason that
does not depend on modelling them: `${#NAME}` and `$((...))` yield
numbers, and a
number cannot name an absolute path.

The same rule decides what the directory walk may carry. A `cd` the
shell may not
have run cannot be carried, and the separator is what says whether it
ran: a bracket
or a pipe puts it in a subshell, `||` runs it only when the one before
failed, `;`
continues whether it failed or not. Only an `&&` chain is knowable, and
the chain is
the unit rather than the step -- it stops at its first failure, so what
follows one
inherits the position after any prefix of it. The walk carries those
prefixes and
unions them where the chain ends.

The directory walk took three rounds, and the second and third findings
were both in
code written to close the one before. Recorded because the shape
repeats: the first
version advanced the walk on every `cd`, so a `cd` into a directory that
does not
exist left the walk a level deeper than the shell and the `cd ..` after
it read as a
return. The second read the separator after each `cd`, which closed that
and four
more spellings of it (`||`, a pipe, a background `&`, a bracket) but
treated `&&` as
proof the `cd` ran -- and a condition earlier in the same chain can skip
it. The third
makes the chain the unit: it stops at its first failure, so what follows
inherits the
position after any prefix of it.

Every one of these was measured against bash rather than reasoned about,
in a
workspace with and without the directory the command names, because
which branch
leaks depends on that. `pushd` was found the same way and is the one
finding here
that came from neither review round.

Widening the fence from "no outside path is named" to "and the command
does not walk
out" is the one deliberate contract change beyond the posture. Without
it the rest is
bypassed by `cd /etc; cat shadow`.

Checked and deliberately not changed:

- **The permission tier.** A read-only command is tiered `allow` and
asks nobody. That
is why this defect mattered, but it is a separate control that applies
to every
install including the ones with the fence off, so changing it is not in
scope for
  making the fence keep its own promise.
- **A symlink inside the workspace pointing outside.** `cat link` names
no absolute
path, so there is nothing for a textual scan to refuse. Both sides of
the comparison
already resolve physically, so a named path cannot be laundered through
a symlink.
- **A path computed at run time**, such as one decoded from base64 in a
substitution.
The sandbox executor is the boundary for both of these, which is what
the code says
  about itself.

Blast radius: `_check_workspace_restriction` returns before any of this
when
`restrict_to_workspace` is off, which is the default, so an install that
did not raise
the fence runs none of the new code.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Branch scoped with three dots against the rebased base:
`git rev-list --left-right --count origin/main...HEAD` reports `0 14`.

- `COLUMNS=200 uv run --frozen --python 3.12 --all-extras pytest -q`
2 failed, 23826 passed, 112 skipped in 636.80s under coverage. A second
run of
  the same head gave 1 failed, 23803 passed.
- The one failure is
`test_install_script.py::test_resolve_node_dir_answers_each_case_it_exists_for`
("node without npm provisions a runtime"). It is not introduced here: it
fails
identically with this change removed, and it fails the same way when its
file is run
  alone, so it is neither this branch's nor an ordering artifact.
- A second failure is intermittent and its attribution is left open
rather than
claimed:
`test_agents_research_tools.py::test_a_thin_dataset_card_falls_back_to_the_metadata_the_table_needs`
(`KeyError: 'served_url'`). It appeared in two of three full-suite runs
of this
branch and not in the one full-suite run taken with this change removed,
which is a
single baseline sample and not enough to rule the branch out. Against
that: the test
passes when its file runs alone both with and without this change, and
it passes
when its file and the new one run together in a single process, so the
new tests do
not reach it directly. The mechanism that does fit is in that file: it
selects a
process-wide `ContextVar` 29 times and never restores the token, so its
state leaks
to whichever tests share its xdist worker, and the default `--dist load`
decides
that per run rather than by a fixed split. Adding tests changes that
distribution
without changing any code the test touches. Not changed here, since the
isolation
  defect is that file's own.
- `uv run --frozen --python 3.12 --all-extras pytest
tests/test_shell_workspace_fence.py -q`
43 tests, 116 cases, all passing. Every case added for a reported bypass
was watched
failing before its fix; the rest are over-blocking guards, which pass on
both sides
  by design and are there to fail if a fix reaches too far.
- Reverting only the first commit's hunk in
`raven/agent/tools/shell.py`, with the new
tests left in place, turned 13 of the then 38 cases red: every case
written for the
defect, and none of the over-blocking guards. The tree was restored
afterwards and
`git diff HEAD` confirmed empty, since a three-way reverse-apply stages
what it
  applies.
- `make coverage-diff` (the gate that failed on the first revision):
  `Diff coverage: 97.89% (186/190 executable changed lines)`.
- Recursion through brace bodies terminates because a body is strictly
shorter than
the text holding it and a value is substituted rather than re-read;
checked at 200
levels of nesting and 500 parameters in one body, both answer without a
  `RecursionError`.
- Fence verdict against a real bash, fourteen shapes, each run in a
workspace with and
without `subdir`: `holes: 0 over-blocks: 0`. A hole is
allowed-and-leaks; an
over-block is refused-while-neither-branch-leaks, which is why both
fixture states
  are needed to claim either.
- The walk's separator rule is held from both sides. Never ending an
`&&` chain turns
10 cases red, every escape class; ending it at every separator turns 6
red, the
proven walks that must keep running. Dropping `pushd`/`popd` recognition
turns 3
  red; recognising `popd` without its worst-case reset turns 1.
- `pytest tests/test_*shell*.py tests/test_permissions*.py
tests/test_*exec*.py
  tests/test_sandbox*.py` -- 958 passed.
- `ruff check` and `ruff format --check` clean on both changed files.
- `scripts/check_commit_messages.py`, `scripts/check_large_files.py` and
  `scripts/check_source_language.py` over `origin/main..HEAD`: all pass.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

The docs box is unchecked because nothing documents the fence's path
semantics. The
only mention outside the code lists `restrictToWorkspace` as a config
field, which
this does not change; there was no published claim describing the old
behaviour.

Not reproduced here: the end-to-end refusal on Windows. A Windows path
is not absolute
to a POSIX `Path`, so a verdict test would turn on the host running the
suite rather
than on the expansion. The percent-expansion tests assert the expanded
text and drive
the Windows branch through a class attribute instead.

## Risk

With the fence on, commands that name an outside path through a
variable, that `cd`
out of the workspace, or that use a parameter expansion the fence cannot
resolve are
now refused where they previously ran. The first two are the control
doing what its
name says. The third is the posture, and it is the one that can refuse a
command that
would have been harmless: `echo ${NAME^^}` names no path and is refused
because
nothing here can prove it does not.

That trade is deliberate. An over-block is visible to the operator and
can be argued
with or reported; an under-block is silent, and eight of them were found
across two
review rounds on this branch before it was ready.

One shape that was refused mid-branch now runs: a `cd` sequence chained
with `&&`
that returns to the workspace, such as `cd subdir && cd .. && ls`. The
`;` spelling
of the same walk is refused, and correctly: when the directory does not
exist the
`cd` fails, the `;` carries on from where the shell stood, and the `cd
..` leaves the
workspace. Measured against bash, it reads the parent.

Nothing changes for a default install, where the fence is off and this
code does not
run.

Over-blocking was the failure mode guarded against throughout, and one
instance of it
was introduced and then removed within this branch: reading `%VAR%` on
POSIX refused
`echo '%HOME%/notes'`, which `sh` prints literally. Workspace-relative
`$PWD` use,
single-quoted text containing a `$`, `cd` inside the workspace and back,
`cd` into an
operator-added extra directory, `cd -- subdir`, a wrapper in front of
inside work, a
nested payload that stays inside, a fallback that resolves inside, a
trim that
resolves to a basename or matches nothing, `${#NAME}`, `$((...))`, `date
+%Y%m%d`,
`printf '%s%s'` and `cut -d/` style delimiter arguments all still run,
each with a
test.

Rollback is reverting these commits; the fence returns to its previous
behaviour with
no migration and no persisted state involved.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
…544)

## Summary

At teardown the store pipeline counted every unsettled record as lost,
including one
a worker had already handed to the memory service. Cancelling that
worker does not
cancel its request, so the service finished the write anyway, and the
host told the
user the turn was gone while it was being indexed.

Measured on three one-shot CLI runs against a live EverOS: the client
gave up at its
budget and `atomic_facts_extracted` landed for the same session 32 to 48
seconds
later, every time. The message was wrong twice over -- the turn was
written, and the
service was not unavailable.

Review found that the first version of this fix replaced that error with
its mirror.
The mark is set on entering `MemoryBackend.store()`, which is not a
handoff: a
backend that persists only after an await, cancelled mid-await, leaves
`DrainOutcome(lost=0, in_flight=1)` with the backend's persistence list
empty, and
the notice then said the turn had reached long-term memory. A slow
connection
cancelled before its request body goes out has the same shape.
`MemoryBackend`
declares no delivery guarantee, so no evidence from the transport is
available
without extending the protocol and asking every backend to implement it.
This branch
takes the other route and preserves the uncertainty instead.

What changed:

- `StorePipeline.drain` returns `DrainOutcome(lost, in_flight)` instead
of one count.
A record is marked in flight around the backend call, and the mark is
read before
the collect await that unwinds the cancellation, because that unwinding
clears
  exactly what is being counted.
- `AgentLoop.drain_backend_stores` returns the outcome and logs the two
apart: a
warning for a turn this side watched fail, and for an unsettled one a
line saying
the drain stopped waiting and cannot see whether the service finished
it.
- `report_dropped_memory_writes` becomes `report_memory_write_outcome`,
since it no
longer reports only what was dropped. An unsettled turn reads "were
still being
written when shutdown stopped waiting; whether the memory service
finished them is
  not known here" -- the outcome, not a verdict on it.
- `DrainOutcome` is published on the memory engine's face beside
`StorePipeline`, so
callers reach it the way `test_memory_engine_face` requires. Its
docstring now
carries both cases, with the EverOS measurement scoped to the one it
measured.
- A test drives the `raven agent` teardown that renders the outcome. It
had no
coverage at all: the shared harness stubs `maybe_build_memory_backend`
to `None`,
  so the branch guarding those two lines never ran.

`in_flight` keeps its name deliberately. This repository already uses
the word for
started and not settled -- `swap_in_flight`, `outbound.in_flight`, a DAG
node in
flight -- with no delivery sense, so renaming it would detach the field
from a
vocabulary that is already right. The delivery claim lived in the prose,
and that is
what changed.

Deliberately unchanged: the 2.0s drain budget. It is far shorter than
the write it
waits on -- one store measured 16.98s end to end, of which the flush leg
was 12.1s --
and raising it to cover that would park every CLI exit for as long while
buying
nothing, since the service completes a request it received either way
and nothing
needs to recall the turn that just ended. That reasoning sits beside the
constant,
because the gap between the two numbers reads like a defect and the next
reader would
otherwise close it.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Run on this branch's head, Python 3.12.13, all extras synced, rebased
onto the
current default branch.

- `pytest tests/test_agent_loop_backend_dispatch.py
tests/test_providers_factory.py
  tests/test_cli_agent_commands.py tests/test_memory_engine_face.py
  tests/test_cli_gateway_commands.py` -- 171 passed, 0 failed.
- Full suite: CI on this head runs the suite in four shards: 23742
  passed, 117 skipped, 0 failed. A local full run reports one failure,

`tests/test_install_script.py::test_resolve_node_dir_answers_each_case_it_exists_for`,
which was run again in a clean tree extracted from the default branch at
the
commit this branch is based on and fails there with the same assertion
and the
same value; the CI runner, which is not root, passes it. 0 introduced
either way. The one failure is

`tests/test_install_script.py::test_resolve_node_dir_answers_each_case_it_exists_for`.
It was run again in a clean tree extracted from the default branch at
the commit
this branch is based on, and fails there with the same assertion and the
same
  value. 0 introduced.
- Diff coverage against the default branch: `Diff coverage: 93.94%
(31/33 executable changed lines)`,
measured by the `coverage gates` job on this head. The ratchet moved
with it:
  line +1.81pp, branch +2.74pp. The remaining uncovered pair is
the same teardown inside `gateway_commands.run`, a 634-line function
whose harness
  would cost far more than the one in `agent_commands`.
- `make lint-python`, `make lint-imports`, `make lint-types` -- all exit
0.
- `make check-commits`, `make check-large-files`, `make
check-source-language` --
  all exit 0.
- The rebase was checked with `git range-diff`: every pre-existing
commit replays
`=`, and none of the commits the base gained touches a file this branch
touches, so
  the changed-line set the coverage gate measures is unchanged.

Neither round's tests are theatre, and each was checked on its own:

- First round: the tests were watched failing on the missing attribute
and the
missing function. Two mutants were then run against them -- moving the
in-flight
read below the collect await reddens two of them, which is the
constraint the
comment beside that read now names; moving it below `cancel()` alone
leaves them
  green, correctly, because `cancel()` only schedules.
- Second round: the notice test was watched failing on the delivery
claim before the
wording changed, and the end-to-end test was watched failing on the log
line.
Removing the `report_memory_write_outcome` call from the `agent`
teardown turns the
new CLI test red on its stdout assertion; restoring it turns it green
with
  `git diff HEAD` empty.

One thing that would have gone unnoticed: `capsys` captures nothing from
loguru here,
because the sink holds the real `sys.stderr` from when the handler was
added, so the
first version of the end-to-end test passed while reading an empty
string. It adds
its own sink now and asserts the sink saw something, so the log half
cannot pass
unread.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

No docs change: the only user-facing text is the shutdown notice itself,
and no
document quotes it.

## Risk

User-visible change is one line of CLI output at exit, and it now states
an unknown
outcome rather than a good one. A user who reads it has nothing to redo
-- re-sending
the turn could duplicate a write the service did complete -- so the line
is
informational by design, not an alarm.

The exit path is otherwise untouched: the budget constant is
byte-identical to the
one on the default branch, and the diff of that file contains no
non-comment line, so
exit timing does not move.

`drain_backend_stores` changes its return type, which is internal. Five
callers:
three render the notice and are updated here, two await it and ignore
the value as
they already did.

Rollback is a revert of the five commits on this branch. Nothing
persists and no
format changes. Reverting only the last two would restore the delivery
claim while
keeping the split, which is the state review rejected, so revert all or
none.

Checked and deliberately not changed:

- An existing assertion changed meaning and is called out rather than
left in the
diff. The test pinning the wording "still finishing in the background",
whose
docstring stated the write had reached the service, is renamed and
rewritten; its
two negative assertions were right and are kept, and a new sibling pins
the other
  direction.
- The memory plugin still prints its own `final-flush sweep hit its 5.0s
budget` line
to stderr at shutdown. It belongs to a different layer and is factually
correct;
  folding it into this notice would widen the diff into the plugin.
- The same teardown in `gateway_commands` stays uncovered. Its enclosing
function is
634 lines against 60 in `agent_commands`, and the gate passes without
it.
- No domain term is coined. `StorePipeline` appears in `CONTEXT.md` only
as a passing
mention with no entry of its own, and `DrainOutcome` is a module-local
type rather
  than a concept the map tracks.

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
## Summary

The parameter sweep case shipped a two-panel matplotlib plot. The same
run
also produced a dashboard, and it carries what the plot could only
assert:
the sixteen measured cells as a grid with recall and latency in each,
the
best cell ringed at top_k 10 and chunk_size 1024, and a footnote naming
the
single warm-up sample behind the 256/top_k=3 outlier. This swaps the
plot
for the dashboard and rewrites the alt text to describe that grid rather
than the two lines.

**Language.** The dashboard was authored in Chinese. Both front pages
share
one image set, and the comparison board beside this cell is already
English,
as is the alt text on both pages, so what lands is an English render
built
the same way that board was. The Chinese original is left untouched on
disk
and nothing in the repository refers to it.

**Widths.** The new render is 2000x1568 against the board's 2000x1547,
so
the two sit almost exactly on one aspect. The board goes back to full
width
and the dashboard takes 99 percent; the board only needed 86 percent
against
the old plot's 2000x1332.

## Type

- [ ] Fix
- [ ] Feature
- [x] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

```
git diff --check origin/main...HEAD
  clean

COMMIT_RANGE=origin/main...HEAD make check-large-files
  exit 0

PYTHONPATH=. uv run --frozen --python 3.12 --extra dev python scripts/check_source_language.py origin/main...HEAD
  exit 0

make check-commits
  exit 0

PR_TITLE="docs: replace the sweep chart with the run's dashboard" make check-pr-title
  exit 0

git merge-tree --write-tree HEAD origin/main
  clean
```

The attachment was fetched anonymously before it was referenced: HTTP
200,
2000x1568, md5 fcd434b3000cbcda9339e98c22bf974b, matching the local
render
byte for byte.

Rendered the section through GitHub's markdown endpoint under the
published
markdown stylesheet and measured every image box. All eight pairs share
a
top edge and match within 2px on height; the Frameworks pair is now
379x293
against 375x294, where before this change it was 326x252 against
375x294.

Both front pages were compared after the edit: each carries the
identical
set of attachment ids, and neither still references the old plot.

No test suite is relevant to a change that edits two markdown files. The
python lint job was not run locally for the same reason; CI runs it on
the
head.

- [ ] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed

## Risk

Documentation only, in both READMEs. One attachment reference is
replaced
and one width attribute is restored; no file is committed to the
repository.
The old plot remains reachable at its own attachment URL, so a revert of
the
commit restores it with no other action.

- [ ] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

Follows #555.

Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
…575)

## Summary

The ppt-engine plugin's outbound hook still wrote a PDF rendering next
to every published deck and announced it on a second `MEDIA:` line. A
delivery therefore carried two files where the user asked for one, and
the transcript showed a PDF tile beside the deck. The earlier change
that stopped the build stage from producing the preview (D64) did not
reach this hook.

The hook now publishes the deck and names it once. The web UI renders
its own preview from the deck on demand, so nothing is lost; the two
helpers that placed and named the copy are removed.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

- `uv run --frozen pytest tests/test_ppt_engine_plugin.py -q` -> 52
passed (after rebase onto main). Four tests that asserted the PDF copy
and the second MEDIA line are flipped to assert their absence.
- `ruff check` / `ruff format --check` on the two files -> clean.
- Observed on a full gateway before the fix: a template-engine delivery
arrived as `deck.pptx` plus `deck.pdf` and two transcript tiles.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

## Risk

- User-visible: a deck delivery has one file, not two. Anyone who relied
on the PDF beside the deck now opens the preview in the web UI or
converts the deck themselves.
- Rollback: revert the squash commit.

- [ ] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A
## Summary

`raven web` cannot start at all on a Windows host whose reserved port
ranges cover 8765..8784. The gateway asks `pick_port(8765)` for the
control-plane port; `pick_port` probes twenty ports forward and raises
`OSError` when every one of them is refused. That raise reached no
handler, so the whole gateway died, the supervisor retried five times
and gave up.

All twenty can be refused at once on Windows. winnat reserves whole
hundred-port blocks (`netsh interface ipv4 show excludedportrange
protocol=tcp`), a host whose dynamic port range starts low collects
those blocks in the 8000s, and a bind inside one fails with WinError
10013 while netstat shows the port unused. On the host this was found
on, the reserved blocks were 8636-8935 and 9014-9580, which swallow the
probe span whole.

The fix takes an OS-assigned port when the span is exhausted, rather
than dying on a port nobody asked for. Nothing needs this port to be
predictable: local clients read it from the lock payload, and
`ControlPlaneServer.start` already reads the bound port back off the
socket so that port 0 works. A warning names what happened, so an
operator who cares can still see that the historical port was
unavailable.

Scoped to the control plane. `pick_port` itself is unchanged, because
its other two callers serve the page, where the port is what an open
browser tab is already pointed at and moving it strands that tab.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Run on the Windows 11 host that exhibits the failure (reserved ranges
8636-8935 and 9014-9580, so 8765..8784 is entirely unbindable).

Unit tests:

```
uv run pytest tests/test_cli_gateway_commands.py tests/test_rpc_control.py tests/test_cli_gateway_page.py -q
71 passed
```

Both new behaviour tests were mutation-checked rather than assumed to
bite:

- restoring the inline `pick_port(8765)` at the call site:
`test_the_gateway_takes_its_control_port_from_the_fallback` fails;
- removing the `except OSError` fallback from the helper:
`test_the_control_plane_falls_back_to_an_os_assigned_port` fails.
- changing `pick_port`'s exhaustion raise from `OSError` to
`RuntimeError`: `test_pick_port_raises_the_class_the_fallback_catches`
fails, and is the only case that does (1 failed, 49 passed). Added in
review: the two cases above replace `pick_port` with a stub that raises
`OSError` itself, so neither reaches the real function, and the
cross-module class contract the fallback rests on went unpinned.

End to end on the same host, with an isolated `RAVEN_HOME` and every
channel disabled:

```
uv run raven gateway --config <isolated config> --port 18795
WARNING  raven.cli.gateway_commands:_control_plane_port - control plane: no free port from 8765; taking an OS-assigned one
INFO     raven.rpc.control:start - control plane: listening on ws://127.0.0.1:9594/ws
OK Control plane: ws://127.0.0.1:9594/ws
OK Health: http://127.0.0.1:18795/health
```

The published endpoint carries the port it actually got (`control_port:
9594` in the lock payload), and both control-plane clients reach it
there:

```
uv run raven gateway status   ->  gateway pid 712, generation 1, page mounted at http://127.0.0.1:18792
uv run raven gateway stop     ->  gateway stopping, both ports released
```

Before this change the same command on the same host ended in `OSError:
no free port in 8765..8784`.

Lint:

```
uv run --extra dev ruff check raven/cli/gateway_commands.py tests/test_cli_gateway_commands.py
All checks passed!
uv run --extra dev ruff format --check raven/cli/gateway_commands.py tests/test_cli_gateway_commands.py
2 files already formatted
```

Not claimed: `make ci` was not run in full. `tests/test_gateway_lock.py`
has 4 failures on this host, all of them POSIX file-mode assertions;
they reproduce identically on unmodified `origin/main` (verified by
stashing this change), so they are pre-existing and Windows-only, not
introduced here.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

## Risk

Behaviour is unchanged on any host with a free port in 8765..8784: the
fallback is reached only where the gateway previously refused to start.
Where it is reached, the control plane answers on a different port each
boot, which is already the case for the token beside it and is why both
are published together rather than assumed.

The control plane is loopback-only and token-gated on its first frame,
and neither property depends on the port number, so an OS-assigned port
is no weaker than the historical one.

Rollback is the single commit; nothing is persisted that outlives the
process beyond the lock payload, which is rewritten on every boot.

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
…571)

## Summary

Two halves of one incident, on 2026-09-20. A Raven-Design run (an ACP
sub-agent) stalled on its model stream; the parent conversation was told
`turn_failed (code -32099)` and nothing else, diagnosed it as an npm
crash and restarted the run. Behind that, the same run had written
58,675 `agent_thought_chunk` frames, one per reasoning token, averaging
5.7 characters. Every frame is a synchronous write plus flush on the
client's pipe, so when the client (the gateway, frozen by a synchronous
`find`, PR #566) stopped reading, the server blocked on that write too.

**Thought chunks are coalesced.** `UpdateTranslator` (the ACP server's
outbound sink, `raven/acp/updates.py`) now holds consecutive
`thinking.delta` updates per session and sends them as one
`agent_thought_chunk` when:

- a 200ms window closes (a loop timer, `THOUGHT_COALESCE_WINDOW_S`), or
- the held text reaches 512 characters (`THOUGHT_COALESCE_MAX_CHARS`),
or
- anything that has to follow them arrives: a chunk of the answer, a
tool call, a latch or an ending event, the session's release, the turn
slot being dropped, or the connection closing.

Order on the wire is kept: held text always precedes whatever released
it, and a delegated agent's thought (a different `_meta` target) never
joins the main agent's held text. The thresholds are measured, not
chosen: replaying the incident's frame log, 200ms/512 gives 4,333 frames
(13.5x fewer); a 100ms window gives 6,210; a 1,024-character cap changes
nothing (4,331), so the window is what does the work. The median gap
between consecutive thought frames in that log was 0.34ms, the p90 92ms.
`agent_message_chunk` (the answer itself) is not coalesced: a reader
watches it stream, and its frame count was 79.

**A failed turn names what failed.** (Second round: the class is
prefixed only when the message does not already open with it, so an SDK
error that names itself does not stutter; and `StreamIdleTimeoutError`
is classified by type, as `stream_idle_timeout`, before any substring
branch can read the budget in its message as a 429 or a 500.)
`Lane._run_turn` built `TurnFailed(error=str(exc))`, and `str()` of
asyncio's bare `TimeoutError` is empty, so the ACP translator's `_error`
had only the wire message to show. `describe_failure(exc)` now carries
the exception class and its message when it has one (`ValueError: boom`,
`TimeoutError`). And the stall itself is named at the source: the
per-chunk idle `wait_for` in `LiteLLMProvider.chat_stream` raises
`StreamIdleTimeoutError(idle=...)`, a `TimeoutError` subclass (so every
`except TimeoutError` arm and retry verdict keeps applying, the same
argument `FirstByteTimeoutError` makes) whose message reads `the model
stream sent nothing for 180s (bound stream_idle_timeout=180s)`.

Not changed: the frame writer is still a blocking write on the loop.
Fewer frames make the stall far less likely, they do not make it
impossible; moving the write off the loop would change the ordering
guarantee `UpdateTranslator`'s docstring states and is not attempted
here. The other providers' idle watchdogs (the Responses and Anthropic
SSE paths) still raise a bare `TimeoutError`; they now surface as
`TimeoutError` through `describe_failure` rather than as nothing, and
giving them the same message is a follow-up.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

At head b6097ea:

```
uv run pytest tests/test_acp_updates.py tests/test_spine_scheduler_lane.py tests/test_litellm_provider_timeout.py -q -n 0
# 233 passed
uv run pytest <24 test files across acp, spine, providers, rpc spine, subagent acp> -q
# 1048 passed (before the last test was added)
uv run pytest -q --ignore=tests/integration
# 23791 passed, 110 skipped at head b6097ea, in a throwaway worktree with COLORTERM=truecolor
make lint-python lint-imports lint-types lint-deps          # all pass
make check-source-language check-large-files                # exit 0 (COMMIT_RANGE=origin/main..HEAD)
```

Revert-to-red, each production file (pair) reverted to `origin/main` on
its own with the tests kept:

```
raven/acp/updates.py reverted                              6 failed
raven/spine/scheduler.py reverted                          2 failed
raven/providers/{first_byte,litellm_provider}.py reverted  1 failed
```

Mutation, each mutant applied alone, control restored between runs:

```
no flush before an update that must follow the thoughts   3 failed
window timer never armed                                  2 failed
512-character cap disabled                                1 failed
different speakers merged into one chunk                  1 failed
release_session keeps the held text                       1 failed
close keeps the held text                                 1 failed
end_turn keeps the held text                              1 failed  (0 before the second commit added its test)
describe_failure drops the class name                     2 failed
idle stall re-raised as a bare TimeoutError               1 failed
control (nothing mutated)                                 230 passed
```

Second round (267 tests across `test_acp_updates.py`,
`test_spine_scheduler_lane.py`, `test_error_classification.py`):

```
the three files reverted to the first-round head          4 failed
idle stall classified by substring                        3 failed
describe_failure always prefixes                          1 failed
settle_turn keeps the held text                           1 failed
sliding window (timer re-armed per chunk)                 1 failed
control                                                   267 passed
```

The two guard clauses the first round's mutation pass showed could not
decide the flush (`or result.latch`, `or _ends_the_stream(event)`) were
dropped rather than tested; the comment says why.

One existing test changed: `test_reasoning_arrives_as_a_thought_chunk`
asserted a lone thought went out immediately; it now sends two thoughts
and an answer chunk and asserts one thought frame precedes the answer,
which is the ordering rule this PR adds.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [x] User-facing docs or screenshots are updated when needed (CHANGELOG
`Unreleased / Fixed`, two entries)

## Risk

User-visible behaviour changes:

- An ACP client sees reasoning text arrive in bursts of up to 200ms or
512 characters instead of per token, so the first thought token of a
turn arrives up to 200ms later than before. Answer text, tool calls and
endings are unchanged and never reordered relative to the thoughts that
preceded them.
- Every failure event (`error` on the wire, `turn_failed` in ACP) now
starts with the exception class, e.g. `ValueError: boom` where it read
`boom`. A client matching the old text exactly would need to match the
suffix.
- The mid-stream idle stall in the LiteLLM path raises
`StreamIdleTimeoutError` instead of a bare `TimeoutError`. It is a
`TimeoutError`, so `except TimeoutError` behaves as before;
`classify_error` names it `stream_idle_timeout` (retryable, fallback)
where it used to say `network`, and `_NON_AUTH_HINTS` has no entry for
it, so the CLI renders its error line alone, as it does for
`first_byte_timeout`.

Rollback: revert the squash commit. No stored data or wire schema
changes; the coalescer is process-local state.

- [x] Security impact considered (held text is the model's own
reasoning, already on this channel; `redact` is applied where it was
before, in the failure text)
- [x] Backward compatibility considered (frame shapes unchanged and
schema-validated in the tests; exception hierarchy preserved)
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
## Summary

The agents page offered a Connect that could not be trusted, in two
separate ways, and this
closes both.

A press held for as long as the server's readiness gate takes -- one
real prompt through that
agent's own backend, measured at 16s against a live adapter -- and
nothing on the row said it
had begun. A reader with nothing to look at pressed again. The second
press then killed the
first: the connection pool keys a connection on its launch arguments,
cwd among them, and cwd
falls back to the calling task's workspace, so a ping's key never
matches a held connection
and acquire retires it. The first press came back "sub-agent did not
answer a test message",
naming the agent for a failure raven had caused. Reproduced on a live
gateway: press one
failed at 3.0s, exactly the click interval; press two succeeded at
17.1s.

Two changes, either of which alone leaves a hole. The page holds a row
while its own write is
in flight, the way the Test button beside it always has -- the writes
only, since test_cancel
has to reach the server while its test is still running. And a readiness
ping now runs on a
connection pool of its own, closed on the way out, which also stops a
ping fired while the
agent is serving a real run from taking that run's connection with it.
The same A/B afterwards:
both presses succeeded, 15.0s and 13.2s, with no relaunch between them.

The second way was the button itself. Four of the eight third-party
agents this machine can
offer would refuse the press and always would have: three have no binary
on PATH, and one
answers the handshake and then declines to open a session without a
credential. Each carried a
Connect identical to a working row's. The page now says which it is and
offers nothing --
"Not Installed" for a command that does not resolve, "Unauthorized" for
an agent that asked to
be signed in. Neither is pressable, because neither remedy is on this
page; the way back is the
card's own Test, which re-measures and replaces the recorded verdict,
plus a restart, which
re-measures a recorded refusal rather than trusting it. A connected row
is untouched, so an
expired credential cannot cost it the Disconnect it still needs.

needs_auth is measured, not inferred, and it takes the refusal's own
words. verify_agent
already read a verdict off the handshake's auth methods and the error
text, spent it on
choosing a status, and dropped it; the advertisement half is not
evidence about any one
refusal, because an agent that works advertises auth methods too -- one
measured here names
four -- so needs_auth takes only the credential-shaped refusal while the
coarse status keeps
the reading it always had. The snapshot then carries the verdict and
nothing downstream
pattern-matches an agent's prose. Two
supporting changes follow: the snapshot backfill now records a
credential refusal as well as a
pass, and covers shipped presets nobody has configured, which is exactly
the set whose Connect
was about to mislead.

Costs and residue, all deliberate:

- The backfill now handshakes the shipped acp presets that are on this
machine, not only the
configured rows: 9 instead of 4 here. Sequential, backgrounded, no model
tokens. A preset
whose command does not resolve is skipped, so nothing is launched to
learn what the free
  probe already answered.
- Recording a non-ready snapshot is narrowed to credential refusals. A
timeout or a crashed
adapter stays unrecorded, because that is a fact about one minute and
freezing it would
  label a working agent broken until somebody pressed Test.
- The snapshot store gains at most one row per shipped preset.
- The TUI roster reads the same rows and calls the same subagents.add /
subagents.toggle. It
does not read needs_auth, so its button is unchanged -- not a
regression, and not fixed here
because it would widen the diff into another surface. Named so a
reviewer does not have to
  find it.
- Recording a preset's snapshot also means a preset row's `stateful` and
its probe status are
read from a measurement rather than defaulted. More accurate in both
cases, and named here
  because it is a second consequence of one line.
- A Test records a snapshot and does not rebuild the agent table. That
is unchanged and
pre-existing for configured rows; a preset row has no roster entry to
rebuild.
- The pre-submit sweep found three defects in this branch's own code and
each is fixed in its
own commit: the page's leading comment still declared "two verbs on a
row, and only two"
after a third state was added; the preset filter resolved commands
against this process's
PATH while _probe_acp answers the row on screen from the login shell's;
and the scheduler's
early return had to go, because whether any shipped preset is on this
machine cannot be
  answered from the synchronous side.

## Type

- [x] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [ ] Other

## Verification

Run in a worktree at this branch's head, rebased onto 8d74fb9.

- `uv run --frozen --all-extras pytest tests/test_subagent_probe.py
tests/test_subagent_acp.py
  tests/test_rpc_subagents.py tests/test_rpc_schema_match.py
  tests/test_subagent_acp_mcp_lifecycle.py tests/test_acp_unprompted.py
tests/test_subagent_manager.py tests/test_acp_backend_cwd.py
tests/test_subagent_third_party.py
  -q -p no:randomly` -- 1157 passed.
- `npx vitest run src/features/xa/` in ui-web -- 91 passed.
- Full ui-web suite, serially, with the localstorage flag this box needs
-- 1840 tests, 2
failed, 0 unhandled errors. Both failures are in
src/features/workspace/, which this branch
does not touch, and the failing ID set is identical to the one recorded
before any of this
work. They are an artifact of node v26 shadowing happy-dom's
localStorage on this machine.
- `make lint-python lint-imports lint-types lint-ui lint-tui` -- all
pass. lint-bridge fails on
a TypeScript deprecation in bridge/tsconfig.json; it fails identically
on an untouched
  checkout and this branch does not touch bridge/.
- `make check-commits check-large-files check-source-language` with the
three-dot range -- all
  exit 0.
- Seventeen new tests, fourteen python and three in the page's own
suite. Every one was
watched failing for the missing behaviour before the code that answers
it was written,
except the three that cover already-written defensive branches, which
are verified by the
coverage measurement below instead. The test_cancel regression that
guards the in-flight lock was proved
load-bearing by widening the lock to the whole row and watching it go
red.
- Review round one, two blocking findings, both reproduced here before
being fixed and both
re-run afterwards. needs_auth from the advertisement alone: forcing the
message heuristic to
False left needs_auth true against the stub, and the stub gained a
`no_session_other` mode
because every existing refusal mode refuses in auth-shaped words, so no
test could separate
the two halves. The unrecoverable row: a recorded refusal, then
`run_test(source="preset")`
returning ok, then the same snapshot still reading needs_auth true; it
now reads false and
ready. Re-measured against the real agents afterwards, Grok Build still
reports needs_auth
  true and CodeBuddy still ready.
- The diff-coverage gate read 88.00% on the first head. Of the 120 lines
this branch adds to
probe.py, none is now uncovered; the four that remain in that file are
the base's own.
- A second sweep over the review-fix commits found one more defect,
fixed in its own commit: a
test comment still gave the removed gate as the reason for an assertion
that stands for a
  different one.
- End to end against real agents on this machine: the backfill reaches
the 9 presets that are
installed and skips Copilot, Qwen and Kimi, which are not; verify
reports CodeBuddy ready
with needs_auth false and Grok Build attention with needs_auth true; and
the rows come back
  reading ready/false, missing/false and attention/true respectively.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

No docs change: the agents page carries no user-facing documentation
page, and the two new
strings are catalogue entries.

## Risk

User-visible behaviour changes, all on the agents page and its server
surface:

- A Connect press disables its own button until the call returns. A
second press for the same
  row is dropped client-side rather than sent.
- A row whose command is not on the machine, and a row whose agent asked
to be signed in, no
longer offer Connect. They show a word and are not pressable. The way
back for the second is
  the card's Test.
- A gateway start now handshakes the installed shipped presets once, in
the background.
- subagents.list rows gain needs_auth. It is optional on the wire and
defaults false, so a
  client that does not read it behaves exactly as before.

Rollback is the revert of this squash commit: nothing here migrates
data, and a snapshot row
recorded by the widened backfill is re-derived or ignored by the
previous code, which reads the
same store and simply never wrote those rows.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A

---------

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
Resolves 64 conflicts against main at cf3c95f.

Twenty-five paths were taken from main unchanged: this branch never edited
them, it only carried stale copies picked up by the hand-squashed catch-up
commit 80ecae6, and main has moved on since. Nineteen delete/modify
conflicts stay deleted, the old frontend tree having been replaced.

Three resolutions made a judgement call, and are the ones to re-read:

- subagent/probe.py: main's #554 records a preset's capability snapshot
  where this branch's #559 refused to. Main's decision stands, through this
  branch's machinery, so record_capabilities remains the one writer and
  _test_record still keeps a previous measurement when a verify fails. The
  boot backfill re-measures on either side's condition.
- rpc/methods/subagents.py: both sides read the same snapshot under
  different names. One read now serves meta, own and needs_auth.
- ui-web check-class-namespace.mjs: the rail ratchet moves 11 to 12 for
  .wdt, which arrives with main's folder chip beside eleven other
  unprefixed classes of that domain.

Two gaps this merge does not close. The rebuilt settings page offers no
timezone control, so the backend half of #506 has no surface; and main's
shell/workdir.ts folder picker is left out, since it imports shell modules
this branch deleted. The rail half of #542 did merge into features/rail.

lib/upload.ts follows raven/rpc/files.py to 100 MB, and the offline fixtures
now send session workdir and subagent needs_auth.

Verified on this head: ui-web type-check and gen:check clean (190 methods),
40 of 40 convention gates, 1279 backend tests, source-language check clean.
The full ui-web suite fails the same files it fails on the unmerged head of
this branch, and no others.

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
Both files took their conflict as a union of the two sides' appended tests,
which left the joins one blank line short of what ruff format wants. Four
blank lines, no test touched.

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
@LivXue
LivXue requested review from 0xKT and gloryfromca and removed request for 0xKT September 21, 2026 08:34

@gloryfromca gloryfromca left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: the rebuilt external-agent UI does not preserve main's authentication-refusal behavior.

Two merge-specific defects are marked inline. Together they undo the user-visible contract of #554: a measured credential refusal must reach the page as a disabled Unauthorized state, including when an agent that worked before later logs out.

I covered the full target-to-head file set and main arrival history, the remerge/combined conflict resolutions, the affected callers and persisted snapshot path, backward compatibility and generated RPC/fixture shapes, test changes for weakening, AGENTS.md and the ui-web CONTEXT/CONTRIBUTING architecture rules. I found no weakened tests or additional architecture violations worth raising.

Verification on this head: the focused backend set passed 593 tests with one environment skip (python-pptx unavailable); ui-web type-check and gen:check passed (193 methods); eslint exited 0 with five existing exhaustive-deps warnings; and all 192 ui-web test files passed (2,573 tests). A direct _test_record reproduction also showed a new needs_auth=True refusal persisted as false when a prior successful snapshot exists.

The author-designated follow-ups remain nonblocking from my side: the rebuilt settings page still lacks the timezone control, and the old workdir folder picker was not ported.

Comment thread raven/rpc/methods/subagents.py
Comment thread raven/acp_client/capabilities.py

@gloryfromca gloryfromca left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: the two open authentication-refusal findings still have to be fixed.

The only delta from the reviewed head is four formatter-required blank lines in tests/test_rpc_console.py and tests/test_rpc_subagents.py. Neither affected implementation path changed, so both existing threads remain applicable and unresolved. I rechecked the resulting target-to-head diff, its callers and history, compatibility and generated-contract surfaces, the prior test coverage, AGENTS.md, and the ui-web architecture rules; I found no new findings and no weakened tests.

Verification on 4113b729c184: uv run pytest tests/test_rpc_console.py tests/test_rpc_subagents.py -q passed 228 tests; uv run ruff format --check reports both files formatted; uv run ruff check passed; git diff --check passed.

LivXue and others added 2 commits September 21, 2026 09:07
…sured

`_test_record` keeps a previous record's capabilities when a verify fails, so
one flaky Test does not cost a row its menu and statefulness. It rebuilt that
record from `previous` and copied only the verdict fields, and `needs_auth` was
not among them -- it arrived on main with #554, after this helper was written.

An agent that worked and has since been logged out therefore stayed
`needs_auth: false` through the very Test that detected the refusal, and the
page went on offering it a Connect the agent would decline.

`needs_auth` travels with the verdict, not with the capabilities: a menu
measured earlier is still the best account of what the agent can do, but
whether it will open a session without a credential is a fact about this
handshake alone.

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
…hake

The server reports `needs_auth` on an acp row whose handshake refused to open a
session without a credential. The rebuilt external-agent domain never read it:
`extAgentRowOf` is a whitelist and did not name the field, so `stageOf` saw an
unconfigured row and answered `add`, and the row rendered Connect and sent
`subagents.add` to an agent that had already said no.

Carries main's #554 contract into the domain that replaced its page: the field
on the row, the two stages that name why a write would fail rather than a write
to make, and a disabled label in place of the button, in the row and in the
sheet, which read one decision so they cannot disagree.

The store refuses those stages too. It maps anything that is not `add` or
`stale` onto `toggle`, and it is reached from the row, the sheet and the wizard
step -- a barred row arriving there would have turned a refused add into a
toggle of an entry that does not exist.

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>

@gloryfromca gloryfromca left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No blockers; this can merge as far as I am concerned.

The new delta fixes both prior blockers: the rebuilt web domain now carries the handshake authentication verdict through its source mapper and stage/UI/store decisions, and preserved capability records take needs_auth from the newest measurement while retaining only the older capability data. I replied to and resolved both settled threads.

Review coverage: the repository rules and UI architecture constraints; the delta against the previously reviewed target diff; affected callers and wire compatibility (the optional field remains safe with older servers); relevant history and intended #554 behavior; and the additive regression tests to confirm no tests were weakened.

Verification:

  • uv run pytest tests/test_subagent_acp.py tests/test_subagent_probe.py tests/test_rpc_subagents.py -q: 384 passed
  • npx vitest run src/features/extAgents/source.test.ts src/features/extAgents/store.test.ts src/features/extAgents/ExtAgentsPage.test.tsx: 81 passed
  • npm run type-check: passed
  • npm run lint: passed with 5 existing warnings in untouched files and 0 errors
  • uv run ruff check raven/agent/subagent/probe.py tests/test_subagent_acp.py: passed

@LivXue LivXue closed this Sep 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.