Skip to content

fix(*): stop the connect button promising what it cannot do - #554

Merged
LivXue merged 8 commits into
mainfrom
fix/xa_connect_press_and_unusable_rows
Sep 21, 2026
Merged

LivXue merged 8 commits into
mainfrom
fix/xa_connect_press_and_unusable_rows

Conversation

@LivXue

@LivXue LivXue commented Sep 20, 2026 •

Copy link
Copy Markdown
Member

Summary

The agents page offered a Connect that could not be trusted, in two separate ways, and this
closes both.

A press held for as long as the server's readiness gate takes -- one real prompt through that
agent's own backend, measured at 16s against a live adapter -- and nothing on the row said it
had begun. A reader with nothing to look at pressed again. The second press then killed the
first: the connection pool keys a connection on its launch arguments, cwd among them, and cwd
falls back to the calling task's workspace, so a ping's key never matches a held connection
and acquire retires it. The first press came back "sub-agent did not answer a test message",
naming the agent for a failure raven had caused. Reproduced on a live gateway: press one
failed at 3.0s, exactly the click interval; press two succeeded at 17.1s.

Two changes, either of which alone leaves a hole. The page holds a row while its own write is
in flight, the way the Test button beside it always has -- the writes only, since test_cancel
has to reach the server while its test is still running. And a readiness ping now runs on a
connection pool of its own, closed on the way out, which also stops a ping fired while the
agent is serving a real run from taking that run's connection with it. The same A/B afterwards:
both presses succeeded, 15.0s and 13.2s, with no relaunch between them.

The second way was the button itself. Four of the eight third-party agents this machine can
offer would refuse the press and always would have: three have no binary on PATH, and one
answers the handshake and then declines to open a session without a credential. Each carried a
Connect identical to a working row's. The page now says which it is and offers nothing --
"Not Installed" for a command that does not resolve, "Unauthorized" for an agent that asked to
be signed in. Neither is pressable, because neither remedy is on this page; the way back is the
card's own Test, which re-measures and replaces the recorded verdict, plus a restart, which
re-measures a recorded refusal rather than trusting it. A connected row is untouched, so an
expired credential cannot cost it the Disconnect it still needs.

needs_auth is measured, not inferred, and it takes the refusal's own words. verify_agent
already read a verdict off the handshake's auth methods and the error text, spent it on
choosing a status, and dropped it; the advertisement half is not evidence about any one
refusal, because an agent that works advertises auth methods too -- one measured here names
four -- so needs_auth takes only the credential-shaped refusal while the coarse status keeps
the reading it always had. The snapshot then carries the verdict and nothing downstream
pattern-matches an agent's prose. Two
supporting changes follow: the snapshot backfill now records a credential refusal as well as a
pass, and covers shipped presets nobody has configured, which is exactly the set whose Connect
was about to mislead.

Costs and residue, all deliberate:

  • The backfill now handshakes the shipped acp presets that are on this machine, not only the
    configured rows: 9 instead of 4 here. Sequential, backgrounded, no model tokens. A preset
    whose command does not resolve is skipped, so nothing is launched to learn what the free
    probe already answered.
  • Recording a non-ready snapshot is narrowed to credential refusals. A timeout or a crashed
    adapter stays unrecorded, because that is a fact about one minute and freezing it would
    label a working agent broken until somebody pressed Test.
  • The snapshot store gains at most one row per shipped preset.
  • The TUI roster reads the same rows and calls the same subagents.add / subagents.toggle. It
    does not read needs_auth, so its button is unchanged -- not a regression, and not fixed here
    because it would widen the diff into another surface. Named so a reviewer does not have to
    find it.
  • Recording a preset's snapshot also means a preset row's stateful and its probe status are
    read from a measurement rather than defaulted. More accurate in both cases, and named here
    because it is a second consequence of one line.
  • A Test records a snapshot and does not rebuild the agent table. That is unchanged and
    pre-existing for configured rows; a preset row has no roster entry to rebuild.
  • The pre-submit sweep found three defects in this branch's own code and each is fixed in its
    own commit: the page's leading comment still declared "two verbs on a row, and only two"
    after a third state was added; the preset filter resolved commands against this process's
    PATH while _probe_acp answers the row on screen from the login shell's; and the scheduler's
    early return had to go, because whether any shipped preset is on this machine cannot be
    answered from the synchronous side.

Type

  • Fix
  • Feature
  • Docs
  • CI / tooling
  • Refactor
  • Other

Verification

Run in a worktree at this branch's head, rebased onto 8d74fb9.

  • uv run --frozen --all-extras pytest tests/test_subagent_probe.py tests/test_subagent_acp.py tests/test_rpc_subagents.py tests/test_rpc_schema_match.py tests/test_subagent_acp_mcp_lifecycle.py tests/test_acp_unprompted.py tests/test_subagent_manager.py tests/test_acp_backend_cwd.py tests/test_subagent_third_party.py -q -p no:randomly -- 1157 passed.

  • npx vitest run src/features/xa/ in ui-web -- 91 passed.

  • Full ui-web suite, serially, with the localstorage flag this box needs -- 1840 tests, 2
    failed, 0 unhandled errors. Both failures are in src/features/workspace/, which this branch
    does not touch, and the failing ID set is identical to the one recorded before any of this
    work. They are an artifact of node v26 shadowing happy-dom's localStorage on this machine.

  • make lint-python lint-imports lint-types lint-ui lint-tui -- all pass. lint-bridge fails on
    a TypeScript deprecation in bridge/tsconfig.json; it fails identically on an untouched
    checkout and this branch does not touch bridge/.

  • make check-commits check-large-files check-source-language with the three-dot range -- all
    exit 0.

  • Seventeen new tests, fourteen python and three in the page's own suite. Every one was
    watched failing for the missing behaviour before the code that answers it was written,
    except the three that cover already-written defensive branches, which are verified by the
    coverage measurement below instead. The test_cancel regression that guards the in-flight lock was proved
    load-bearing by widening the lock to the whole row and watching it go red.

  • Review round one, two blocking findings, both reproduced here before being fixed and both
    re-run afterwards. needs_auth from the advertisement alone: forcing the message heuristic to
    False left needs_auth true against the stub, and the stub gained a no_session_other mode
    because every existing refusal mode refuses in auth-shaped words, so no test could separate
    the two halves. The unrecoverable row: a recorded refusal, then run_test(source="preset")
    returning ok, then the same snapshot still reading needs_auth true; it now reads false and
    ready. Re-measured against the real agents afterwards, Grok Build still reports needs_auth
    true and CodeBuddy still ready.

  • The diff-coverage gate read 88.00% on the first head. Of the 120 lines this branch adds to
    probe.py, none is now uncovered; the four that remain in that file are the base's own.

  • A second sweep over the review-fix commits found one more defect, fixed in its own commit: a
    test comment still gave the removed gate as the reason for an assertion that stands for a
    different one.

  • End to end against real agents on this machine: the backfill reaches the 9 presets that are
    installed and skips Copilot, Qwen and Kimi, which are not; verify reports CodeBuddy ready
    with needs_auth false and Grok Build attention with needs_auth true; and the rows come back
    reading ready/false, missing/false and attention/true respectively.

  • Relevant tests pass locally

  • Relevant lint / type checks pass locally

  • User-facing docs or screenshots are updated when needed

No docs change: the agents page carries no user-facing documentation page, and the two new
strings are catalogue entries.

Risk

User-visible behaviour changes, all on the agents page and its server surface:

  • A Connect press disables its own button until the call returns. A second press for the same
    row is dropped client-side rather than sent.
  • A row whose command is not on the machine, and a row whose agent asked to be signed in, no
    longer offer Connect. They show a word and are not pressable. The way back for the second is
    the card's Test.
  • A gateway start now handshakes the installed shipped presets once, in the background.
  • subagents.list rows gain needs_auth. It is optional on the wire and defaults false, so a
    client that does not read it behaves exactly as before.

Rollback is the revert of this squash commit: nothing here migrates data, and a snapshot row
recorded by the widened backfill is re-derived or ignored by the previous code, which reads the
same store and simply never wrote those rows.

  • Security impact considered
  • Backward compatibility considered
  • Rollback path is clear for risky changes

Related Issues

N/A

@LivXue
LivXue requested review from 0xKT and gloryfromca September 20, 2026 09:45

@gloryfromca gloryfromca left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: needs_auth can be falsely assigned and can remain permanently stale after sign-in.

I found two blocking paths, marked inline. I covered the full diff, the UI/RPC/runtime callers, the relevant history, AGENTS.md and context-map constraints, backward compatibility of the wire field, connection-pool architecture, and whether tests were weakened. The added tests are meaningful, but neither failure path below is covered.

Verification on this head: uv run pytest tests/test_subagent_probe.py tests/test_subagent_acp.py tests/test_rpc_subagents.py tests/test_rpc_schema_match.py -q passed (722); npm test -- src/features/xa passed (91); npm run type-check passed; npm run gen:check passed (178 methods). git diff --check github/main...HEAD also passed.

Comment thread raven/acp_client/capabilities.py
Comment thread raven/agent/subagent/probe.py
LivXue and others added 8 commits September 20, 2026 11:07
The agents page holds a Connect press for as long as the server's
readiness gate takes -- one real prompt through that agent's own backend,
measured at 16s against a live adapter -- and nothing on the row said it
had begun. A reader with nothing to look at pressed again, and the second
press reached the server as a second connect for the same agent.

That second press then killed the first. The connection pool keys a
connection on its launch arguments, cwd among them, and cwd falls back to
the calling task's workspace; ping_agent gives every ping a fresh
temporary directory, so a ping's launch key never matches a held
connection. acquire therefore retired the first press's connection as a
stale launch config, and the first press came back "did not answer a test
message" -- naming the agent for a failure raven had caused.

Two changes, either of which alone leaves a hole.

The page now holds a row while its own write is in flight, the way the
Test button beside it always has: the store drops a second call for a
name already in flight, and the button says so rather than sitting there
looking pressable. The writes only -- test_cancel has to reach the server
while its own test is still running, and locking the row for the length
of a two-minute test would make Stop unpressable.

And a readiness ping now runs on a connection pool of its own, closed on
the way out. That also keeps a ping fired while the agent is serving a
real run from taking that run's connection down with it, which no amount
of gating inside one page can prevent.

AcpAgentBackend takes its pool as an argument and resolves the default
lazily, so close_pool still means what it says. The session-open timeout
test hands one in rather than patching the module global; the behaviour
it asserts is unchanged.

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
Four of the eight third-party agents this machine can offer refused the
press and always would have: three have no binary on PATH, and one answers
the handshake and then declines to open a session without a credential.
Each carried a Connect identical to a working row's, so the only way to
learn which kind of row it was was to spend a launch and read a sentence
written for that agent's own maintainer.

The page now says which it is and offers nothing: "Not Installed" for a
command that does not resolve, "Unauthorized" for an agent that asked to be
signed in. Neither is pressable, because neither remedy is on this page --
installing the command and signing in both happen elsewhere, and the way
back is the card's own Test, which re-measures and lets the row return to
Connect. A connected row is untouched: an expired credential must not cost
it the Disconnect it still needs.

needs_auth is measured, not inferred. verify_agent already read it off the
handshake's auth methods and the error text, spent it on choosing a status,
and dropped it; what survived was auth_methods, which a perfectly usable
agent advertises too -- one measured here names four. So the snapshot
carries the verdict, the row carries it out, and nothing downstream has to
pattern-match an agent's prose.

Two supporting changes fall out of that. The snapshot backfill now records
a credential refusal as well as a pass: that verdict was the one worth
remembering and the only one being thrown away. Every other failure stays
unrecorded, because a timeout or a crashed adapter is a fact about one
minute, and freezing it into a snapshot would label a working agent broken
until somebody happened to press Test. And the backfill now covers shipped
presets nobody has configured, which is exactly the set whose Connect was
about to mislead; a preset whose command does not resolve is skipped, so
nothing is launched to learn what the free probe already answered.

rpc-schema/openrpc.json is the contract both clients generate from, with
raven/rpc/models.py as its mirror, so the field is added in all three.

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
The file opens by declaring its own contract -- "two verbs on a row, and
only two" -- and the change before this one put a third thing there. A word
where a button was is not a third verb, but the comment has to say so, or
the next reader takes the sentence at face value and reads the label as a
button somebody forgot to wire.

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
The snapshot backfill skips a preset whose command is not on the machine,
which is right, and it was asking the wrong PATH: this process's, while
_probe_acp answers the row on screen from the login shell's. The two differ
on any install that puts an agent on the interactive path only -- the page
would report the agent installed and the backfill would quietly decline to
verify it, so the row it was added to inform stayed unverified for the one
reason nobody could see.

The repo already had this exact finding for the cli probe
(test_cli_probe_resolves_on_the_login_shell_path_not_ravens); the filter
added here is the same question asked a second time.

Capturing it needs an await, so the selection moves inside the task that was
already async, and the scheduler no longer returns early on an empty
registry -- whether any shipped preset is on this machine is precisely what
it cannot answer from the synchronous side.

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
`initialize` lists the auth methods an agent supports, and an agent that works
lists them too: the stub advertises one, CodeBuddy advertises four and is ready.
So `bool(handshake.auth_methods)` says nothing about whether THIS session refusal
was about a credential, and taking it meant any unrelated remote error produced
`needs_auth`. Reproduced by forcing the message heuristic to False against the
stub: `needs_auth` stayed True, so the advertisement alone was deciding it. Once
the backfill persists that, the page shows a disabled "Unauthorized" for a
transient model or configuration failure, with no way back.

The advertisement keeps the job it was already doing. The coarse status only
decides whether a reader should look at this row, and being an agent that can
want signing in is good enough evidence for that, so `status` is computed exactly
as before and no existing verdict moves. What narrows is `needs_auth`, which is
persisted, rendered as a disabled control and read as an instruction: it now
takes the refusal's own words.

Re-measured against the two real agents this separates: Grok Build still reports
needs_auth true ("Authentication required"), CodeBuddy still ready.

The stub grows a `no_session_other` mode, because every existing one that refuses
a session refuses it in auth-shaped words -- which is why no test could tell the
two halves of the disjunct apart.

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
The page offers no press on such a row -- the remedy is not on the page -- so
its whole recovery path is: sign in, then press Test on the card. That path did
not exist. `_test_acp` records on `source == "config"` alone, and every row this
label appears on is an unconfigured preset, so the Test passed and the recorded
refusal outlived it. Reproduced: a recorded snapshot with needs_auth true, then
`run_test(cfg, source="preset")` returning ok, then the same snapshot still
reading needs_auth true. Restart did not recover it either, because staleness
asks whether the launch config moved and signing in does not move it.

Both halves close here. A preset's Test records, and the gate's own worry is
answered by how the store is keyed rather than by refusing to write:
`SnapshotStore.load` matches on a fingerprint of the launch fields, so a row
recorded under a preset's name is returned only to a config that launches the
same way -- one the user never adds is a row nothing asks for, and one they add
under another name does not match it. And the startup backfill re-measures a
recorded credential refusal instead of trusting it, which is the one verdict the
user is expected to go and change; a fresh pass is still taken on trust, because
re-proving it would spend a process per row per boot for an answer that cannot
have moved.

Verified on the same reproduction afterwards: the Test now leaves
needs_auth false, status ready.

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
The diff-coverage gate read 88.00% against a 90% threshold, and every uncovered
line was a defensive branch of the preset selector: a command no shell can split,
a machine where none of them resolves, and a table the schema will not read.

Each runs at startup, before anything the user did, so a packaging fault there is
not their typo to see and must not take the gateway down. That is worth asserting
rather than assuming: the branches now answer with the rows they could build,
which for these inputs is none. Measured afterwards with coverage over
raven.agent.subagent.probe: of the 120 lines this branch adds to that file, none
is uncovered; the four that remain are the base's own.

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
The assertion is unmoved and still right; the reason above it stopped being
true one commit ago. It read that `_test_acp` records no snapshot for a preset,
which was the gate that made a signed-in row unrecoverable and is now gone. What
actually holds, and held all along, is that a Test writes a capability snapshot
and never a roster entry, so there is nothing for the table to be rebuilt from --
for a preset exactly as for a configured row, which has always recorded one here
and never applied either.

Found by the pre-submit sweep over the review-fix commits, grepping the removed
rule's own wording rather than its symbols.

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
@LivXue
LivXue force-pushed the fix/xa_connect_press_and_unusable_rows branch from 80e865b to 9d06def Compare September 20, 2026 11:08

@gloryfromca gloryfromca left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No blockers; this can merge as far as I am concerned.

The two prior blockers are fixed and their threads are resolved. I reviewed the rebased full diff and the four new fix/test commits, including the UI-to-RPC-to-snapshot recovery path, connection-pool behavior, project rules and context terms, history, wire backward compatibility, and whether tests were weakened. I found no new defect worth raising.

Verification on this head: uv run pytest tests/test_subagent_probe.py tests/test_subagent_acp.py tests/test_rpc_subagents.py tests/test_rpc_schema_match.py -q passed (726); npm test -- src/features/xa passed (91); npm run type-check passed; npm run gen:check passed (178 methods); git diff --check github/main...HEAD passed; and the merge-tree check against current github/main was clean.

@0xKT

0xKT commented Sep 20, 2026

Copy link
Copy Markdown
Member

Not a blocker. The split this PR draws is the right one, and the private
pool in ping_agent fixes a real collision. Three observations about what
happens to a needs_auth after it is written, none of which holds the PR. Read
on head 9d06def against base 0cb8b79.

1. A recorded needs_auth is a one-way latch, and the per-boot re-measure cannot clear it

Two new lines, read together:

probe.py:678  if snapshot is not None and not snapshot.stale
                 and not getattr(snapshot, "needs_auth", False): continue
probe.py:685  if result.status == "ready" or getattr(result, "needs_auth", False):
                  store.record(result)

The first makes a needs_auth snapshot permanently un-fresh, so it is
re-verified at every boot -- a child process each time, which the comment above
it argues is worth paying. The second is the only door to record, and it is
open for exactly two outcomes. So the one case the re-measure exists to catch --
the agent is no longer refusing over a credential -- is discarded unless it
comes back fully ready. An agent that now fails for a plainly different reason
leaves needsAuth: true on disk, and the row keeps rendering a disabled
"Unauthorized" for a verdict the code measured to be wrong and threw away, one
process per boot, forever.

Measured through schedule_snapshot_verification with the real store and a real
child adapter, launch fingerprint held constant so the snapshot compares as
fresh rather than stale:

installed, never signs in        boot 1..5   spawns=1 each, needsAuth stays true
CONTROL: an agent that passes    boot 1 spawns=1, boots 2..5 spawn=0
latch: refusal changes to a
  non-credential one             boots 1..3  spawns=1 each, needsAuth stays true
CONTROL: re-measure comes
  back ready                     boot 2 clears it, boot 3 spawns=0

The control matters: a passing agent is skipped after the first boot, so the
repetition is the new not needs_auth clause and not the harness.

The way out exists -- the card's Test records unconditionally
(probe.py:536) and canTest is !builtin && !vendored, so the row can be
cured by hand, which is why this is not blocking. But the label the user is
shown says "Unauthorized", which tells them to go and sign in; signing in is
exactly what will not help, because the launch config has not moved and the
boot-time path cannot write the cure.

The input that puts a row into this state is _looks_like_auth, now
load-bearing for a persisted field. It is a substring match over
("auth", "unauthor", "credential", "api key", "apikey", "login", "401"), so
.../scratch-4012/... (a path), unknown author field and
/bin/login-shell-wrapper exited 127 all read as credential refusals. Being
fair about the evidence: those three are constructed, and against the shipped
adapters no false positive appeared -- Claude Code and Codex came back ready,
and Pi's "executable not found" correctly gave needs_auth=False. The latch is
measured; the trigger is, so far, hypothetical.

2. A stale snapshot's needs_auth is served as a present-tense verdict

acp_snapshot_for loads with allow_stale=True and logs a warning when what it
returns is stale. subagents.py:342 then emits
bool(getattr(caps, "needs_auth", False)) with no stale guard between them, and
XaPage.tsx:137 turns it into the unauthorized stage, which :180-183 renders
as a disabled button. For the same row, base served off with a Connect button:
a control the reader had is withdrawn on the strength of a measurement the code
itself has already flagged as taken against a launch config the agent no longer
has.

The page has a stale stage and uses it elsewhere (:347), so the distinction
exists here -- it just is not consulted on this path.

One more ordering detail: :136-138 returns unauthorized before :139's
if (!row.configured) return 'add', so an unconfigured preset row that picked up
a latched needs_auth from the new boot-time verify loses its Add button, not
only its Connect.

3. An unconfigured preset that comes back neither ready nor unauthorized is retried every boot with no memory of having tried

Same two lines as (1), from the other side. A preset has no snapshot, so :678
never skips it; if verify_agent returns anything but ready or a credential
refusal, :685 writes nothing, and the next boot is identical. For a preset whose
command is an npx -y <pkg>@<ver> fetch, that is the fetch paid again at every
gateway start, with nothing on disk to show for it. Measured: three boots, three
npx -y pi-acp@0.0.33 spawns, snapshot file unchanged.

Recording the failure is not obviously right either -- the comment at :681-684
explains why not, and it is a good reason. But the two decisions together leave
a class of row that is re-probed forever and never settles, and nothing bounds
the repetition.

@0xKT
0xKT self-requested a review September 21, 2026 05:58
@LivXue
LivXue merged commit cf3c95f into main Sep 21, 2026
29 of 32 checks passed
@LivXue
LivXue deleted the fix/xa_connect_press_and_unusable_rows branch September 21, 2026 07:21
LivXue added a commit that referenced this pull request Sep 21, 2026
…sured

`_test_record` keeps a previous record's capabilities when a verify fails, so
one flaky Test does not cost a row its menu and statefulness. It rebuilt that
record from `previous` and copied only the verdict fields, and `needs_auth` was
not among them -- it arrived on main with #554, after this helper was written.

An agent that worked and has since been logged out therefore stayed
`needs_auth: false` through the very Test that detected the refusal, and the
page went on offering it a Connect the agent would decline.

`needs_auth` travels with the verdict, not with the capabilities: a menu
measured earlier is still the best account of what the agent can do, but
whether it will open a session without a credential is a fact about this
handshake alone.

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
LivXue added a commit that referenced this pull request Sep 21, 2026
…hake

The server reports `needs_auth` on an acp row whose handshake refused to open a
session without a credential. The rebuilt external-agent domain never read it:
`extAgentRowOf` is a whitelist and did not name the field, so `stageOf` saw an
unconfigured row and answered `add`, and the row rendered Connect and sent
`subagents.add` to an agent that had already said no.

Carries main's #554 contract into the domain that replaced its page: the field
on the row, the two stages that name why a write would fail rather than a write
to make, and a disabled label in place of the button, in the row and in the
sheet, which read one decision so they cannot disagree.

The store refuses those stages too. It maps anything that is not `add` or
`stale` onto `toggle`, and it is reached from the row, the sheet and the wizard
step -- a barred row arriving there would have turned a refused add into a
toggle of an entry that does not exist.

Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
gloryfromca added a commit that referenced this pull request Sep 21, 2026
## Summary

Second catch-up of `refactor/ui_web_architecture` with `main`, this time
to
`6ebc2baa` (#592): the 31 commits `main` gained since #529 took it to
`dbe259f6`. Synced main to `6ebc2baa`; the next round starts from that
SHA.

#529 landed as a squash, so git still saw the two branches diverge at
`5a29ea22` and re-raised every conflict #529 had already settled (68
paths
instead of 33). The first commit here merges `dbe259f6` with `-s ours`
--
tree unchanged, ancestry recorded -- so the second commit merges only
what
is new; a third takes #528 and a fifth takes #592 (READMEs only, no
conflict), both of which landed on `main` while the earlier commits were
being verified. The same steps work for every later round.

Thirty-three paths conflicted in the second commit, one in the third;
the
fourth commit is what an independent read of the merge turned up (the
sheet's Connect, `capabilities_wanted`'s fourth case, the dead styles);
the
sixth merges this branch's own tip (#584, #591, #586), where only the
fixture-shape gate's UNSENT list met (both sides kept). How each
conflict
was settled:

- 17 legacy ui-web files (`live/`, `demo/`, `features/xa`,
  `features/knowledge`, `scripts/boot-order.test.mjs`,
  `composer/open-conversation.test.ts`), edited on `main` for #542, #549
  and #554: deleted on this branch, and the deletions stand. `main`'s
  `shell/workdir.ts` and its test, which git relocated into `lib/`, are
  dropped for the same reason -- see the follow-up below.
- `main.tsx`, `page.html`, `state/session/registry.test.ts`: this
branch's
version. `main`'s edits there were the legacy chip wiring; the registry
conflict was a rename git guessed from
`scripts/page-switch-live.test.mjs`.
- `raven/agent/subagent/probe.py`: both imports; Test records a snapshot
  for a preset too (`main`, #554) through `record_capabilities` (this
  branch, #559); a kept record carries `needs_auth`; the boot backfill
  re-measures an outdated model menu (this branch) and a recorded
  credential refusal (`main`).
- `raven/rpc/methods/subagents.py`: `needs_auth` reads the snapshot this
  branch already loads per row.
- `raven/agent/subagent/manager.py` (#528, third commit): both sides
kept
at both hunks -- `main`'s two shutdown refusals beside
`SPAWN_REFUSED_PREFIX`
  with this branch's `_row_pin` after them, and the spawn announce that
types its mark and carries `node_id` (this branch) then returns early
when
  `_inject` was refused (`main`).
- `tests/test_subagent_probe.py`, `test_subagent_acp.py`,
`test_rpc_subagents.py`, `test_rpc_console.py`, `test_config_update.py`:
  both sides' new tests, in order. This branch's two backfill tests pin
  `_unconfigured_acp_preset_rows` the way `main`'s do, and the
`_agent_home` helper both sides defined is one function passing a
`Path`,
  as the config property does.
- `README.md`, `README.zh-CN.md`: `main`'s figure and showcase section.
- `ui-web/src/rpc/generated.ts`, `ui-tui/src/rpc/generated.ts`,
  `ui-tui/src/i18n/messages.generated.ts`: regenerated from the merged
  schema and catalogue.

Carried onto the rebuilt tree, since the backend halves merged cleanly:

- #554: the agent hub's rows, its sheet and the wizard step name a
refused
  agent `Unauthorized` (disabled) instead of offering Connect; the sheet
offers Test beside it, the press that re-measures the handshake and lets
the row return to Connect after a sign-in. `capabilities_wanted` treats
a
recorded refusal as wanting a re-measure, the same fourth case the boot
  backfill gained from `main`. Four tests.
- #542: rail rows carry the session's `workdir`, so the folder group
`RailPage.tsx` gained in the merge can fill. The tag's class moves under
  the rail namespace (`rail-wdt`) with its rule in the new
  `features/rail/styles.css`, as check-class-namespace requires.
- #549: the page-side upload ceiling follows `raven/rpc/files.py` to 100
MB.
- The offline fixtures send `needs_auth`;
`session.list.sessions[].workdir`
is pinned unsent in fixture-shape until the chip lands. The folder
chip's
thirty lines of popover rules that `main` added to `styles/page.css` are
  dropped with the chip; the follow-up brings its own.

Follow-up, not in this PR: the folder chip itself (#542's
`shell/workdir.ts`: the composer chip, the `fs.dirs` browser, `workdir`
on
`session.create`) has to be rebuilt on this tree's chip pattern
(`state/perm.ts` + `chrome/PermChip.tsx`). The backend, the i18n keys
and
the rail half are already here.

## Type

- [ ] Fix
- [ ] Feature
- [ ] Docs
- [ ] CI / tooling
- [ ] Refactor
- [x] Other (catch-up merge)

## Verification

Run in the worktree on the merged tree:

- `uv run pytest tests/test_subagent_probe.py tests/test_subagent_acp.py
tests/test_rpc_subagents.py tests/test_rpc_console.py
tests/test_config_update.py tests/test_rpc_session.py
tests/test_rpc_contract_shapes.py tests/test_rpc_schema_match.py
tests/test_rpc_registration.py tests/test_i18n_boundary.py
tests/test_subagent_manager.py tests/test_subagent_dag_runner.py
tests/test_rpc_bootstrap.py tests/test_cli_tui_commands.py
tests/test_cli_gateway_commands.py tests/test_rpc_transport.py
tests/test_cli_serve_commands.py tests/test_importer_orchestrator.py
tests/test_cli_gateway_page.py -q` on the final head -- 1893 passed (the
same files, run per commit as the branch grew: 1141, 590, 501).
- `make lint-python` -- ruff check clean; ruff format clean after
formatting the three spliced test files.
- `npm test` in `ui-web` on the final head -- 189 files, 2468 tests
passed (before the sync merge: 193 files, 2609; the first run had failed
two gates, check-class-namespace and fixture-shape, on the merged `.wdt`
class and the three new optional wire fields, both fixed as described
above).
- `npm run type-check`, `npm run lint` (0 errors, 5 pre-existing
warnings), `npm run gen:check` (193 methods) in `ui-web`.
- `npm run lint:i18n`, `npm run lint:rpc`, `npm run type-check`, `npm
test` in `ui-tui` -- all clean.
- `make build-ui` on the final head -- page built, `boot-snapshot: OK`
on both snapshots (235 nodes, the golden #591 set).
- `scripts/check_source_language.py
origin/refactor/ui_web_architecture..HEAD` and
`scripts/check_large_files.py origin/refactor/ui_web_architecture..HEAD`
-- both exit 0; `git merge-tree --write-tree origin/main HEAD` -- clean.

The full Python suite is left to CI.

- [x] Relevant tests pass locally
- [x] Relevant lint / type checks pass locally
- [ ] User-facing docs or screenshots are updated when needed

## Risk

No source change of its own beyond the resolutions and the three carries
above. On the page: a refused external agent now shows a disabled
`Unauthorized` button where it showed Connect; conversations pinned to a
folder group under their own rail heading with the folder's name on the
row; uploads up to 100 MB are sent instead of refused at 25 MB. All
three
match what `main` ships today.

This branch is squash-only, so the merge base stays at `5a29ea22` after
this lands; the next round repeats the `-s ours 6ebc2ba` step first.
Rollback is reverting the squash commit; the branch is back at
`99c37e74`.

- [x] Security impact considered
- [x] Backward compatibility considered
- [x] Rollback path is clear for risky changes

## Related Issues

N/A. Unblocks #475.

---------

Co-authored-by: xfng-sd <xufang@shanda.com>
Co-authored-by: Claude (claude-opus-5) <noreply@anthropic.com>
Co-authored-by: Handsome-wzw <68996445+Handsome-wzw@users.noreply.github.com>
Co-authored-by: Dizhan Xue <dizhan.xue@evermind.ai>
Co-authored-by: Kevin Hu <kevinhu.sh@gmail.com>
Co-authored-by: admin <admin@SH-HuKai.local>
Co-authored-by: 江国庆/Forrest <mr.jianggq@163.com>
Co-authored-by: 江国庆 <guoqingjiang@deepglint.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: ypflll <ypflll@163.com>
Co-authored-by: yao pengfei <yaopengfei@shanda.com>
Co-authored-by: zhanghui <huizhang1995@gmail.com>
Co-authored-by: Zuyi Zhou <144661423+ZuyiZhou@users.noreply.github.com>
Co-authored-by: Tong Li <litong02@shanda.com>
Co-authored-by: litong <238663200+TongLi31@users.noreply.github.com>
Co-authored-by: Tchen-data <176354753+Tchen-data@users.noreply.github.com>
Co-authored-by: gloryfromca <23442919+gloryfromca@users.noreply.github.com>
Co-authored-by: Dizhan Xue <xuedizhan17@mails.ucas.ac.cn>
Co-authored-by: silverLXT <25195409+silverLXT@users.noreply.github.com>
Co-authored-by: xiaotian.luo <xiaotian.luo@thetahealth.ai>
Co-authored-by: userName20260323 <zhao.wang@evermind.ai>
Co-authored-by: zhao.wang <270284818+userName20260323@users.noreply.github.com>
Co-authored-by: arelchan <1239372199@qq.com>
Co-authored-by: arelchan <204152633+arelchan@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants