Skip to content

[upstream #12874] feat: add Reasonix as first-class TUI agent - #148

Closed
innocarpe wants to merge 1075 commits into
mainfrom
fix/reasonix-tui-agent
Closed

[upstream #12874] feat: add Reasonix as first-class TUI agent#148
innocarpe wants to merge 1075 commits into
mainfrom
fix/reasonix-tui-agent

Conversation

@innocarpe

Copy link
Copy Markdown
Owner

Portfolio mirror of my contribution to upstream stablyai/orca.
Exhibition only — the real review/merge target is upstream.

Upstream

Summary

Description Register Reasonix (DeepSeek-native coding agent CLI) as a first-class TUI agent for detect, launch, catalog, and icon. ## Focused fix - In scope: TuiAgent + TUI_AGENT_CONFIG, catalog entry, icon assets, process/title recognition, telemetry kind mapping, mo

Note

  • Do not merge this into innocarpe/orca main until the upstream PR is merged.
  • After upstream merges: sync fork from upstream, then close this mirror PR.
  • This open PR exists so visitors see in-flight work on this fork's Pull requests tab.

AmethystLiang and others added 30 commits August 2, 2026 23:27
…yai#12233)

* fix(terminal): route remote-runtime link clicks to the system browser

Terminal link clicks classified ownership from the global
activeRuntimeEnvironmentId, which is null when runtimes are bound per
workspace, so a link clicked in a remote-hosted pane opened a local-only
Orca browser tab and never reached the host. Thread each pane's resolved
runtimeEnvironmentId into openHttpLink as sourceOwner across the OSC 8,
WebLinksAddon, and click-fallback paths.

Co-authored-by: Orca <help@stably.ai>

* fix(terminal): route link clicks based on pane ownership, not global sta

Clicking links on remote-hosted panes was routing based on global runtime state, causing unexpected reconnections. Now link routing decisions (where to open: Orca vs system browser) are based on the actual pane's owner — local, SSH connection, remote runtime, or unknown — regardless of whether any runtime is globally active. This ensures a local pane can route to Orca while another pane's remote runtime is active, and a remote pane always routes to the system browser.

---------

Co-authored-by: Orca <help@stably.ai>
…tablyai#12103) (stablyai#12204)

* fix(onboarding): run skill setup in the configured Windows runtime (stablyai#12103)

Onboarding was the one skill-setup surface that did not route its install
command through the resolved runtime. Settings, the feature-wall panels and
the Linear prompt all wrap theirs as `wsl.exe -d <distro> -- sh -c ...` and
pass a matching shell override; onboarding spawned a bare terminal and handed
it the raw `npx skills add ...`. With Node inside WSL, npx is not on the
Windows PATH, so the install failed.

The runtime resolver had a second gap behind that: it only consulted
per-project settings, and onboarding runs before any project exists. With no
project it returned undefined and fell through to the Windows host, ignoring
a global WSL default entirely. `getLocalAgentPreflightContext` already had a
no-project fallback for PATH detection; the skill-install path had none.

- extract that fallback as `getGlobalWindowsExecutionRuntimeContext` and
  rewire the existing agent-preflight branch through it so the two cannot drift
- adopt it in `useActiveProjectSkillRuntime` when no project is active. WSL
  only: a windows-host default already matches the old no-project behavior,
  and resolving it would hand skill discovery a target where it had none,
  re-triggering scans for every host-default user
- build the onboarding terminal's command for the runtime and pass its shell
  override
- register the CLI in WSL rather than on the host, so `orca` lands on the PATH
  the install actually runs on, and wrap the copied command to match

* fix(onboarding): keep skill setup runtime consistent

* test(onboarding): satisfy runtime settings contract

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
…rktrees (STA-3123) (stablyai#12235)

* fix(mobile): surface worktree catalog failures instead of showing 0 worktrees (STA-3123)

A connected host whose worktree.ps request fails now shows an explicit
catalog-failure state (with the RPC error code) on the host page, and
'Worktree list unavailable' on the home host card, instead of silently
rendering as a healthy host with zero workspaces.

* fix(mobile): mark cached worktree catalogs unavailable
…#11987)

* fix(terminal): expand variables in Windows PATH

* fix(terminal): preserve expanded Windows PATH at spawn

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
Adapted from stablyai#11616 commit 9f449d7 and stablyai#12120 commit 290f42a.

Co-authored-by: holdn2 <club.makersfarm@gmail.com>
… Enter-keyup newline

On Windows, the Enter-keyup synthesis path inferred the modified-Enter
chord from release-time modifier state. A plain committing Enter
(Process/229, no modifiers) followed by a rolled-over Shift for the next
doubled consonant made the keyup report shiftKey=true and synthesized a
Shift+Enter the user never chorded; a directly-sent Shift+Enter could
likewise send a second newline from its keyup once the next composition
started. Record observed Enter keydowns (code -> timeStamp) and let the
keyup synthesis run only for presses whose keydown the IME swallowed
entirely; a balancing keyup that copies the keydown timeStamp keeps the
evidence for the later physical release.

Refs stablyai#11878

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A single slot per code let a rapid second Enter press go unguarded: the
first release found a mismatched timeStamp, dropped the only entry, and
the second release then synthesized the Shift+Enter this guard exists to
prevent. Track one entry per press and drain exactly one per physical
release, so every press stays guarded until its own release. Auto-repeat
keydowns do not stack an entry, since the whole run ends in one release,
and the list is bounded so a press whose release never arrives cannot
grow it without end.

Refs stablyai#11878

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adapted from the physical Windows event trace contributed on stablyai#11878.
Backspacing away an entire Pinyin preedit ended the composition with
empty data, no textarea residue, and no input/keypress events — yet
_sendPendingComposition fell back to the last non-empty
compositionupdate data and typed its first character into the PTY.
Only trust that fallback when observed input evidence corroborates it;
a composition with no evidence in any channel was cancelled.

Fixes the macOS Pinyin regression from stablyai#11293 (stray letter left after
deleting a preedit); same fix covers IBus/fcitx Backspace cancellation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Bound exclusive host navigation to a generation-aware latest-wins
single-flight so bulk open and switch fan-out stay responsive on large
remote fleets. Add freeze repro harnesses and navigated settlement.
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
Co-authored-by: yoke233 <yoke2012@gmail.com>
…ablyai#12278)

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
Co-authored-by: Hyunggyun Lyou <hg.lyou@miraeasset.com>
)

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
Co-authored-by: fengxinzi1814 <55725821+fengxinzi1814@users.noreply.github.com>
… snapshots (stablyai#12207)

* test(runtime): prove mobile session publication rebuilds every worktree

buildMobileSessionTabSnapshots consults its per-worktree cache after building
the content, so a republish saves the fanout but none of the work. With 300
worktrees, an unchanged republish still does 601 units of per-worktree work,
and a single changed worktree does 602.

Publication is keyed on agentStatusByPaneKey/agentStatusEpoch, so this runs on
every agent status tick. On a multi-client runtime host with 381 worktrees this
allocated ~350 MB/min and rode the renderer into repeated 4 GB OOMs.

Tests are marked it.fails so the branch stays green; drop .fails when the build
loop skips worktrees whose inputs are unchanged.

* refactor(runtime): make mobile session snapshot inputs explicit per worktree

Every per-worktree builder in buildMobileSessionTabSnapshots took the whole
AppState, so a worktree's real input set was the transitive closure of seven
helpers and could not be memoized safely. Introduce MobileSessionWorktreeInputs
— built once per worktree — and thread it through the group projection and the
terminal/markdown/file/browser tab builders so the compiler proves the input
set. Tab- and pane-keyed slices are narrowed to this worktree's tab ids, file
ids, browser workspace/page ids, and pane keys; agent statuses are bucketed per
worktree once per publication via a tab-id index.

No behavior change. Dropping AppState from the projection path also removes the
second per-worktree read of browserTabsByWorktree, so the publication-cost
counter falls from 601 to 1 per publication and its two cases now pass.

* fix(runtime): skip unchanged worktrees before building mobile session content

buildMobileSessionTabSnapshots consulted its per-worktree cache only after
building that worktree's three Maps, group projection, and full tab array, so
the cache suppressed the fanout but none of the computation. Every agent-status
tick therefore rebuilt every worktree, which drove sustained 4 GB renderer
working sets on a host holding 381 worktrees.

Cache MobileSessionWorktreeInputs alongside each snapshot and reuse the snapshot
when every input field is reference-equal, before any intermediate structure is
allocated. Worktrees with a mounted TerminalPane always rebuild: their live
DOM/PaneManager state is invisible to store references. Absent per-worktree
slices now resolve to shared empty values so an empty worktree can compare equal
to its last publication. jsonContentEquals stays as the backstop on the rebuild
path for inputs that churn by reference without changing output.

With 300 worktrees, per-worktree content builds go from 300 to 0 on an unchanged
republish and from 300 to 1 when one worktree changes.

* test(runtime): cover agent status publication cost

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
* fix(terminal): recover degraded daemon spawn routing

* fix(terminal): preserve fresh-session recovery semantics

* fix(terminal): avoid retaining exited recovery sessions
… runtime (stablyai#9541)

* fix(sidebar): route local folder adds to the intended host, not the global runtime

Adding a local folder while connected to a remote runtime failed with
"<path> was checked on <host>, but that host did not report a usable folder"
because addRepoPath decides local-vs-remote purely from the global
settings.activeRuntimeEnvironmentId when no explicit host is passed.

Two local-add flows relied on that global fallback and got misrouted:

- useAddRepoLocalFolderFlow (native picker / drag-drop): the Add Project
  host selector can display "Local" (selectedRuntimeEnvironmentId = null, so
  the guard passes) while the global still points at an unavailable runtime.
  Native-picked/dropped paths are always local, so force local routing.

- AddProjectFromFolderDialog ("Add folder as project" on a subfolder): a
  subfolder lives on the active repo's host, so carry that host through the
  modal data and route by it — local for local projects, the owning runtime
  for runtime projects — instead of the globally-active runtime.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sidebar): route runtime server-path adds by selected runtime, not global

The Add Project "server path" step (reached only when a runtime host is
selected) called addRepoPath(path, kind) with no explicit host, so it
inherited the global settings.activeRuntimeEnvironmentId. When that global
diverged from the dialog's selected runtime, the add was misrouted off the
host the user picked — the same root cause as stablyai#9541, opposite direction.

Route the server-path add by the dialog's selected runtime explicitly.
Co-locate selectedRuntimeEnvironmentId in useAddRepoHostSelection next to
selectedSshTargetId so the dialog reads it from one place.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(sidebar): route the pre-add git server-path scan by the selected runtime

For kind === 'git', scanNestedRepos runs before addRepoPath and can
early-exit the flow into the nested-repo review, but it still routed by the
global active runtime — so it could scan the wrong host even after the add
itself was correctly routed to the selected runtime (CodeRabbit).

- scanNestedRepos accepts an optional runtimeEnvironmentId in its controls;
  when present it routes by that host, else falls back to the global
  (existing callers unchanged).
- useAddRepoServerPathFlow passes the selected runtime into the scan and
  derives runtimeKind/streaming support from it instead of the global-reading
  getNestedRepoRuntimeKind(null), so telemetry and the nested review target
  the same host as the add.

Adds renderer- and store-level regression tests covering scan routing to the
selected runtime, the null-override local case, and the nested-review handoff.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(sidebar): make nested-scan cancellation route by the scan's owning host

Follow-up to CodeRabbit review of the scan-routing change:

1. scanNestedRepos treated `{ runtimeEnvironmentId: undefined }` as an explicit
   local override via the `in` check. Only null or a string is now explicit;
   undefined falls back to the global (matches getAddRepoPathRouteSettings).

2. scanNestedRepos gained a routing override but cancelNestedRepoScan still
   routed by the global — an asymmetric contract where an override-routed scan
   could be un-cancellable if the global diverged mid-scan. cancelNestedRepoScan
   now takes the same override, and useAddRepoNestedReviewState remembers each
   scan's owning host by scanId (set when the scan is registered) so both stop
   and reset cancel on the host the scan actually ran on. The local folder flow
   routes its scan explicitly local so scan, cancel, and add all agree.

Adds store-level regression tests (explicit override wins, undefined falls back
to global, cancel routes by override) and a new useAddRepoNestedReviewState test
covering cancel-by-owning-runtime for stop and reset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(sidebar): keep subfolder adds on their owning host

* fix(onboarding): keep completion on captured host

* chore(review): drop unreachable onboarding recovery

* fix(sidebar): preserve paired runtime checkout ownership

* fix(runtime): index paired worktrees by logical owner

* fix: fail closed on worktree owner alias collisions

* docs(sidebar): clarify host-routing intent flagged in review

Two Greptile P2 notes, addressed as comments (no behavior change):
- project-added-default-checkout.ts: the runtime branch's `hostId === executionHostId`
  is NOT unreachable — a colliding repo id can carry a runtime-qualified hostId with
  no runtimeOwnerEnvironmentId (see the "repo IDs collide" test). Documented why the
  comparison is reachable and load-bearing rather than replacing it.
- AddProjectFromFolderDialog.tsx: note that omitting the runtimeEnvironmentId spread
  intentionally signals local (NonGitFolderDialog coerces absence to null).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(activity): control portal readiness observer delivery

---------

Co-authored-by: fanyunqian.1 <fanyunqian.1@bytedance.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
…2370) (stablyai#11956)

The runtime RPC WebSocket listener bound to 0.0.0.0:6769 at startup, so a
desktop with no paired device was reachable from the whole LAN before the
user opted in. Default the bind to 127.0.0.1 and widen to all interfaces
only on an explicit opt-in:

- createMobilePairingOffer / getRuntimePairingUrl widen (ensureNetworkExposure)
  before advertising a LAN endpoint; the rebind reuses the resolved port so an
  already-issued offer stays valid, and concurrent offers share one rebind.
- orca serve and E2E set exposeNetworkByDefault to bind wide at startup.
- A previously-connected device (lastSeenAt > 0) rebinds wide at startup so
  reconnect after restart keeps working; a pending/never-connected offer does
  not persist exposure across a restart.

The advertised pairing endpoint still resolves to a concrete interface address,
never the 0.0.0.0 bind host.
…ion (stablyai#12340)

The cell stamps invite expiry at exactly now+10min from its own clock while
the desktop rejected anything past now+10min from the local clock with zero
tolerance, so any cell clock ahead of the machine by more than network
transit made every Relay pairing code fail with an opaque toast. Same
defect class as the host-proof freshness incident; same 30s leeway.

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
AmethystLiang and others added 26 commits August 5, 2026 15:29
* Reorder source control to show staged changes first by default

Stages are closest to the commit action and most relevant to the
commit workflow. Merges untracked files into Changes visually while
preserving their Git area. Removes the untracked-first preset and
includes migration logic for existing user settings.

* Drop source control group order user preference

Remove the sourceControlGroupOrder setting and related UI, migrations, and persistence logic. The source control view now always displays sections in the order: staged changes, unstaged changes, untracked files.

* Reorder source control to show changes before staged

Aligns with the edit-stage-commit workflow by showing unstaged
changes (active edits) before staged changes (queued for commit).
…12646)

* Display SSH worktrees immediately using persisted metadata

Users can now see known worktrees for SSH hosts without waiting for the
provider connection to establish. Worktrees are fetched from local metadata
and displayed as non-authoritative, then merged without replacing richer
live data once the provider becomes available.

* Show SSH folder workspaces immediately via persisted metadata

Add safeguards for metadata fallback: track authoritatively removed
worktrees per host to prevent resurrection, position new rows within
the host block to avoid jumping on authoritative scan arrival, and
preserve co-owner detection status during merge. Coalesce concurrent
metadata fetches to dedupe overlapping queries.
Add a new census module that tracks pending and retained OSC sequences
across all active PTY output processors. Each processor registers a gauge
at creation and unregisters it on dispose, detach, or destroy — this
prevents retained gauges from inflating later heap high-water profiles
and allows the memory profiler to detect stalled processors as a sign of
leaks.
…isting surface (stablyai#11576)

* fix(file-explorer): sort numbered file names naturally

The File Explorer compared names with bare localeCompare, so numbered
files listed 100, 200 before 99. Hoist the numeric collator Source
Control file rows already use (stablyai#10850) into src/shared and apply it to
the local and runtime directory listings, the name-filtered view, and
Source Control directory nodes, which were inconsistent with the file
rows one line below (stablyai#11426).

* fix(file-explorer): natural sort on SSH funnels, relay, and pickers

Adversarial-review round 1 rework:
- Both readDir funnels short-circuited to the SSH filesystem provider
  before the patched sort, so SSH workspaces kept lexicographic order;
  re-sort locally after the provider returns (the remote relay may be an
  older build), and fix the relay's own comparator for relay-native
  consumers.
- sortDirEntries (shared, unit-tested) owns the directories-first +
  natural-order listing contract used by every funnel.
- compareFileNames breaks numeric-collation ties ('2' vs '02') by code
  units so sibling order stays total instead of readdir order, and pins
  the collator locale to 'en' so every host produces one order.
- The SSH folder browser and runtime server dir picker now match the
  Explorer they browse into.
- Ordering pinned by tests at the relay, source-control tree, and shared
  helper.

* fix(mobile): natural sort in the mobile file explorer

Mobile re-sorted host readDir results with bare localeCompare, undoing
the host funnel's natural order (round-2 review). Reuse the shared
comparator and pin the order in the mobile suite.

* fix(file-explorer): natural sort at the renderer choke point and remaining ties

Round-3 review: the remote-runtime RPC and paired-web routes return the
host's order verbatim, so re-sort in readFileExplorerDirectory where
every desktop route converges; pin the SSH funnel with a handler-level
test; and route Source Control path compares through compareFileNames so
numeric-collation ties share one total order with the Explorer.

* docs(file-name-sort): state the real perf baseline in the hoist comment

* refactor(source-control): drop the dead collator export; pin the test oracle locale

* fix(file-listings): cover remaining natural-sort surfaces
…ve (stablyai#12791)

`orca serve` publishes a ready graph under HEADLESS_RUNTIME_WINDOW_ID with no
BrowserWindow behind it. `shouldCreateInBackground` only degraded when the
create was renderer-backed, so any focus-requested create fell through to
getAuthoritativeWindow() and threw "No renderer window available" — leaving
`terminal create --focus` with no workaround on a remote server (stablyai#10333).

With a worktree selector and no renderer window, a background spawn is the only
usable path, so collapse the renderer-backed window check into a plain
"no window" check. That is the existing rendererBacked clause plus exactly the
missing focus case, and it drops the confusing `rendererWindow === null`
indirection (rendererWindow is already gated on rendererBacked).

Focus is not lost by the degrade: the spawned pane is still published to the
session-tab model and revealed with `activate: true`, which is how a paired
client learns about it. Mirrors the in-tree precedent in
runCreateMobileSessionTerminal.

Headed hosts are unaffected — the clause only fires when no window exists.

Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
…ablyai#12495)

* fix(mobile): keep the cached transcript visible while reconnecting

A manual retry closes the client and opens a fresh one, so the chat session
hook saw a new client under an unchanged identity, dropped its settled read,
and handed out an empty list — the transcript collapsed to a full-screen
spinner until the swapped client's snapshot landed.

Hold the last settled list per identity (captured post-commit) and keep
rendering it while the re-read is in flight. `transcriptLoading` still gates
consumers that decide from an empty transcript, so the launch-draft seed is
unaffected. The held list is keyed by a new `sourceIdentity` (host/workspace)
in addition to agent/session/transcript, so it can never serve another
source's messages.

Refs STA-3333.

* test(mobile): assert the whole reconnect window, not just its first frame

The re-subscribe lands a commit after the first render of the swap, so a
regression that cleared the held list there left frame 0 green and still
blanked the transcript. Verified: clearing the cache in the subscribe
cleanup now fails this test, where before only the view-toggle test caught it.

* fix(mobile): don't derive a tappable ask card from the held transcript

The cache this PR adds keeps the previous list rendered while a swapped
client re-reads. useMobileNativeChatPrompts was the one consumer reading
`messages` without honouring `transcriptLoading`, so an ask answered on
the terminal resurrected as a live, tappable card during that window.

Gating on `transcriptLoading` is exactly base behaviour: `setRead` only
ever stores 'ready'/'error', so status==='loading' implied an empty list
before this PR. The live `askFromStatus` path is untouched.

* chore: keep merge formatting scoped
…ed socket (stablyai#12790)

The paired-runtime WS heartbeat terminated a client after a single unanswered
15s ping. One missed pong is UNKNOWN, not proof the peer is gone: a cellular or
Tailscale blackhole, or a stalled TCP retransmit, routinely swallows one pong
from a peer that is still there. Users on flaky paths saw constant drops, each
costing a full redial plus E2EE re-handshake and subscription replay.

Reap now needs MISSED_PROBE_LIMIT (3) consecutive unanswered probes, counted per
socket rather than timed. Any proof of life -- pong or any inbound frame -- clears
the count, as does a resume from a server-loop pause, since a gap the client was
never given a chance to answer must not top up its budget. Missed sweeps still
re-probe, so a recovered path proves itself on the next tick.

Three matches the liveness budgets already in the product: the web client gives
45s (25s idle + 20s probe grace) and the relay control gives 75s. The paired
transport's single miss was the outlier.

Also gives the web client's redial the one-sided jitter the shared-control path
already had, so a fleet dropped by one shared blip does not re-dial in lockstep;
the helper is extracted to src/shared/reconnect-jitter.ts and shared by both.

STA-3320, stablyai#12327

Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
…ablyai#12815)

stablyai#12793 fixed PR status being hidden by workspace activity by relocating
it: prDisplay was dropped from WorktreeCardStatusSlot and re-rendered as
WorktreeCardReviewStatus in the title-row indicator group, at the right
edge of the card.

Reverting restores the review glyph to the left status lane. Because
stablyai#12658 still narrows the passive-identity set to {inactive}, this alone
brings PR status back only for inactive workspaces; reverting stablyai#12658
widens it to active and done.

No behavior change for branch identity, and stablyai#8813's guard stays intact.
…oo (stablyai#12825)

Stacked on the stablyai#12793 revert. Widens the passive-identity set from
{inactive} back to {active, done, inactive}, so the PR/check glyph
returns to the left status lane for workspaces that are actively being
worked, not just idle ones.

Tradeoff, deliberate: stablyai#12658 was not purely a regression. It also fixed
stablyai#8813, where an active workspace with branch identity and no PR showed
the grey branch glyph instead of the emerald Active dot. This revert
reintroduces that, and removes its e2e guard.

The left lane holds one glyph, so activity, branch identity, and review
status cannot all be shown. This picks review status.
…stablyai#12841)

Post-merge review of stablyai#12790 demonstrated a real leak: the resume-from-pause
pardon cleared banked misses outright, so a host whose sweep stalls once every
three ticks reset the budget forever and a dead socket was never reaped. The
reviewer ran 300 sweeps against a permanently dead peer with a >1.5x gap every
third tick and observed zero terminate() calls.

Pre-stablyai#12790 that required a stall on *every* tick; the counter widened the
pathological window 3x, and the failure mode is permanent non-reaping — the
MAX_WS_CONNECTIONS leak the reaper exists to prevent.

A stalled tick now charges no miss, which is all the original rationale needed
(the client had no chance to answer that probe), but no longer forgives the
misses already banked. A live client still clears its own count by answering
the probe that is still sent on the stalled tick.

The tolerance test's pause case is rewritten to assert the new contract rather
than the old forgive-everything one, and a new test pins the leak directly: a
host stalling every third tick must still reap a dead socket.

Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
…tablyai#12776)

* fix(terminal): stop transient probe blips from erroring restored panes (STA-3536)

terminal_pane_owner_unverified fired for every restored pane whenever one
liveness probe answer went missing: a cold-start daemon draining an attach
stampede misses the 2s getSize deadline, and a wedged superseded daemon
(protocol upgrades leave them running) turns every unmapped fan-out probe
null forever.

- probePtyOwners now skips legacy daemons whose startup inventory listing
  succeeded: fresh sessions never route to them, so they provably don't own
  an unmapped id and one wedged zombie can't poison every pane's verdict.
- attachStablePaneOwner retries the probe over a short backoff ladder before
  surfacing unverified, so a single missed deadline resolves to a verdict.
- The renderer replaces the raw error code with actionable copy.

* fix(terminal): stop retrying definitive owner probes

* fix(terminal): recover live panes after renderer restart

* refactor(terminal): share owner resolution abort guard
…blyai#12775)

* fix(mobile): keep hosts visible when credentials are unavailable

* fix(mobile): guard unavailable host recovery

* fix(mobile): protect replacement credential writes

* fix(mobile): retain superseded cleanup intents

* fix(mobile): preserve credential cleanup authority

* fix(mobile): make host cleanup crash-safe
* fix(terminal): fence daemon endpoint ownership

* fix(terminal): clean failed daemon PID claims

* fix(terminal): close daemon ownership review gaps

* test(daemon): release startup IPC in boot smoke

* test(daemon): mirror production stdio in boot smoke

* fix(daemon): exit after rpc shutdown cleanup

* fix(terminal): make the socket name the daemon endpoint authority

The reported failure was a live daemon hosting PTYs that nothing could
reach: terminals acknowledged input and never ran it, listings diverged
from reality, and restarting the app never helped because the detached
helper survived. The ownership fence added for it could not fire in the
sequence that produces the split brain.

libuv unlinks the pathname a server bound to when that server closes,
with no ownership check. A daemon that lost its endpoint name therefore
deleted whichever socket then sat at that path — including a live
replacement's — stranding a daemon that still hosted every session.
Bind a private same-directory name and hard-link it into place instead:
libuv can only ever unlink our own bind name, the exclusive link is a
kernel-enforced endpoint claim, and the canonical name is removed only
under an inode ownership check. The bind name replaces the basename
rather than extending it, so it cannot overflow sun_path.

killStaleDaemon removed the PID record unconditionally immediately
before every fork, so the exclusive PID claim was always uncontested at
bind time. It also unlinked a live daemon's endpoint whenever a connect
probe merely timed out, and treated a `ps` timeout as proof of PID
recycling. Now only positive evidence of a dead endpoint authorizes
reclaiming it, SIGKILL is confirmed rather than assumed, and a daemon
that cannot be proven stopped keeps its record and endpoint while the
launcher refuses to fork beside it.

A daemon whose endpoint was taken over now retires itself, draining
rather than killing, so an unreachable orphan stops being permanent.

A repaired PID record re-derives entryPath, appVersion and the Linux
incarnation markers from the authenticated owner instead of dropping
them; without appVersion a healthy daemon read as a permanently stale
bundle and, on Windows, went unpinned against daemon-host pruning.
Repair failure now fails open — abandoning a healthy daemon over a pid
file write cost every persistent terminal on the machine.

Also: treat only ENOENT as an unclaimed record so a Windows file lock is
not reported as an ownership conflict; settle start() before close() so
an accepted connection cannot defer it forever; sweep abandoned claim
and bind names; and type the endpoint-identity seam so a rename cannot
silently disable the fence.

Adds a real-process handover smoke that reproduces the failure with two
daemons racing one endpoint, and wires it into the native-smoke job.

* fix(daemon): retire only on proven endpoint ownership loss

The ownership watchdog read a null identity for any stat failure, so a
transient EACCES or EIO on the runtime directory would retire a daemon
that was still serving every terminal on the machine. Distinguish "the
entry is gone" from "the probe failed" and act only on the former.

Also require the loss to persist across two polls: a replacement
publishes by unlink-then-link, and a single observation can land in that
gap.

* fix(daemon): source repaired ownership metadata from the authenticated hello

Adversarial review found three defects in the previous two commits.

Re-deriving entryPath from the owner's command line truncated it at the
first space. A command line is a single space-joined string, so
`C:\Program Files\Orca\...` and `/Applications/Orca 2.app/...` came back
as `"C:\Program` and `/Applications/Orca`. getDaemonLaunchIdentity treats
a present entryPath as authoritative, so a healthy daemon read as
`different_app_path` and was killed and re-forked — worse than the
missing-metadata case the derivation was added to fix. Carry entryPath
and appVersion as optional fields on the daemon hello identity instead:
the daemon already has both from its own argv, and per
docs/reference/remote-wire-compatibility.md a new optional field is safe
because every reader falls back when it is absent. This also removes a
synchronous `ps` spawn from the Electron main thread during startup.

`start()` rolled back the PID record even when it never published one.
Losing the endpoint link now runs that path, and the ownership-checked
unlink briefly renames the incumbent's record aside — enough to strand a
live daemon's ownership. Roll back only what we actually wrote.

publishDaemonSocketPath read its identity from the canonical name after
linking, so a concurrent unlink returned null: no ownership watchdog and
no endpoint cleanup on any shutdown path. Read it from the bound name
before linking, which shares the inode.

Refusing to fork beside an unconfirmed daemon left the user with no
daemon at all and no in-app recovery, since restart re-entered the same
fence. We have just proved something answers the endpoint, so adopt it
in degraded mode: live sessions keep working, fresh terminals run
locally. SIGTERM is also individually guarded now — an EPERM fell into
the blanket catch and reported "nothing alive", authorizing the very
duplicate this fence exists to prevent.

Also reset the ownership-loss streak on an inconclusive probe so the
confirmations are consecutive, and sweep scratch names before the launch
so a failed launch still reclaims them.
…rame (stablyai#12860)

* fix(native-chat): locate Claude's model row by frame structure

The scraper assumed the model descriptor sits within three rows of the
`Claude Code vX` line. It does not: Claude prints it near the bottom of the
startup frame with the welcome art and release-notes panel in between — eight
rows down at 100 columns. Narrow panes degrade the frame further, dropping the
version from the title row entirely below ~70 columns and wrapping the billing
tail onto its own row. Any one of those made the scrape return null, so the
model picker showed no current selection at all.

Search the frame from its bottom border upward for the row carrying model
metadata, read only the leftmost frame cell so release-notes prose can never
win, accept the frame corner as header proof when the version is gone, and
tolerate the effort suffix being elided to an ellipsis. Catalog families now
match as a leading word, which both survives the resolved-name suffix
("Opus 5 (1M context)") and keeps custom slugs like company/my-haiku-v2 from
being claimed as haiku; an unrecognized name is reported as a custom model.

Fixtures are real: captured from a live claude 2.1.220 by replaying the PTY
bytes through @xterm/headless and serializing exactly as TerminalPane does.

* fix(native-chat): resolve the scraped model against the host's real catalog

The scraper matched the static seed while the picker lists what stablyai#12369
discovers from the host CLI, so the two spoke different id spaces. On a current
CLI `list_models` returns `opus[1m]`, `sonnet`, `sonnet[1m]`, `fable` and
`haiku` — no plain `opus` — while the seed only knows families. Reporting
`opus` therefore selected a row the picker had to invent, dropping the host's
own effort and fast-mode descriptors with it. Locating the model row correctly
made this the normal case rather than a rarity, since the scrape now succeeds.

Resolve against the discovered list first, falling back to the seed for aliases
a host no longer lists and to the raw name for genuinely custom models. Matching
requires the family to lead as a whole word and the label's remaining tokens to
appear in order, so `Opus 5 (1M context)` picks `opus[1m]`, plain `Sonnet 5`
keeps `sonnet` instead of being captured by the 1M-context row, and
`opus-internal-v3` stays custom. Most specific label wins.

The hook keeps the screen that parsed so a discovery landing after the first
read re-resolves it, rather than stranding a family id once the frame has
scrolled out of the buffer.

* fix(native-chat): identify option-less models on narrow panes

Live capture at 60 columns: a Haiku session prints a bare `Haiku 4.5` row with
its billing wrapped to the next line. No middot, no effort suffix — nothing
marks it as the model, so it reported nothing. Claude always closes the frame
with the working directory and prints at most the descriptor plus a wrapped
billing line above it, so fall back to walking up from there when no row
carries descriptor metadata. The height bound is what keeps the walk from
climbing into the welcome art.

* test(native-chat): pin re-resolution when discovery lands after the read

Covers the wiring the parser tests cannot reach: the frame is visible at mount
and gone by the time the host's model list arrives, so only the cached screen
can drive the second resolve. Fails against a listener that merely replaces the
models.

* fix(native-chat): prevent stale Claude model reports
…row at SessionStart (STA-3386) (stablyai#12859)

* fix(agent-hooks): give resumed Claude sessions a sidebar row at SessionStart (STA-3386)

Claude's hook set never registered SessionStart and normalizeClaudeEvent
dropped it at ingest, so a resumed session that idled produced zero hook
traffic and earned no sidebar agent row until the first prompt.

- Register SessionStart in CLAUDE_EVENTS (local + remote installs).
- Map lead SessionStart (startup/resume/clear) to an idle 'done' row,
  resetting stale roster/task/cron/tool/prompt state like the Codex path;
  compact restarts and child-attributed SessionStart stay dropped.
- Thread hookEventName through the agent-status IPC payload so the
  completion coordinator can tell a session connect from a turn result;
  a SessionStart 'done' no longer raises agent-task-complete.

* fix(agent-hooks): mark SessionStart rows as session boundaries, not completions (STA-3386)

Review follow-up: represent the idle connect as a first-class
sessionBoundary flag on the status payload instead of gating one
renderer consumer on hookEventName.

- sessionBoundary rides AgentStatusPayload/AgentStatusEntry (done-only,
  clamped like interrupted); drops the hookEventName IPC threading.
- Completion-reactive consumers ignore session boundaries: the
  completion coordinator (task-complete notifications), automation
  dispatch observers (a connecting agent no longer completes the run
  and closes its tab), activity unread counts, and the dashboard
  finished timestamp; the status slice keeps boundaries out of
  stateHistory and preserves the flag across done->done repaints.
- SessionStart sources are allowlisted (startup/resume/clear) so
  compact restarts or unknown sources fail closed mid-turn.
- A live SessionStart now un-retires a reusable pane like a fresh
  prompt, so resume-in-reused-pane earns its row too.

* fix(agent-hooks): keep session-boundary dones out of teardown and completion history (STA-3386)

Review round 2:
- A boundary done no longer deletes the pane's launch-config registry
  entry, so a resumed idle TUI keeps its registered-launch-agent
  identity evidence.
- A boundary landing on a REAL done pushes that completion into
  stateHistory so the finished timestamp and unread badge survive a
  resume//clear right after a finish.
- The done->done flag carry yields to turn evidence (assistant message
  or changed prompt) so a genuine completion can never be suppressed.
- Star-nag value-moment observer and the server's OSC-equivalence
  dedupe now discriminate the flag.

* fix(agent-hooks): keep a displaced completion unread in the sidebar badge (STA-3386)

Review round 3: sidebar-badge mode counts only the live entry, so a
session boundary landing on an unacknowledged completion silently
dropped the sidebar badge while the agent-events count kept it. Count
the displaced completion from history for boundary rows, and pin the
behavior with countActivityUnread tests.

* fix(agent-hooks): prevent SessionStart completion side effects (STA-3386)

* fix(agent-hooks): preserve SessionStart through renderer IPC (STA-3386)
* fix(mobile): bound home host auto-connect fanout

* fix(mobile): clarify bounded host connection state

* fix(mobile): close unused settings host clients

* fix(mobile): close released host clients

* fix(mobile): sync host release policy in effect

* fix(mobile): preserve focused host client ownership

* test(mobile): cover manual Home client release

* Fix React Doctor array type check
…(STA-3433) (stablyai#12839)

Mouse events posted with CGEventPostToPid reach the target app with no
window association, so AppKit never routes the press to a view: hover
states fire but the control is never activated, and the mouseUp is
dropped outright when posted back-to-back. Post click events to the HID
event tap instead (as keyboard synthesis already does), pace them, and
stamp mouseEventClickState so multi-clicks register.

Synthetic clicks now also report verification unverified/synthetic_input
from the helper itself, matching the other synthetic actions.
…efetches (STA-3343) (stablyai#12830)

* fix(tasks): hold dialog-confirmed issue state over stale list refetches (STA-3343)

Closing an issue from GitHubItemDialog patched workItemsCache directly
with no mutation-registry record, so a search-lagged Tasks refetch
(GitHub search index eventual consistency + gh's ~120s URL cache)
silently reverted the row to Open. Record the confirmed state as
registry authority (same mechanism the list-row mutations use) so list
fetch paths re-assert it until search catches up; quiet adopt still
releases it on match, so external reverts win once the index is fresh.

Covers issue close/reopen, PR close/reopen, and PR merge in the dialog.

* fix(tasks): preserve newer state authority on rollback
… (STA-3448) (stablyai#12831)

* fix(browser): recover browser tab when guest WebContents is destroyed (STA-3448)

A <webview> whose guest WebContents died without render-process-gone
(detach/reattach race, guest-side close) stayed attached and painted
black forever; reload and focus on the dead guest threw uncaught.

- listen for the webview 'destroyed' event at both layers, mirroring
  render-process-gone: registry marks recovery pending (survives pane
  unmount), BrowserPane triggers guest recovery immediately
- route reload-on-dead-guest into guest recovery instead of throwing
  (toolbar, Cmd+R renderer+IPC paths, context menu)
- guard webview.focus() against Electron's null-internals throw

* fix(browser): close guest reload destruction race
…(STA-3442) (stablyai#12838)

httpProxyUrl is the only network setting stored via safeStorage. Two
failure modes silently killed the configured proxy on macOS:

- A keychain reset/denial makes decryptString throw at load; the raw
  ciphertext then masqueraded as the configured proxy URL, so
  applyElectronProxySettings silently fell back to DIRECT and the
  garbage re-persisted forever (no self-heal).
- safeStorage.isEncryptionAvailable() throwing (keychain/API errors,
  pre-ready use) was uncaught in encrypt/decrypt, failing the entire
  state save - the data file never gained the proxy keys at all.

Load now validates the decrypted value and clears undecryptable
ciphertext (plaintext URLs still pass, preserving the pre-encryption
upgrade path), the availability check is exception-safe, and startup
logs when persisted proxy settings are invalid instead of silently
using direct networking.
…yai#12812)

* fix(skills): match official skill files despite local sidecars

Scope known-snapshot matching to manifest-listed files so agent-written
sidecars (e.g. agents/openai.yaml) no longer mark a package unrecognized
and block updates when official bytes still match.

Preserves fail-closed detection when a listed file's content drifts.

Fixes stablyai#12694

* fix(skills): scope lock trust and convergence to official files too

Sidecar tolerance stopped at the snapshot match, leaving three disk-vs-official
comparisons still judging the whole folder.

The lock-comparable hash covered every observed file, so a clean update beside
agents/openai.yaml reported as failed and read 'may be modified'. It is now
carried both whole and scoped to the current bundle's paths, and either may
satisfy the lock: the sidecar case only ever matches scoped, while an upstream
revision that ADDS a file only ever matches whole, so publishing one alone
would trade this bug for stablyai#11220.

Convergence re-derived the disk revision from that same whole-folder digest,
which no revision matches once a sidecar lands, retiring the stuck-lock gate
and arming an update the command provably cannot perform; it now honours the
revision observation already resolved.

Subset matching also let an older revision launder drift on a file the current
bundle lists, since that revision does not list it and so read it as a
neighbour. Identity now keys tolerance on what the current bundle owns.

---------

Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
* feat(dashboard): add experimental agent map view

* fix(dashboard): harden agent map behavior

* fix(dashboard): harden agent map recovery

* fix(dashboard): close map selection on view change

* fix(agent-map): center sparse layouts

* fix(agent-map): align completion and workspace actions

* Polish agent map interactions and repo labels

* feat(agent-map): add worktree lineage and project actions

* fix(agent-map): use marker for unread agents

* fix(agent-map): compact orchestrated families

* fix(dashboard): harden agent map actions and layout

* fix(agent-map): bound layout work and preserve interactions

* fix(agent-map): move unread marker to ring top-right

* fix(agent-map): seat unread marker on the ring's top-left edge

* feat(agent-map): restore the agent launcher and declutter map labels

Three gaps in the experimental Agent Map:

- The "start a new agent" picker was split onto a preserved branch during the
  08-02 rebase (47829cb) and never re-landed. Restores that commit and its
  pop-out IPC, keyed on the raw worktree id rather than the map identity.
- Workspace labels draw at a fixed screen size with no collision handling, so a
  zoomed-out map stacked dozens of names on each other. Adds a declutter pass
  that seats project names first, then workspace names by attention, then
  project counts in whatever room is left.
- The pop-out had no workspace right-click at all: its renderer has no store, so
  the shared sidebar menu cannot mount there. Adds a snapshot-driven menu with
  the launcher and Sleep, relayed to the main renderer.

* refactor(agent-map): fold the map's filter rail into the shared toolbar filter

The rail duplicated the toolbar's project filter and cost the canvas 14rem of
width on the surface that needs it most. Agent states move into the toolbar's
Filter dropdown (map view only — the board's columns already separate them) and
count toward its badge; project filtering falls back to the toolbar's own. Show
all is the dropdown's Clear all, and Fit already lives in the viewport controls.

* fix(agent-map): isolate map work from main renderer

* perf(agent-map): stream status updates to popout

* fix(i18n): add agent map catalog entries

* feat(agent-map): glow working entities

* fix(agent-map): prioritize attention ring status

* fix(agent-map): distinguish subagent connectors
Register reasonix for detect/launch, catalog, icons, and process title
recognition so the DeepSeek-native coding agent can be launched from Orca.

Preserves: existing agent config contracts and prompt injection modes for
other agents. Interactive launch uses stdin-after-start (print/run exit).

Evidence: vitest agent-kind, detection, startup, selection suites.

Fixes stablyai#12200
@innocarpe innocarpe added the enhancement New feature or request label Aug 6, 2026
@innocarpe innocarpe closed this Aug 6, 2026
@innocarpe
innocarpe deleted the fix/reasonix-tui-agent branch August 6, 2026 06:21
@innocarpe

Copy link
Copy Markdown
Owner Author

Closing portfolio mirror: upstream stablyai#12874 closed in favor of stablyai#12330 (first mover on Reasonix agent registration).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.