fix(daemon): restore previously recoverable scrollback from durable history - #14193
Conversation
…istory Keep the 1000-row live daemon window so session count cannot grow the grid without bound. Rebuild, remount, restart, and reattach now reconstruct the desktop 5000-row depth from durable history instead of compacting that live window. STA-4091
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (2)
🚧 Files skipped from review as they are similar to previous changes (2)
📝 WalkthroughWalkthroughThe daemon reconstructs bounded terminal snapshots from cold-restore data, durable history, and pending records. It preserves terminal metadata, output sequencing, parser state, and frame-restore data. Checkpoint handling tracks live and continuity requirements and supports fallback and retry states. Reattach and remount flows apply durable-history overlays. Tests cover restoration depth, replay, compaction, sequencing, fallback, and checkpoint lifecycle behavior. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🧹 Nitpick comments (1)
src/main/daemon/daemon-restore-scrollback-depth.test.ts (1)
120-128: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winAssert that
live.writeSyncsucceeded before the negative containment check.Line 126 ignores the
writeSyncreturn value. Line 128 is a negative assertion. IfwriteSyncfails emulator-wide, the snapshot is empty and line 128 passes without proving that the live window evictedPREVIOUSLY_RECOVERABLE_LINE.The positive assertions in the sibling tests at lines 168-170 cannot pass vacuously, so only this arm needs the guard.
Based on learnings: "
writeLargeScrollbackmust assert that allHeadlessEmulator.writeSynccalls succeeded. Otherwise, theendedAtcold-restore eligibility test can pass against an empty 416-byte checkpoint".♻️ Proposed fix
- live.writeSync(numberedOutput(DESKTOP_TERMINAL_SCROLLBACK_ROWS_DEFAULT)) + expect(live.writeSync(numberedOutput(DESKTOP_TERMINAL_SCROLLBACK_ROWS_DEFAULT))).toBe(true) const liveSnapshot = live.getSnapshot() expect(snapshotText(liveSnapshot)).not.toContain(PREVIOUSLY_RECOVERABLE_LINE)Source: Learnings
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro Plus
Run ID: 5bd54c7c-ba0f-4c5a-bde4-510c04604a22
📒 Files selected for processing (11)
src/main/daemon/daemon-durable-history-snapshot.tssrc/main/daemon/daemon-pty-adapter.test.tssrc/main/daemon/daemon-pty-adapter.tssrc/main/daemon/daemon-restore-scrollback-depth.test.tssrc/main/daemon/daemon-restore-scrollback-depth.tssrc/main/daemon/daemon-session-scrollback-window.tssrc/main/daemon/history-reader.tssrc/main/daemon/session.test.tssrc/main/daemon/session.tssrc/main/daemon/terminal-history-cold-restore-info.tssrc/main/daemon/types.ts
| const restoreInfo = await this.historyReader.detectColdRestore(sessionId, { | ||
| ignoreCleanEnd: true, | ||
| wslDistro: this.wslDistrosBySessionId.get(sessionId) | ||
| }) | ||
| if (!restoreInfo) { | ||
| return liveSnapshot | ||
| } |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Return the committed checkpoint snapshot when no durable history is found.
Line 1308 prefers checkpoint.snapshot over liveSnapshot when the checkpoint does not commit. Line 1318 does the opposite after a successful commit: it returns liveSnapshot.
liveSnapshot was captured before takeSnapshotAndCheckpoint drained the pending queue, so it carries an older outputSequence and older bytes than checkpoint.snapshot. A bounded-depth remount on a session with no readable durable history therefore receives a staler snapshot than the one the adapter already committed to disk.
checkpoint.snapshot is the fresher self-consistent pair for this path.
🐛 Proposed fix
if (!restoreInfo) {
- return liveSnapshot
+ return checkpoint.snapshot
}📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| const restoreInfo = await this.historyReader.detectColdRestore(sessionId, { | |
| ignoreCleanEnd: true, | |
| wslDistro: this.wslDistrosBySessionId.get(sessionId) | |
| }) | |
| if (!restoreInfo) { | |
| return liveSnapshot | |
| } | |
| const restoreInfo = await this.historyReader.detectColdRestore(sessionId, { | |
| ignoreCleanEnd: true, | |
| wslDistro: this.wslDistrosBySessionId.get(sessionId) | |
| }) | |
| if (!restoreInfo) { | |
| return checkpoint.snapshot | |
| } |
There was a problem hiding this comment.
Actionable comments posted: 1
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro Plus
Run ID: fc97d10c-543f-4289-bad9-d2cbc6e3fa01
📒 Files selected for processing (9)
src/main/daemon/daemon-checkpoint-file.tssrc/main/daemon/daemon-pty-adapter.test.tssrc/main/daemon/daemon-pty-adapter.tssrc/main/daemon/daemon-restore-scrollback-depth.test.tssrc/main/daemon/history-manager.tssrc/main/daemon/terminal-checkpoint-serializer.tssrc/main/daemon/terminal-history-checkpoint-reader.tssrc/main/daemon/terminal-history-cold-restore-info.tssrc/main/daemon/terminal-history-session-writer.ts
🚧 Files skipped from review as they are similar to previous changes (3)
- src/main/daemon/terminal-history-cold-restore-info.ts
- src/main/daemon/daemon-pty-adapter.test.ts
- src/main/daemon/daemon-restore-scrollback-depth.test.ts
| ...(opts?.pendingOutputSeq !== undefined | ||
| ? { pendingOutputSeq: opts.pendingOutputSeq } | ||
| : {}), |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win
Persist only the sequence included in the checkpoint.
pendingOutputSeq is documented as the last pending-output batch included in the checkpoint. In src/main/daemon/daemon-pty-adapter.ts, the live-snapshot fallback commits this checkpoint before it appends take.records at take.seq. Persisting take.seq in that path marks a post-checkpoint batch as already included. Recovery can skip the tail or reject the log as non-contiguous.
Compute the value from the records actually replayed into the snapshot. Use the prior sequence or omit the field when take.records remains a post-checkpoint log tail. Add a restart regression that verifies the tail is replayed exactly once.
The supplied src/main/daemon/daemon-pty-adapter.ts:2196-2256 flow shows the append occurs after the checkpoint.
…3) (#14346) PR #14193 routed warm reattach and deep-buffer snapshots through overlayDurableRestoreSnapshot, which joined a single process-wide checkpoint tail with no deadline. One never-settling checkpoint blocked every daemon-backed terminal from reattaching. Checkpoint exclusivity only protects one session directory's tmp-write/rename pair, so serialize per session instead. Reattach now waits on a bounded deadline and degrades to the daemon's live window; the abandoned compact keeps running and still commits, so a blown deadline costs restore depth for one reattach, never durable history.
* test(e2e): stabilize triaged release failures (#14242)
* test(e2e): stabilize triaged release failures
* rm file
* feat(worktrees): add per-source visibility controls (#14189)
* feat(worktrees): add per-source visibility controls
* fix(worktrees): explain unsupported visibility hosts
* fix(worktrees): keep add location form inline
* fix(worktrees): align source visibility across runtimes
* test(worktrees): cover Windows drive-relative roots
* Add golden E2E tests for fresh terminal and shell commands (#14302)
* test(e2e): add golden tests for fresh terminal and shell commands
Adds regression tests for terminal initialization in fresh profiles and shell command execution to the golden test suite, integrated across Linux, macOS, and Windows CI.
* test(e2e): bracket shell command output between markers
The echoed command can wrap or be clipped by the buffer tail. Bracket output between begin and end markers to reliably identify real output, and strip ANSI escape sequences that interfere with parsing.
* test: add golden E2E tests for source control workflows (#14260)
* test: add golden E2E tests for source control workflows
- Tests core source control interactions: file edit/save, commit staging, and diff viewing
- Integrated into CI/CD pipelines for Linux, macOS, and Windows
- Includes helper utilities for test setup and worktree management
* test(e2e): verify golden commit author and fix test flakiness
- Configure git author name/email at worktree level during setup
- Verify commits are made with correct author details in assertions
- Add explicit timeouts to file visibility waits and git status polling
- Fix test ordering to seed edits after source control is open
- Simplify git status refresh logic to rely on automatic updates
* Add rollback to createGoldenWorktree on setup failure
Cleanup callbacks only register after setup succeeds. When a config
command fails, the half-built worktree and branch leak into later
test runs, causing flakiness. Now we roll back immediately and
re-throw the setup error.
* test(e2e): match explorer rows after the git status badge appears
The golden file-save spec used an exact /^README.md$/ filter. After save,
the explorer row text becomes "README.md M", so reopen clicked nothing.
* test: strengthen golden worktree setup verification
- Track working directory in git call inspection to verify correct execution context
- Verify user.name/email config applies to worktree-specific settings, not repo
- Add exhaustive setup call sequence assertions to catch setup/rollback leaks
* test(e2e): add golden E2E tests for workspace session management (#14304)
* test(e2e): add golden E2E tests for workspace session management
- Restore exact file and terminal state after quit/relaunch
- Verify terminal file link activation and external edit detection
- Test worktree creation and switching with isolated terminals
- Isolate test repo paths between concurrent CI runs with UUIDs
* Add platform-aware marker echo command utility
- Create splitMarkerEchoCommand() to generate shell commands that
safely echo test markers across Windows and Unix platforms
- Split markers into prefix/suffix fragments so output assertions
prove execution, not just shell echo-back
- Consolidate SORTABLE_TAB export and improve tab bar locator logic
- Refactor terminal link helpers to extract client point calculation
* test: add golden e2e tests for agent TUI launch and shell recovery (#14258)
* test: add golden e2e tests for agent TUI launch and shell recovery
Add test fixtures and E2E tests to verify agent TUI functionality:
- Stub agent implementation supports cross-platform execution (Unix/Windows)
- Test verifies multiline composer with Shift+Enter support in agent TUI
- Test verifies clean shell resumes after agent exit without state leakage
* test: add golden e2e tests for agent TUI launch and shell recovery
Add agent TUI launch and shell-recovery tests to the golden (release-blocking)
E2E suite, covering agent initialization and shell availability after agent
exit. Improve escape sequence handling in the stub agent to prevent stray key
reports from contaminating test output. Add terminal input readiness checks to
ensure commands execute reliably before verification.
* test: coerce golden stub stdin chunks for type-aware lint
Node types the stdin data event as string | Buffer even after
setEncoding('utf8'), so restrict-plus-operands failed CI.
* test: fix golden stub agent Windows batch files and add Ctrl+C support
- Store batch files with CRLF to avoid Windows 512-byte parser boundary bug
- Handle Ctrl+C (0x03) in raw mode as alternative to Ctrl+D (0x04)
- Update release notes documenting golden test skip behavior on older tags
* Remove Windows batch file gitattributes workaround
The -text whitespace=cr-at-eol rule preventing CRLF conversion for
.cmd files is no longer needed. Allow batch files to use normalized
line endings.
* Move worktree labels to right badge rail with truncation tooltips (#14313)
* refactor(palette): move worktree labels to right badge rail with tooltip
Move worktree and branch names from the tab title line to a dedicated right-side badge rail, preventing long titles from being truncated. Add smart tooltips showing full names when truncated, and distinguish between workspace names and branch names based on whether they're auto-generated labels.
* fix(palette): disconnect worktree rail ResizeObserver on unmount
Move truncation observation into a layout effect so React Doctor sees a cleanup path and the subscription cannot leak after unmount.
* Fix mobile HTML report rendering (#14196)
* fix(mobile): render local HTML reports at phone width
* fix(mobile): satisfy browser URL lint
* feat(agent-map): make an unread finish visible and let a glance demote it (#14197)
* feat(agent-map): make an unread finish visible and let a glance demote it
A finished agent was the quietest mark on the map. `status-glow` was defined
only for working/waiting/blocked, so a finish got a 1.7px emerald stroke on a
6px mark and no halo — invisible across a 200-agent fleet.
Worse, `dashboardCardDisplayState` folds `done && !unseen` into `idle`, so
opening an agent erased it: finished-but-unlanded work looked exactly like a
workspace that never ran. Looking at something is not the same as dealing
with it.
Split the two on the map only, via a local `AgentMapNodeStatus` at the single
`agentMapNodeStatus` seam. `DashboardCardDotState` crosses the pop-out bridge
and is unchanged, so there is no wire change and bucket counts are untouched.
- done + unseen: filled emerald core, emerald halo, one-shot 1.4s flare
- done + seen: hollow emerald, no halo — still yours to land
- workspace ring turns green only once the whole workspace has settled
The flare is gated on a wall-clock recency window, not the map's `now` prop:
`now` ticks every 30s to refresh relative timestamps, so measuring a 1.4s
window against it fired at random moments instead of on the transition.
* fix(agent-map): bound completion paint work
* perf(agent-map): avoid fleet flare allocation
* fix(agent-map): localize seen completion label
* Remove unrelated merge formatting changes
* test(agent-map): guard completion burst paint budget
* fix(mobile): stop republishing stale launch agent identity (#14244)
* fix(mobile): keep terminal identity scoped to leaves (#14247)
* fix(mobile): scope terminal identity to leaves
* fix(mobile): preserve leaf title projection
* fix(agent-map): hold orchestration chevrons at a fixed pitch (#14256)
* fix(dashboard): hold agent map orchestration chevrons at a fixed pitch
CHEVRON_SPACING only chose how many chevrons to draw; placement then divided
the edge evenly, so the pitch grew with the distance between agents and, past
the 32-chevron cap, grew without bound. Step at a literal 8px pitch instead,
centering the run so it never overhangs either node.
Fixed pitch makes length drive the chevron count, so cap coverage now decides
how far the run reaches: at 32 it spanned only 248px and every longer link
would have shown a chevron cluster stranded mid-edge. Raise it to 256 (~2048px,
past any real in-project link) and memoize the generated path per endpoint
coordinates, since the scene rebuilds every path string on each zoom frame and
on 4Hz snapshot refreshes while world positions stay put.
* test(agent-map): isolate lineage path cache coverage
* fix(orchestration): preserve direct user authority after worker_done (#14192)
* fix(orchestration): preserve direct user authority
* test(orchestration): assert settled dispatch boundaries
* fix(terminal): clear the SGR pen on hidden-output restore and abandon (#14241)
* fix(terminal): clear the SGR pen on hidden-output restore and abandon
The hidden-delivery gate drops renderer-bound PTY bytes while a pane has no
visible view. The renderer's xterm is a separate emulator from the daemon
model, so when the dropped span contains the sequence closing an attribute run
(e.g. the ESC[22m ending a bold run) the renderer's pen stays latched while the
daemon model stays correct. Neither recovery path cleared it:
- buildMainModelSnapshotReplayWrites reset the pen on the two alt-screen
branches but not on the normal-buffer branch, and replayed scrollbackAnsi
ahead of the reset it did emit, so replayed content inherited the stale pen.
- abandonHiddenOutputRestoreAndDrainPendingForeground declares the dropped
bytes unrecoverable (it writes a user-visible warning) and then drained the
queued foreground chunks straight into xterm under that same unknown pen.
Add RESET_GRAPHIC_RENDITION and emit it ahead of replayed content in every
branch, and on both abandon exits. The existing profiles all clear DEC mode
bits and none touched SGR.
* fix(terminal): also restore charset designation after a dropped-byte gap
A gap can strand more than the pen: a dropped `ESC(B` leaves line-drawing
selected and ordinary text renders as box characters. Route both recovery
paths through one RESET_AFTER_BYTE_GAP profile covering SGR + charset.
Deliberately not a soft reset (DECSTR): xterm's DECSTR wipes kitty flags and
stacks (terminal-kitty-keyboard-mode-tracker applySoftReset), which would
silence Option chords for a live agent that negotiates them only at startup.
Reset what a gap strands and no running TUI re-asserts on its own; leave the
rest to its next repaint.
* fix(terminal): close the emulator state gap where the drop is announced
The restore-needed marker is the single point where "renderer-bound bytes
were dropped" is known. The handler already resets the transport's
cross-chunk parser state there for exactly this reason — a partial escape
spanning the gap would corrupt the next chunk. The emulator carries state
across chunks in the same way, so reset it in the same place.
That makes restore, abandon and overflow all start from a known pen by
construction, instead of each recovery path having to remember.
* fix(terminal): fully ground byte-gap recovery state
* fix(terminal): reset state when remote restore re-arms
* fix(terminal): keep the gap reset on the warning abandon path
The reset had been folded into an else of the unavailable-warning branch, so
the primary abandon path relied on the marker's earlier reset still standing.
It does not always: this function captures a replayingSnapshot, so it can run
after a partially-applied replay has already moved the pen, and the warning
itself is plain text carrying no SGR. Restore the unconditional write, guarded
only against the remote re-arm which writes its own.
* fix(terminal): scope the byte-gap reset to the pen and skip it under flood
Two regression risks in the widened recovery reset, both removed:
- The profile had grown to cancel partial escapes, close OSC 8 and re-designate
all four ISO 2022 registers. Each changes what a live TUI sees on a path that
runs in production, and none has a reported symptom behind it — a legitimately
line-drawing TUI that does not re-designate after recovery would render box
characters as ASCII. Scope back to SGR, which is what the field reports show.
- The marker-time reset ran before the flood-backpressure guard, so a flood
wrote one reset per marker in exactly the case that guard exists to damp. Move
it after; the flood path repaints through buildMainModelSnapshotReplayWrites,
which grounds the pen itself, so no coverage is lost.
Coverage verified non-vacuous: blanking RESET_AFTER_BYTE_GAP fails 5 tests
across all four paths (replay branches, marker, abandon-with-warning, remote
re-arm).
* test(e2e): stabilize release triage coverage (#14328)
* test(e2e): stabilize release triage coverage
* test(e2e): avoid post-selection remount race
* test(e2e): wait for New tab search focus
* test(e2e): focus New tab search without pointer input
* test(e2e): reopen voice settings after device change
* test(renderer): cover Agent Dashboard status replay fanout (#14036)
* test(native-chat): deflake watcher rebind (#14335)
* fix(orchestration): retry silent mail pointers (#14332)
* fix(orchestration): retry silent mail pointers
* fix(orchestration): bound mail pointer repair
* fix(docs): replace stale preload typecheck reference (#14298)
* docs: fix stale preload typecheck reference
Signed-off-by: HoonDongKang <d159123@naver.com>
* docs: keep .d.ts guidance canonical
---------
Signed-off-by: HoonDongKang <d159123@naver.com>
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
* fix(sidebar): remove non-Orca source icons (#14345)
* feat(dashboard): give the agent question state its own color and glyph (#14248)
* feat(dashboard): give the agent question state its own color and glyph
On the agent map, 'working' (yellow-500) and 'waiting' (amber-500) sat 16 hue
degrees apart and rendered as identical halo rings, so the only cue separating
"busy, leave it alone" from "it is asking you something" was a hue step most
people cannot resolve at map zoom. Every other surface distinguishes the two by
shape (spinner vs question glyph); the map had dropped that.
Move 'waiting' to orange-500 — midway between working-yellow and blocked-red —
and give the map node the same question glyph the sidebar and tabs already use,
so the state reads by shape when hue fails (low zoom, red-green CVD).
Both live behind one token, --agent-question, plus a shared AgentQuestionIcon,
so the sidebar, terminal tabs, kanban, toolbar and map cannot drift apart again.
The unread amber pip is deliberately left alone: "new output" and "needs an
answer" are different states and now read as different colors.
* test(dashboard): retarget the agent-row question assertion at the shared token
DashboardAgentRow renders the glyph through AgentStateDot, so the component
already moved with the token — only its assertion still pinned text-amber-500.
Missed locally because I ran components/dashboard-popout but not
components/dashboard.
* fix(dashboard): raise light question marker contrast
* perf(dashboard): keep question badges on the SVG paint path
* fix(worker-start): match Codex effort ceilings (#14281)
Honor the advertised reasoning-effort ceilings for Codex models, preserve conservative unknown-model handling, and localize the new ultra effort label.
* feat(sidebar): replace the external-worktrees inbox list with one card (#14341)
The Non-Orca worktrees modal already lists every hidden external worktree
with search, virtualization, and per-row Show. Its filter is a strict
superset of the sidebar inbox's, so the expanded sidebar list was a second,
worse copy that grew to ~900px at 24 worktrees.
The inbox is now a single clickable card stating the count, which opens that
modal. Drops the expand state, nested list, per-row Import, and the
Keep hidden / Import all footer.
* fix(native-chat): retain watcher identity during initial drain (#14344)
* fix(skills): use exported recipe id in environment guide (#14280)
* fix(skills): use exported recipe id in environment guide
* fix(skills): keep recipe-derived Vercel names valid
---------
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
* feat(agent-map): flare fresh questions and finishes (#14349)
* fix(ai-vault): make the merged scan stamp independent of leg order (#14270)
* fix(ai-vault): make the merged scan stamp independent of leg order
The all-host merge picked its stamp with a strict `stampMs > latestMs` and
echoed the winning leg's verbatim string. Two legs reporting the same instant
in different legal ISO shapes ("...:05Z" vs "...:05.000Z") therefore resolved
by position in the results array, i.e. by host-enumeration order (local, then
SSH, then runtime). The prior lexicographic max was order-independent, so this
was a regression with no test covering it.
A merge has no single scan instant, so its stamp is derived data rather than
any one leg's string: return the canonical ISO form of the newest accepted
instant. That is order-independent and format-independent, and drops a
variable instead of adding a tie-break branch.
Also share one request resolver between main and the renderer so the
renderer's merged-scope predicate is equivalent to main's routing by
construction, rather than by a comment that overclaimed it.
* docs(ai-vault): scope the merged-predicate comment to the desktop IPC path
The replacement comment still asserted the result is always several hosts'
legs. The paired web transport drops executionHostScope and serves one host,
so 'all' there is a single scan. State that the predicate is deliberately
over-inclusive and why erring the other way would be unsafe.
* test(ai-vault): pin the merged-stamp Date range boundary
new Date(ms).toISOString() throws RangeError outside +/-8.64e15. That is
unreachable only because Date.parse applies TimeClip, so the NaN guard alone
constrains the argument. Nothing pinned that. Dropping the guard now fails
these two cases with the RangeError they exist to prevent.
* refactor(ai-vault): route session-title scope through the shared resolver
The last character-for-character copy of the request-scope default. Leaving
it would make the shared resolver the single source of truth for two of three
sites, which is the drift this change exists to remove. No behavior change.
* fix(browser): fence the cookie-clear fallback to the pre-clear snapshot (STA-4170) (#14343)
* fix(browser): fence the cookie-clear fallback to the pre-clear snapshot (STA-4170)
The post-rejection fallback re-read the live cookie jar and removed everything
removable at that moment, while restore only ever covered the pre-clear
identity snapshot. A cookie that arrived mid-clear -- a login the user had just
completed -- was therefore deleted with no identity able to put it back, and a
later removal failure still reported "existing cookies were restored".
Fix the removal plan at the same point as the identity snapshot so the mutated
set can never exceed the restore set. Arrivals are not touched at all, which
matches what the successful bulk-clear path already does.
* test(browser): lock same-coordinate arrival rollback to the pre-clear value (STA-4170)
* fix(daemon): isolate durable-history checkpoints per session (STA-4173) (#14346)
PR #14193 routed warm reattach and deep-buffer snapshots through
overlayDurableRestoreSnapshot, which joined a single process-wide
checkpoint tail with no deadline. One never-settling checkpoint blocked
every daemon-backed terminal from reattaching.
Checkpoint exclusivity only protects one session directory's
tmp-write/rename pair, so serialize per session instead. Reattach now
waits on a bounded deadline and degrades to the daemon's live window;
the abandoned compact keeps running and still commits, so a blown
deadline costs restore depth for one reattach, never durable history.
* revert(ssh): back out #13326 and #13928 — reconnect loses every tab (#14361)
* Revert "fix(daemon): stop killing live coding agents when the daemon can't report its sessions (#13928)"
This reverts commit 2e8cf589de2bee2365d3798b471f39bc7e13989b.
Reverted together with #13326: the 1.4.182-daily.202608131439 build carrying
both loses every tab on an SSH disconnect/reconnect cycle. Reverting first so
main stays releasable and the P0 fixes in the wild remain cherry-pickable,
rather than fixing forward on a shipped regression.
* Revert "fix(ssh): stop SSH reconnect from multiplying terminals and resuming agents twice (STA-3077)" (#13326)
This reverts commit 3ab8b6a11769666b7ac4d43a38a82c771dffaad6.
Reported on 1.4.182-daily.202608131439: connect to an SSH worktree, disconnect
the host from the hosts popup, reconnect — every tab is gone. That is worse than
the behaviour this PR set out to fix, where most tabs were retained.
Reverting rather than fixing forward, so main stays releasable and the P0 fixes
already out in the wild stay cherry-pickable. STA-3077 stays open.
* fix(terminal): resolve legacy wrapper fallbacks from PATH only (STA-4169) (#14365)
The retained git/gh compatibility wrappers resolved their real binary with an unqualified lookup. On Windows that searches the current directory before PATH, so a repository-local git.exe/gh.exe ran with the user's arguments. The POSIX wrapper had the same exposure through empty PATH elements, which mean the current directory.
All three wrappers now resolve only against the cleaned PATH: cmd and PowerShell walk it explicitly instead of delegating to a lookup that includes the cwd, and the POSIX filter drops empty elements. Candidates inside the wrapper directory stay rejected.
* fix(orchestration): wait for Claude composer render (#14342)
* fix(orchestration): make late task readiness atomic (#14163)
Derive initial dependency readiness inside the task INSERT so concurrent completion and creation cannot strand a task in pending.
Add deterministic state-machine coverage and a built-CLI RPC/runtime persistence E2E.
Fixes #14143
* fix(orchestration): resolve explicit worker worktrees directly (#14275)
* fix(orchestration): resolve explicit worker worktrees directly
* fix(orchestration): share worker workspace resolution
* fix(runtime): reject cross-host path ambiguity
* fix(orchestration): share federated workspace resolution
* refactor(runtime): share worktree host identity
* test(orchestration): align worker lifecycle fixtures
---------
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
* fix(terminal): make the bold font weight its own setting (#14368)
* fix(terminal): make the bold font weight its own setting
Deriving bold as max(700, regular + 200) silently destroyed bold. A family
exposes only a few real faces: the monospace the default chain resolves to on
macOS has exactly two, splitting at 600. Measured by rasterizing each weight to
a canvas — 100-500 are byte-identical (ink 3023) and 600-900 are byte-identical
(ink 3855), at every weight the same advance. So any base weight at or above 600
put both values in the same face and bold stopped existing, on 4 of the 9
positions the slider offers, with no error and nothing the user could do.
Arithmetic cannot fix it — on a two-face family there is no heavier face to
escape to. So bold is now user-owned: a new terminalFontWeightBold setting with
its own control, defaulting to 700. The default pair (500/700) straddles the
boundary, so existing profiles render exactly as before; a collision is now a
choice the user can see and undo.
The old test asserted 800 -> {800, 900} as 'keeps bold heavier', which is where
this hid: numerically heavier, identically rendered.
* fix(terminal): surface bold face collisions accurately
* fix(serve): exit cleanly after headless Linux signals (#14334)
* fix(serve): keep owned Xvfb alive through Electron teardown
* test(serve): gate packaged signal shutdown
* test: harden headless shutdown lifecycle gate
* fix(serve): isolate Xvfb from foreground signals
* docs(serve): preserve Xvfb during systemd stop
* test(serve): pin shutdown policy to owned Xvfb unit
* test(serve): harden shutdown gate portability
* test(serve): bound systemd unit parsing
* Fix paired remote terminal browser links
STA-4181
* Narrow paired browser link fix to registration race
* fix(terminal): reject cwd-resolving PATH entries and harden the legacy wrappers (#14370)
* fix(terminal): guard cmd wrapper PATH walks against an empty variable
An empty PATH or cleaned PATH left the substitution with an unbalanced quote, desynchronizing cmd parsing so the not-found branch emitted a parse error instead of its message. Reproduced and fixed on real Windows.
* fix(terminal): reject relative PATH entries and drop the dirname dependency
Adversarial review found the STA-4169 fix incomplete: it dropped empty PATH elements but kept relative ones, which resolve against the current directory identically. A repo-local git was still executed via PATH=. or node_modules/.bin, reproduced in all three wrappers.
Also removes the external dirname call in the POSIX wrapper (an unresolvable dirname silently made wrapper_dir the cwd, so the wrapper failed to exclude itself and reported git missing while git was on PATH), compares against the cached wrapper dir in the cmd subroutine where %~dp0 is rebound by CALL, and probes .cmd after .exe so a non-.exe git is still found.
* fix(terminal): test rooted paths in pure batch, not via an external tool
The rooted-path guard shelled out to findstr, which cmd resolves from the current directory first — so a repository-local findstr could run, and a malicious one could report success for every entry and defeat the guard entirely. Same hijack the guard exists to prevent. Now a pure-batch substring test with no external process.
Splits the Windows wrapper templates into their own module to stay under max-lines without a suppression.
* fix(terminal): reject drive-relative PATH entries and fix no-slash wrapper dir
Round-2 review: PowerShell IsPathRooted accepts drive-relative C:foo, which still resolves against the current directory on that drive. IsPathFullyQualified is absent on Windows PowerShell 5.1 (verified 5.1 on the test host), so match the same prefixes the cmd wrapper accepts.
The POSIX %/* strip yields the file name when the path has no slash, so wrapper_dir became the name, self-exclusion missed the shim dir, and the lookup resolved back to the wrapper (spurious 127). Handled with a case split.
Also probes .bat, since the replaced lookup honored PATHEXT.
* fix(agent-status): preserve Codex escape interruption (#14372)
* fix(vm): make hidden SSH cleanup retryable (#14351)
* feat(vm): add provisioned root recipe contract (#14352)
* fix(vm): adopt provisioned SSH checkout roots (#14353)
* feat(vm): create workspaces from provisioned SSH roots (#14359)
* feat(vm): use recipe-provisioned SSH roots
* fix(vm): preserve ordinary create failure timing
* test(vm): prepare provisioned root SSH fixture
* ci(vm): enable SSH setup for provisioned root E2E
* fix(macos): recover severed terminal TCC attribution after updates (#13992)
* fix(macos): recover severed terminal TCC attribution after updates
When a packaged update leaves the daemon healthy but TCC-severed (spawning
binary gone), surface a Manage Sessions toast and replace the daemon before a
new terminal only when zero live sessions remain. Does not broaden FDA or
auto-kill sessions. Addresses the Orca-specific path of #13594.
* fix(macos): coalesce severed-TCC toast checks and sync i18n keys
Add catalog entries for the Manage Sessions toast strings and guard overlapping
mount/focus probes with an in-flight latch so only one infinite toast can fire.
* test(macos): harden severed TCC recovery coverage
* fix(macos): clear recovered TCC warning
* fix(macos): clarify severed TCC warning copy
* fix(macos): bound TCC attribution health checks
* fix(macos): preserve bounded attribution checks after merge
* test(macos): use real execFile callback contract
* fix(macos): cover legacy daemon attribution
* fix(macos): clarify affected Orca terminals
---------
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
* feat(dashboard): give the agent map a real filter panel (#14338)
* feat(dashboard): filter the agent map by multiple hosts
The map's host control was a single-choice segmented row in its own header:
pick All, Local, SSH, WSL, or Remote — never two at once. It also sat outside
the shared filter menu, so the menu's badge never counted it and "Clear all
filters" could not reach it.
Hosts are now checkboxes in the filter menu alongside agent state, so a fleet
split across a Mac and a Windows runtime can show both and hide the rest.
Hosts that contribute no agents are omitted, and the section hides entirely
for a single-host fleet.
Lifting `hostFilter` out of AgentMap also drops `compact`, which existed only
to hide the segmented control next to an open terminal panel.
The filter-option builders move to agent-dashboard-filter-options.ts to keep
AgentDashboardToolbar under the 400-line cap; that move is mechanical.
* feat(dashboard): give the agent map a real filter panel
The map's filters were a single-choice host row in its own header plus a few
rows borrowed from the board's dropdown. This replaces them with one panel that
owns every map facet.
- Quick views: Everything / Needs me / Stuck / Unread / Last 30 min /
Long runners / Stale > 3d / Orchestration. Each replaces the filters
wholesale rather than stacking on whatever was set.
- Provider (agent type) checkboxes.
- Three time ranges — session lifespan, since last message, time in current
state — on a non-linear scale so minutes and days are both reachable.
- Orchestration flows select the coordinator *and* the agents it dispatched;
a children-only filter hid the half that explains the flow.
- Sections collapse, each showing its value when closed. A section that is
filtering forces itself open so a collapsed row can never be the reason the
map looks emptier than the filters claim.
- The shown-count moves into the panel header, next to the controls that
change it, and now reflects the map's own facets rather than the board's.
The map's surface becomes a Popover, not a DropdownMenu: the panel holds range
sliders and a Radix menu swallows the arrow keys those need. The board keeps
its dropdown untouched via a new `filterControl` slot on the shared toolbar.
`ui/slider` renders one thumb per value so a two-value range works; the two
existing single-value callers are unchanged.
* fix(dashboard): preserve agent map filter semantics
* refine agent map filter facets
* perf(dashboard): keep map filters out of main view
* fix(terminal): bake the wrapper interpreter and emit CRLF for cmd (#14382)
* fix(terminal): bake the wrapper interpreter and emit CRLF for cmd
Round-3 review: the shebang resolved bash through the inherited PATH before any script hygiene, so a relative or empty PATH element let an untrusted checkout supply the interpreter. Bakes an absolute verified interpreter with an env fallback.
Normalizes trailing separators in the cmd comparisons; without it a wrapper-dir entry spelled with a trailing backslash escaped self-exclusion and the wrapper tail-chained to itself forever.
Found while verifying on Windows: cmd resolves call targets by byte offset and fails on LF-only files once they grow (worked at 2.4 KB, failed at 3.7 KB), so the Windows wrappers now emit CRLF. Also avoids two cmd parser traps: a trailing backslash before a closing quote in an if comparison, and percent expansion inside rem comments.
* fix(terminal): resolve the wrapper interpreter from absolute PATH entries too (STA-4226)
The well-known-candidate list would fall back to an ambient env lookup on distributions that place bash elsewhere (NixOS, Guix), restoring the exposure. Search absolute PATH entries before giving up, skipping relative and empty ones because those mean the current directory.
* fix(ssh): reapply the reattach-identity work and stop the fallback fence stranding moved panes (#14384)
* Reapply #13326 and #13928 (un-revert #14361)
Restores the SSH reattach-identity and daemon-occupancy fixes. Reverting them
reintroduced their P0s, filed as STA-4224, STA-4225, STA-4227, STA-4230,
STA-4232, STA-4233 and STA-4234 against #14361.
The tab loss that motivated the revert is fixed in the commits that follow, so
this reapplication is not a straight redo.
* fix(relay): stop the fallback attach fence refusing a pane that moved tabs
The primary fence was moved to the shell's own incarnation precisely because
paneKey/tabId froze the pane's LOCATION at spawn and refused panes that had
merely moved. The fallback that older clients fall into kept the old rule, so
the correction never reached it — the same 'the rule exists, but this path does
not ask it' leak this work has hit repeatedly.
A refusal here is not recoverable: an identity mismatch never grounds a respawn,
so the pane keeps a live shell it can no longer reach and renders blank.
Narrowed to paneKey, which is the identity; the tab is a location. Restoring the
tabId comparison reddens the new test.
* Fix Codex hook trust before manual shell launches (#14326)
* fix codex hook trust before shell launch
* fix packaged cli preflight dependency
* fix codex shell preflight safety
* fix Codex shell preflight settings and startup safety
* fix(updater): recover renderer shutdown checkpoint (#14373)
* fix(updater): recover renderer shutdown checkpoint
* test(updater): cover checkpoint recovery in Electron
* fix(updater): keep staging failures blocking
* fix(terminal): name the POSIX interpreter directly when generating on Windows (#14386)
The resolver could never succeed on Windows — no absolute candidate exists there and the PATH search split on the POSIX delimiter — so it always fell back to the ambient lookup it exists to avoid, on a wrapper Git Bash and WSL panes do execute.
Also rejects interpreter candidates containing whitespace (a shebang cannot quote) or that are directories (X_OK alone is true for those), and replaces the sentinel trailing-separator strip, which corrupted paths containing the sentinel, with a comparison against a variable holding the separator.
* fix(mobile): self-heal host opens and harden session liveness (#14333)
* fix(terminal): clear CDPATH when resolving the wrapper directory (#14387)
* fix(terminal): clear CDPATH when resolving the wrapper directory
cd consults CDPATH for a relative operand and echoes where it landed, which the command substitution captured — wrapper_dir came out wrong, the lookup resolved to the tombstone itself, and git died at 127 with a working git on PATH.
Also pins the self-reference guard that turns that failure into a clean 127 rather than repeated self-exec; deleting it previously left the suite green. Both new tests are mutation-verified: an earlier version of each passed with the bug reintroduced because the CDPATH fixture mirrored only the first path segment.
* test(terminal): pin the wrapper guards that survived mutation
A mutation campaign found several security properties with no coverage. Most important: nothing asserted the Windows *reject* path for relative candidates, so deleting it reopened the cwd hijack while every accept-path assertion stayed green.
Also pins the separator variable value (any other character silently un-pins the trailing-separator fix), and grounds the interpreter-search test in the running process cwd — a relative PATH entry resolves against that, so the previous fixture in a tmpdir passed whether or not the guard existed. Each is mutation-verified.
* Revert "fix(ssh): reapply the reattach-identity work and stop the fallback fence stranding moved panes" (#14395)
* perf(worktrees): bound the duplicate-id scan in reuseEqualCatalogRows (#14271)
* perf(worktrees): bound the duplicate-id scan in reuseEqualCatalogRows
Rows sharing an id are scanned linearly with a deep compare each, so a bucket
of k duplicates costs O(k^2). Both callers key on ids that are unique by
construction, so this is a bound on damage rather than a live fix.
Cap the scan instead of adding a second index: reuse is only an optimization,
so a missed match yields a new object identity, never a wrong row. A fingerprint
index would buy a little more reuse in a case nothing reaches, at the cost of a
second equality implementation that must stay in step with catalogValuesEqual
with no automated guard.
Worst-case duplicate bucket, no matches: k=1000 134ms -> 1.2ms, k=2000 541ms ->
2.2ms. Unique-id path unchanged (2000 rows: 0.86ms both).
* docs(worktrees): lead the duplicate-id cap comment with its reachability
A reader hitting MAX_DUPLICATE_ID_SCAN should learn first that no caller
produces duplicate ids today, so the cap reads as bounding future damage rather
than fixing something live.
* test(terminal): pin the wrapper security properties that survived mutation (#14394)
* test(terminal): pin the seven wrapper security properties that survived mutation
The PowerShell drive-relative pin was a prefix that also matched the insecure regex, so re-accepting C:foo left the suite green. Now asserts the full pattern.
Also pins all four PowerShell TrimEnd sites, the fallback wrapper-dir skip, the cmd legacy-dir trailing-separator strip, and C:/-style drive acceptance, plus a behavioural test for the POSIX legacy-dir reject with a distinct ORCA_ATTRIBUTION_SHIM_DIR. Each verified by reverting the property and confirming the suite fails.
* test(terminal): cover the trailing-separator scrub in the path matcher
isLegacyTerminalShimPathEntry sits outside the two generator modules, so the mutation campaign never re-ran it. Dropping its trailing-separator strip survived the suite: a PATH entry spelled with one would not match, leaving the legacy shim directory on the spawned PATH and the wrapper reachable. Mutation-verified on both the POSIX and Windows spellings.
* ci(windows): cover the worktree admin fingerprint on the Windows runner (#14378)
The fingerprint gate added in #14207 reads Git's administrative layout directly -- `.git` as a file or directory, `commondir`, and per-worktree `HEAD`, `gitdir`, and `locked` -- instead of shelling out to `git worktree list`. That makes it depend on Windows path resolution, CRLF inside those files, and whether `worktree move`/`lock` and deleting a live checkout behave as they do on POSIX.
PR CI runs the vitest suite on ubuntu-latest only, so none of that was exercised. Both suites were verified by hand on a real Windows host (Git 2.55.0.windows.3, Node 24.18.0) and pass 25/25, but nothing kept them passing.
Add them to the existing curated `Test Windows-specific boundaries` step rather than standing up a new job: the `package (windows)` job already checks out and installs dependencies, so this costs only the tests themselves.
* Center newly opened tabs at end, keep close controls visible (#14314)
* Center newly opened tabs at end, keep close controls visible
- Add end padding to tab strip to preserve close button visibility against fade
- Center scroll position for newly revealed end tabs, avoiding hard scroll-to-end
- Only pin-to-end for new tabs that are also active; restored mid-strip tabs stay in place
* test(tab-scroll): clarify when tabs center vs settle at end
Add lastTabGeometry helper to correctly derive scrollWidth from the pad policy. Update test expectations to document that center requests only land centered when the strip is within 2x the inset of the last tab's width; wider gaps settle at the end with inset.
* fix(tab-bar): sync active tab id ref after render
React Doctor rejects mutating refs during render. Keep the latest
active tab id in a layout effect so overflow navigation still sees
it without an impure render write.
* fix(orchestration): release federation ack checkpoints once a dispatch settles (#14380)
Checkpoints were inserted per synced federated dispatch and never removed; the
only eviction dropped the whole map, and none of its three call sites fire in
normal operation. A long-running federated coordinator retained one small
object per dispatch for the process lifetime.
Prune from the existing syncOrchestrationFederatedDispatch finally, which is
the one hook covering all five paths that create a checkpoint — including the
two RPCs that sync an already-terminal dispatch with no timer to prune after.
Refs STA-4014
* fix(worktree-scan): stop a stalled admin probe from poisoning repo refresh (#14379)
The scan cache stored the Git-admin fingerprint as an unsettled promise, so a
readdir/stat that never returns (hard NFS/SMB mount, dead sshfs, wedged cloud
FileProvider) left every later refresh awaiting it. The in-flight entry was
never cleared, so the repo fell back to persisted rows indefinitely — and since
computeResolvedWorktrees awaits all repos together, one wedged repo added 5s to
every snapshot for all of them.
Cache the settled value instead, filled by an identity-guarded writeback, and
bound the one branch that actually awaits the probe. withTimeout cannot cancel
a readdir, so an outstanding-probe guard keeps a wedged mount from issuing a
fresh probe every refresh and pinning every libuv fs thread.
Refs STA-4171
* fix(workspaces): gate the GitLab palette number match on repo identity (#14381)
* fix(workspaces): gate the GitLab palette number match on repo identity
After a full GitLab identity compare failed, matching fell through to an
unguarded iid compare, so a pasted URL could activate a workspace from another
project or host. GitLab iids are per-project and start at 1, so small-number
collisions across projects are the norm.
Mirror the GitHub shape rather than rejecting outright: gate the number-only
fallback on a tri-state repo identity check that stays permissive when identity
is unresolvable, so forks and host aliases keep matching. A stored URL that
parses to a different project with the same type and number is a direct
contradiction and is rejected even when identity is unknown.
Refs STA-4155
* fix(workspaces): keep the GitLab repo gate permissive for fork checkouts
deriveGitRemoteIdentity keeps a single remote and ranks `upstream` above
`origin`, so a fork checkout resolves to the upstream project and the fork's
own origin is invisible to the gate. Rejecting on that mismatch dropped MR
URLs from the fork the user actually checked out — a false negative the
tri-state was meant to prevent. Treat an upstream-derived identity as unknown.
Refs STA-4155
* docs(workspaces): note that the GitLab repo identity is a one-shot snapshot
* test(terminal): fix ordering pins that pass when the needle is absent (#14409)
Five pins compared indexOf positions without asserting presence, so a missing needle returned -1 and the assertion passed. Deleting the legacy-shim-dir capture in both Windows wrappers stayed green that way, disabling legacy-dir PATH removal. All five now route through a helper that asserts both operands exist first.
Adds coverage for four properties proven load-bearing by execution: the path_entry_kept guard (without it an empty cleaned PATH resolves a cwd-local git), the cleaned PATH on exec, and multi-separator matching in the shared path matcher.
* Forward Unix launcher signals to orca serve supervisor (#14071)
* fix(cli): forward Unix launcher signals to serve
* test(cli): retry incomplete listener state writes
* test(cli): cover both Unix launcher termination signals
* test(cli): cover macOS launcher signal oracle
* refactor(cli): keep Unix launcher exec atomic
---------
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
* fix(terminal): stop leaking raw PTY lifecycle tokens into the error toast (#14415)
* fix(terminal): stop raw PTY-not-found and session-expired tokens reaching the error toast
A reattach the host answers "no such session" for surfaced its wire token
verbatim — `SSH_SESSION_EXPIRED: orca:<conn>@@pty-N` or the relay's raw
`PTY "..." not found` — including the internal PTY id, and invited the user
to file an issue for an ordinary lifecycle event.
Humanize both in the toast's existing daemon-boundary seam and mark them
explained so the issue link is suppressed. The copy stays silent on whether
the remote shell died: absence from the host is not proof of exit.
Refs STA-4238
* style(terminal): apply oxfmt to the toast humanization tests
* fix(terminal): treat the humanized session copy as literal replacement text
* fix(repos): re-probe git remote identity so a stale snapshot stops misjudging identity gates (#14414)
* fix(repo-identity): re-probe resolved git remote identities on a long TTL
A resolved gitRemoteIdentity was written once and frozen for the life of the
repo record, so adding an `upstream` remote later — or a project rename or
transfer — left identity gates judging against the path the repo had when it
was added. Re-probe resolved repos on a 6h TTL, seeded 5 minutes after a repo
is first seen in a process and capped at 4 refreshes per sweep so a restart
cannot fan out a subprocess per repo. Only a successful probe that yields a
different canonicalKey overwrites; failures and no-remote answers leave the
existing identity alone.
Also explain why the worktree-scan admin fingerprint timeout deliberately
exceeds its caller budget, and log when that probe expires — expiry was
silent and indistinguishable from "fingerprint unavailable".
Refs STA-4247
* fix(projects): carry project state across derived project id changes
A project id is derived from repo identity, so a remote re-probe (or a
repo:->git:->github: promotion) rewrites it. The compatibility merge matched
prior rows by id only, dropping the user's localWindowsRuntimePreference and
leaving a ghost project row that independent host setups still pointed at.
Both merge sites now fall back to the prior row whose sourceRepoIds overlap and
re-point independent setups at the surviving project.
* fix(workspaces): gate the GitHub palette number match on repo remote identity (#14413)
* fix(workspaces): gate the GitHub palette number match on repo identity
`repoMatchesGitHubSlug` returned the permissive `'unknown'` whenever the repo
displayName was not in `owner/repo` form and no upstream metadata existed — the
common basename-named non-fork case. The caller only rejects on `false`, so a
pasted issue/PR URL could activate a workspace in a different repo that happened
to share the number, since issue/PR numbers are per-repo.
Mirror the GitLab gate from #14381: fall back to the probed
`gitRemoteIdentity.canonicalKey` before giving up, comparing host and owner/repo
after normalizing port, `www.`, and case. An `upstream`-derived identity stays
`'unknown'` because `deriveGitRemoteIdentity` ranks `upstream` above `origin`, so
a fork's own origin is invisible and rejecting would drop URLs from the fork the
user actually checked out.
The canonicalKey compare runs after the displayName branch: displayName is
compared host-agnostically, so mirrors and host aliases of the same owner/repo
keep matching as they do today, and the probed remote only fills in where no
name evidence exists.
Refs STA-4237
* fix(workspaces): keep SSH host aliases matching in the palette identity gate
`git remote -v` reports ssh.github.com, www., and ~/.ssh/config `Host` aliases
verbatim, so comparing a probed canonicalKey against a pasted URL host rejected
legitimate GitHub/GitLab remotes. Normalize the alias hosts both sides can fold
offline, and downgrade a host-only mismatch to 'unknown' when the probed host is
dotless (an unexpandable OpenSSH alias); dotted hosts like ghe.example.com still
lose. Lifts the GitHub host normalizer into shared instead of a third copy.
* fix(repos): keep the www host fold out of the derived project identity
getProjectIdentityKey feeds the persisted Project id, so folding www. there
re-keyed existing projects on upgrade and dropped localWindowsRuntimePreference.
Restrict the fold to the palette's URL-vs-remote comparison, and pin the derived
id for a www. remote so it cannot drift silently again.
* refactor(shared): split shared/types.ts into per-domain type modules (#14397)
`src/shared/types.ts` was 3,981 raw lines (2,825 counted, 9.4x the 300-line
budget) behind an `eslint-disable max-lines`, and is imported by 2,092 files —
the single widest contract surface in the repo.
Move all 320 top-level declarations into 46 per-domain modules
(`repo-types.ts`, `worktree-types.ts`, `github-pr-types.ts`, ...) and reduce
`types.ts` to an explicit re-export barrel, so the 2,092 import sites are
untouched.
`HostSettingOverrides` moves into the pre-existing `host-setting-overrides.ts`
alongside the accessors that operate on it, which also removes that module's
circular import back into `types.ts`.
Named re-exports only, never `export type *`: with star re-exports a name
exported by two modules is silently dropped, which would surface as a confusing
"has no exported member" at a random call site.
Verified lossless mechanically, not by inspection:
- export parity — the module's resolved export set through the TS checker is
identical before and after (396 names, no additions, no removals)
- declaration parity — all 320 declarations compare character-identical modulo
comments and whitespace, so no optionality, union order, or generic
parameter drifted
- `tsc --noEmit` green on the node, cli, and web projects
- `oxfmt --write` is byte-identical, so the barrel is format-stable
Drops the `max-lines` bypass and its baseline entry (ratchet 346 -> 345).
* refactor(shared): group worktree, github, and linear modules into folders (#14437)
`src/shared` is a flat directory of ~1,150 entries. The worktree, github, and
linear domains accounted for 71 of them, so finding the module you wanted meant
scanning a wall of same-prefixed filenames.
Move each domain into its own folder and drop the now-redundant prefix:
src/shared/github-pr-types.ts -> src/shared/github/pull-request-types.ts
src/shared/worktree-id.ts -> src/shared/worktree/id.ts
src/shared/linear-links.ts -> src/shared/linear/links.ts
This follows the existing `network/` and `new-workspace/` convention in the
same directory, which also drop the prefix inside the folder.
Whole clusters move, including tests. Foldering only part of a domain would be
worse than flat: a reader would have to check both `github/` and the flat
directory, and `github-auth-types.ts` / `github-project-types.ts` are type
modules that belong with the rest. No files with these prefixes remain flat.
Import specifiers were rewritten by resolving each one to an absolute path and
recomputing it, not by string substitution, so the `@/../../shared/...` alias
forms are handled correctly. 501 specifiers across 298 files.
Two things `tsc` cannot catch, handled explicitly:
- `github-project-types.ts` carries its own `max-lines` bypass, so its baseline
entry is REPOINTED to the new path rather than pruned. Pruning would drop the
bypass and then flag the new path as a fresh violation. Ratchet stays at 345.
- `mobile/` is outside `pnpm typecheck` and cannot be typechecked here
(`mobile/node_modules` is empty). Instead every relative specifier in the repo
was resolved against the filesystem: 174 unresolved before this change and 174
after — identical, so nothing broke in mobile either.
The pinned `tests/e2e/.cross-version-checkouts` fixtures are deliberately NOT
rewritten; they are a snapshot of an older release and still reference the old
paths.
Verified: cold `tsc --noEmit` green on node, cli, and web (buildinfo deleted
first — these projects are `composite: true` and reuse stale caches).
* refactor(preload): split the preload contract into per-domain api modules (#14403)
`src/preload/api-types.ts` was 3,752 raw lines (3,533 counted, 11.8x the
300-line budget) behind an `eslint-disable max-lines`. Almost all of it was a
single `PreloadApi` object type whose ~83 namespace properties were declared
inline, so any IPC surface change meant editing one 2,600-line type.
Give each namespace a named type in its own module under `src/preload/api/`
(`pty-api.ts`, `filesystem-api.ts`, `github-pull-request-api.ts`, ...) and
recompose `PreloadApi` from those names. `api-types.ts` keeps the `declare
global` Window augmentation and re-exports every moved name, so all 52 import
sites are untouched.
Two shapes needed care to stay type-identical rather than merely compatible:
- Three keys (`gh`, `git`, `ui`) are composed from two modules each. A plain
intersection is NOT identical to the original flat object literal, so those
use a `Merged<T>` mapped type; a negative control confirmed that dropping it
fails the parity assertion.
- Keys whose module groups several namespaces use indexed access
(`fs: FilesystemApi['fs']`) to preserve exact identity and source order.
`config/tsconfig.web.json` and `tsconfig.tc.web.json` enumerate files by path,
so they need `src/preload/api/**/*` alongside the existing `api-types.ts` seed
or the web projects fail TS6307.
Verified by exact type identity, not assignability: 41 assertions of the form
`Equals<Now.X, Before.X>` against a frozen pre-split snapshot, covering every
exported name, plus a per-key pass over all 83 `PreloadApi` keys. All three
projects typecheck clean with those assertions active.
Verification note: these tsconfigs are `composite: true`, and `tsc --noEmit`
will reuse a stale `.tsbuildinfo` and report clean for a state that genuinely
fails. Every result above was produced after deleting the buildinfo, including
a negative control confirming the gate still fails on deliberate drift.
Drops the `max-lines` bypass and its baseline entry (ratchet 346 -> 345).
* fix(stats): wrap usage breakdown model names (#14065)
* fix(test): deflake relay exec env, project boundary, and speech resume tests (#14446)
Three test-suite problems, all root-caused in the tests rather than in
production behavior.
1. src/relay/agent-exec-handler.test.ts (real failure, not a flake)
The two spawn-argument assertions failed with "Number of calls: 1" — spawn
ran, but the env differed. Cause: both assert
`expect.objectContaining({ ...process.env, ... })`, which demands that every
ambient variable reach the child verbatim. #7986 (1a6abc87d11) changed both
sides at once: it rewrote the assertion from `env: process.env` to that
objectContaining form, and in the same commit made the handler apply
`applyTerminalGitCredentialPromptGuard`, which appends its own entries to
Git's indexed-config protocol (GIT_CONFIG_COUNT / KEY_n / VALUE_n).
So whenever the test runner's own environment already carries that protocol —
exactly what Orca exports into its agent terminals — the snapshot expects
GIT_CONFIG_COUNT=2 while the correctly guarded child gets 4. The test passes
on a bare CI shell and fails when run from a guarded terminal.
The implementation is right: appending the guard after the caller's config is
the documented contract, and "guards wrapped agents after atomically replacing
inherited indexed config" already covers it. Fixed the test instead, by
clearing the guard-owned keys (GIT_CONFIG_* protocol and WSLENV) from the
ambient env for the duration of the suite and restoring them afterwards, so
the passthrough baseline is deterministic. No assertion was weakened or
removed.
2. project-view-wrapper-source-context-boundary.test.ts (flake: 30s timeout)
`buildProjectWorkItem` is a pure function, but it lived in
ProjectViewWrapper.tsx, so importing it pulled in the store, sonner, lucide,
and the whole UI kit — ~8.8s of transform and module evaluation for one
assertion, which tipped past the 30s limit under parallel load.
Extracted it to project-work-item.ts (its only dependency is
githubProjectHost) and pointed the test there. Both test cases are unchanged.
Also dropped the now-unneeded happy-dom environment, since nothing in the file
touches the DOM any more. 9.15s -> 0.12s.
3. model-manager-download-resume.test.ts (flake: 30s timeout)
"bounds a server that advances by pathologically tiny segments forever"
drives the loop to the MAX_TOTAL_DOWNLOAD_REQUESTS ceiling of 4096. Each
iteration did a real writeFileSync plus two statSync calls through
getPartialDownloadBytes — ~12k synchronous filesystem syscalls in a tight
loop. Fast on an idle disk, but it serializes against every other vitest
worker on a loaded machine, which is what blew the per-test timeout.
Stubbed getPartialDownloadBytes to read the byte counter the test already
maintains, so the loop is pure CPU. The file was only ever a stand-in for that
counter. Ceiling and rejection assertions are unchanged: 332ms -> 15ms.
The two remaining ~1.1s cases in that file spend their time in the real 1s
retry backoff around real stream and file-write plumbing; they are left on
real timers because faking them would mean faking the transport too, and 1.1s
leaves ample headroom.
* Revert "Center newly opened tabs at end, keep close controls visible" (#14450)
* fix(native-chat): keep caller abort ahead of WSL gate refusal (#14445)
Why: a hook-path or scan refusal that races the caller's abort was
converted into a WSL unavailability error, so cancellation looked like
a stalled distro.
* refactor(store): unify the duplicated catalog equality and identity-key helpers (#13804)
* refactor(store): unify the catalog structural-equality walks
Three near-identical structural deep-equality walks had landed independently in
the same window: areValuesEqual (#13744, repo-identity-reconcile.ts),
areCatalogEntriesEqual (#13770, repos.ts — already folded into the first on this
branch's base) and catalogValuesEqual (#13662,
worktree-catalog-reconciliation.ts). All three walk plain records and arrays and
fall back to reference equality for anything exotic.
They are not interchangeable. Two axes genuinely differ, and each caller depends
on its own side:
- Own-key set. #13744/#13770 require equal own-key counts plus hasOwnProperty,
so an absent key differs from a key present and holding `undefined`. #13662
compares the union of both sides' keys, so those are equal. The strict side is
load-bearing: the repo/project merges branch on
`'localWindowsRuntimePreference' in project` (repos-project-runtime.test.ts
"clears stale local runtime preferences"), and projects are now reconciled
with this comparator. The loose side is test-pinned by
worktree-catalog-reconciliation.test.ts "reuses rows with equivalent nested
catalog data", where a locally built row carries `optional: undefined` that
the host omits.
- Leaf comparison. #13744/#13770 use `===` (NaN never equal, 0 equals -0);
#13662 uses `Object.is` (the reverse).
So instead of picking a winner, src/shared/structural-value-equality.ts holds
one walk parameterised by those two axes and exports the two policies as
`structuralValuesEqual` and `structuralValuesEqualIgnoringUndefined`. Every
caller keeps its exact current semantics; the ~40 duplicated lines and the
silent divergence go away. src/shared/persisted-ui-equality.ts (a fourth copy
with a Set branch and no plain-object guard) is deliberately left alone: it
gates a disk write in main with no direct test coverage.
Also folded, all provably behaviour-identical:
- The `${hostId}\0${repoId}` composite key had three copies
(getRepoHostIdentityForParts, repoOwnerKey, getEntryKey) that must agree or
repos silently stop reconciling. Moved to src/shared/repo-host-identity.ts
because one of them lives in src/shared; the renderer module re-exports it.
- mergeFetchedReposForHost's hand-inlined upsert loop now calls mergeByIdentity.
mergeByIdentity additionally skips replacing a structurally equal row, which
cannot change the result here: reconcileFetchedRepos runs immediately after
over the same identities in the same order and restores exactly those rows.
- Renamed repos.ts's `catalogRowsUnchanged` to `arrayElementsUnchanged`. It is a
pure element-identity compare, two files away from
`catalogRowsEqual`, which is a full structural compare.
src/shared/structural-value-equality.test.ts pins both policies over arrays,
nested records, null-prototype records, absent-vs-undefined keys, symbol keys,
and non-plain objects (Date/Map/Set/class) falling back to reference equality.
* fix(store): keep merged sourceRepoIds order host-independent
Prefixing the cross-host remainder made a cross-host project's sourceRepoIds
order a function of the refreshing host, so the projects reconcile never reused
the row. Also pins the repo-derived host-id contribution the new per-project
slice feeds the host-id resolvers.
Co-authored-by: Orca <help@stably.ai>
* refactor(store): migrate call sites that landed after this branch
github.ts and ai-vault-session-identity.ts began using areValuesEqual on main
while this branch was stale, and repo-identity-reconcile's record reconciler
still called its own deleted walker. All three now use structuralValuesEqual;
reuseEqualCatalogRows keeps its duplicate-id cap and calls the ignoring-undefined
variant, which is the key-union semantics catalogValuesEqual had.
---------
Co-authored-by: Orca <help@stably.ai>
* fix(worktree-scan): keep the admin-fingerprint probe inside the caller's per-repo budget (#14454)
* fix(worktree-scan): keep the admin-fingerprint wait inside the caller's per-repo budget
The awaited probe was capped at 10s while `computeResolvedWorktrees` gives each repo
5s, so a slow mount always blew the budget: the caller gave up and republished
persisted rows. The resolved snapshot was then stamped from the *start* of the
compute, so a compute longer than its 1s TTL published an already-expired entry and
the next poll repeated the whole 5s wait — deterministically, on every TTL expiry.
Cap the probe at 2s so the remaining budget still covers the fallback
`git worktree list`, and stamp the snapshot on completion.
* fix(worktree-scan): derive the probe deadline from the caller budget
A flat 2s cut reuse for hosts whose probe lands between 2s and 5s, which used to fit
the caller's budget — trading the stall for a repeating `git worktree list`. Subtract
a fallback-scan allowance from RESOLVED_WORKTREE_REPO_TIMEOUT_MS instead, so the
invariant holds by construction and only probes that could not have fitted are cut.
Tests now pin both ends: too large fails the budget invariant, too small fails reuse
for a slow-but-healthy probe.
* Require an absolute Orca CLI for the agent-teams tmux shim (#14438)
The generated tmux shim fell back to a bare orca / orca.cmd / orca-ide, and cmd.exe resolves an unqualified command against the current directory before PATH (sh does the same via ./empty PATH entries), so a stray orca.cmd in an agent's checkout could run with the agent-teams team id and token in its environment.
Resolve only absolute paths, honor the Windows Path env spelling, degrade to in-process teammates when no CLI can be qualified, and make both shims exit 127 instead of guessing. Verified on macOS, Linux (dash + bash), and Windows (cmd.exe + Git Bash).
Fixes STA-4215.
* fix(browser): acknowledge paired tab before navigation (#14402)
* refactor(shared): drop the shared/types barrel and import from the real modules (#14447)
#14397 split `shared/types.ts` into 46 per-domain modules but kept the path as
a re-export barrel so the import sites did not have to change. This removes
the barrel: every consumer now imports from the module that actually declares
the type, and `src/shared/types.ts` is deleted.
Barrels hide where a type lives, make every consumer look like it depends on
the whole domain, and let an unrelated edit invalidate a module that ~2,000
files transitively import.
2,323 import declarations across 2,321 files. Rewritten mechanically: each
specifier was resolved to an absolute path via the TypeScript AST and
recomputed, rather than string-substituted, so alias forms (`@/../../shared/
types`) and per-specifier `type` modifiers survive.
Four cases the mechanical pass had to handle, each found by a gate rather than
by reading the diff:
- Modules inside `src/shared` import the barrel as `./types`, not
`shared/types`. A pre-filter on the latter string skipped 176 of them and
left imports dangling at a deleted file, which surfaced as confusing
`Property 'x' is optional in type 'Repo' but required in Pick<Repo, ...>`
errors rather than "module not found".
- The barrel RENAMED one type on the way through
(`WorkspaceSource as WorkspaceCreateTelemetrySource`), so the original name
in the owning module has to be re-aliased at each consumer.
- Three test files put `;(globalThis as ...)` on the line after the import.
TypeScript parses that `;` as the import statement's terminator, so
replacing through `statement.getEnd()` deletes it and breaks ASI. The
rewrite now stops at the module specifier.
- A file that already imported directly from a module got a SECOND import
from it, because the barrel re-exported those same names — which trips
`import/no-duplicates` under `--deny-warnings`. A post-pass merges
declarations sharing a specifier and type-only-ness; the `import type` plus
`import` pair from one module is left alone, since that form is allowed.
Splitting one barrel import into several genuinely adds lines, which pushed
`terminal-layout-pty-ownership.ts` t…
Summary
STA-4091. After #13061, daemon rebuild / remount / restart / remote reattach restored only the 1000-row live window. Earlier builds restored 5000 rows.
This keeps the 1000-row live RAM window so an unbounded session count cannot grow the grid without bound. Rebuilds reconstruct the desktop 5000-row depth from durable history (checkpoint + incremental log + drained pending records).
Live emulator = bounded RAM. Disk history = rebuild source. No attention heuristics or heap-pressure trimming.
takePendingOutput.drainedRecordsis optional. Older daemons omit it and compact keeps the live snapshot. No new stream opcodes.Fixes STA-4091. Follow-up to #13061.
Reproduction on origin/main
Writing 5000 numbered lines then compacting the live snapshot restored from about
LINE_03978. Incremental history without that compact already reconstructedLINE_00001/LINE_01000. Red test:daemon-restore-scrollback-depth.test.ts.Testing
pnpm lint(oxlint + type-aware oxlint on touched files; max-lines ratchet OK)pnpm typecheck:nodepnpm test(focused daemon suites, not full repo)pnpm buildFocused: restore-depth, adapter, session, history recovery, incremental restore — 263 passed.
SSH
Throwaway Linux Docker SSH target (not localhost). Relay built (
linux-arm64). Isolated Electron identity:Orca: brennanb2025/sta-4091-restored-scrollback. SSH connected, relay deployed (relay.js, linux-arm64),/work/demo-projectregistered from a local git bundle after GitHub clone was blocked. Remote PTY spawned; warm reattachisReattach: true; relay process and bash PTYs observed on the container. Cleanup removed the containers.Follow-up exhausted the remaining routes. The SSH product path never constructs
DaemonPtyAdapter:getProvider(connectionId)returns the registeredSshPtyProvider;getProviderForPtyrefuses to fall through to the hub-local provider forssh:IDs.connectionIdis absent (SSH/headless don't use the desktop daemon).relay.jstalks tonode-ptyviapty-handler.ts. Grep of the live remote relay found noDaemonPtyAdapter/daemon-entrystrings.SshPtyProvider.canProvideAuthoritativeBufferSnapshotisfalse; remount uses the Mac headless/renderer buffer, not durable daemon history.orca servewould installDaemonPtyAdapterfor that host's local PTYs, but pairing to this PR's Linux serve requires a Linux Electron binary of this worktree. Starting this PR'sdaemon-entry.json the Linux target (using the relay's linuxnode-pty) bound a unix socket; no SSH runtime RPC or relay client connects to it.The 1000-vs-5000 contract is therefore out of scope for
SshPtyProvider/ relay replay (100KB). It is proven on the daemon path in unit tests. Residual from the first SSH lane: a 5000-linecatdid not appear in SSHgetMainBufferSnapshot(61-char prompt; shell-ready / visible-vs-headless buffer race). That residual does not exercise the changed adapter.Residual risk
drainedRecordskeeps the previous live snapshot.Visual proof (Grok Electron / Playwright CDP)
Isolated local daemon-backed terminal on this PR HEAD. Wrote 5000 uniquely numbered
STA4091_L#####lines, parked the worktree (xterm destroyed), then revealed it so remount used the authoritative snapshot path.STA4091_L05000).STA4091_L00001), not a 1000-row live-window flatten.STA4091_FRESH_AFTER_RESTOREwrite.