Repository navigation
session.resume of a self-closed host fails with EEXIST on <id>.lifecycle.ready.json #6261
Description
Activity
Triage: A (actionable, fix being prepared) · bug · P1
Why: A detached-idle host exits without removing its own
<id>.lifecycle.ready.json. The next host for that id fails on theO_CREAT | O_EXCLopen inwriteSessionLifecycleReady. After one failed resume the effect markers no longer match, so the session can't be resumed again unless someone deletes the file by hand.Checked against current
dev(42eb13a); the line numbers in the report still match:packages/coding-agent/src/commands/sdk.ts:1270-1299:exitAfterSessionDisposalchecks the endpoint file but never removes the ready file.packages/coding-agent/src/sdk/broker/lifecycle.ts:1921-1925: the exclusive open of the placeholder atreadyPath.lifecycle.ts:1667-1670:inspectedis incremented before the.lifecycle.jsonname filter, so the 64-entry limit is used up by unrelated files.lifecycle.ts:1679andlifecycle.ts:1686: the 60-minute age gate and the exact-marker match. Together they keep the self-same-id leftover from being reaped.
Planned fix scope: suggested fixes 1 and 2 together. Fix 1 stops new leftovers. Fix 2 retires a same-id ready/marker pair whose recorded owner is proven exited before spawning, which also repairs state dirs that are already stuck. Fix 3 (count only
*.lifecycle.jsontoward the limit) is a one-liner and goes in the same change. Each fix gets a failing regression test first.Note: open PR #6148 also touches
sdk.tsandlifecycle.ts(not this code path). Whichever lands second will need a rebase.Thanks for the detailed root-cause report.
- added a commit that references this issue
on Oct 3, 2026 Additional independent reproduction on gjc 0.18.5, Linux aarch64 using an otherwise empty workspace. This is not limited to detached-idle exit:
sdk session raw global --op session.create --idempotency-key <create-key> --json-input '{"cwd":"<fixture-workspace>","readinessTimeoutMs":60000}'succeeds. Exact inspect reports the original ID andlive:true; create ledger reachesterminal_ok.- Official global
session.closewith{sessionId:<same-id>,cwd:<fixture-workspace>}succeeds and its ledger reachesterminal_ok. - Official global
session.resumewith the same ID, workspace and saved transcript path, separate resume key andreadinessTimeoutMs:60000returns exit 1 / envelopeerror.code=unavailable. - The exact launch failure marker reports
phase=startup,reason=failed,EEXIST: file already exists, open '<fixture-workspace>/.gjc/state/sdk/<same-id>.lifecycle.ready.json'. Its rollback evidence hasfenced,hostStopped,runtimeRemoved,brokerRegistrationReleasedall true. Subsequent exact inspect reportslive:false,deleted:false.
Important evidence distinction: the resume ledger observed here remains at
awaiting_ready; it does not haveterminal_error.response.error.code=EEXIST. EEXIST is from the correlated failure artifact, whereas the CLI exposesunavailable. This may also warrant a regression for failure-artifact-to-ledger settlement.Shared broker PID and model-profile fingerprints were unchanged. No prompt was submitted, no existing production operation was replayed, no SDK marker was removed, and no shared daemon was stopped. Gateway rollout remains blocked rather than masking this with a replacement session or permissive readiness. Please include the explicit close → same-ID resume sequence in the regression coverage, alongside detached-idle exit and mismatched ready/primary markers.
- added 4 commits that reference this issue
on Oct 3, 2026 - added a commit that references this issue
on Oct 3, 2026 Fixed on dev by #6265 (merge
e29e39e2): stale lifecycle ready markers are retired, sosession.resumeno longer trips over them. Closing by hand because closing keywords don't fire ondev. The fix ships with the next release.—
[repo owner's gaebal-gajae (clawdbot) 🦞]
Version: gjc 0.18.6, darwin arm64
Summary
When an SDK session host exits on its own (detached-idle shutdown after
SESSION_HOST_DETACHED_IDLE_GRACE_MS= 30 min), it leaves<cwd>/.gjc/state/sdk/<id>.lifecycle.ready.json(and<id>.lifecycle.json) on disk. A latersession.resumefor the same id spawns a new host. That host'swriteSessionLifecycleReadycreates the ready path withO_CREAT | O_EXCL, getsEEXIST, and dies during startup with an uncaught exception. The resume caller seesgjc sdk request failed: unavailable.After the first failure the session can never be resumed again: the failed launch rewrote
<id>.lifecycle.jsonwith its own pid and effect marker. The leftover ready file still holds the old host's marker, soreapDeadLifecycleMarkersskips the pair forever.Reproduction
session.create, cwd =~/workspace), then let it go idle until the host exits (detached_idle).gjc sdk session inspect <id>reportslive: false, deleted: false, and that~/workspace/.gjc/state/sdk/<id>.lifecycle.ready.jsonstill exists with a deadpid.EEXIST, and the session stays closed. Repeating step 3 fails the same way.Deleting only the stale
<id>.lifecycle.ready.json(its pid is gone) before step 3 makes the resume succeed (live: true). An older<id>.lifecycle.failure.<marker>.jsonfrom the failed attempt does not block it.Cause (line numbers from 0.18.6
src/)commands/sdk.ts:1270-1299(exitAfterSessionDisposal, reached fromstop("detached_idle")at 1301): disposes the session, checks that the endpoint<id>.jsonis gone, and exits. It never removes<id>.lifecycle.ready.json. The ready file is only removed on the publication-failure path (sdk/broker/lifecycle.ts:2001) or by broker-driven retirement and cleanup.sdk/broker/lifecycle.ts:1921-1925(writeSessionLifecycleReady):fs.open(readyPath, O_CREAT | O_EXCL | O_RDWR)for the placeholder. A leftover file from the previous host makes this throwEEXIST.sdk/broker/lifecycle.ts:6653: the launch path callsreapDeadLifecycleMarkers(launch.root)before spawning. The reaper rarely clears the leftover:1669-1670:inspectedcounts every directory entry before the name filter, so only the firstBROKER_DEAD_REGISTRATION_SWEEP_LIMIT(64,lifecycle.ts:8433) entries in readdir order are considered. In a state dir with 196 entries, only 1 of 24*.lifecycle.jsonmarkers fell in that window.1679: a marker younger thanDEAD_LIFECYCLE_MARKER_EXPIRY_MS(60 min,lifecycle.ts:124) is skipped, so a resume within an hour of the last launch always hits the leftover.1686: the ready file must carry the same effect marker as<id>.lifecycle.json. After one failed resume,lifecycle.ts:6813has rewritten<id>.lifecycle.jsonfor the crashed child. The markers then differ permanently, and the pair is never reaped.Observed state: 24 ready files in one workspace state dir, 21 of them with dead pids. The 3 sessions that had already crashed once all show mismatched markers and one failure record each.
Crash log excerpt (
~/.gjc/agent/gjc-crash.log)(9 such records over 3 days, same fingerprint.)
Suggested fixes (any one would do)
exitAfterSessionDisposal), remove the host's own<id>.lifecycle.ready.jsonthrough the same exact-identity unlink used on the failure path.session.resume/createof<id>, retire that id's own ready/marker pair when its recorded owner is proven exited (observeProcess(...) === "exited"). Do this regardless of age, the sweep window, or a marker mismatch caused by a crashed launch of the same id.reapDeadLifecycleMarkers, count only*.lifecycle.jsonentries toward the inspection limit.