Skip to content

session.resume of a self-closed host fails with EEXIST on <id>.lifecycle.ready.json #6261

Description

@Dabinlee1

Version: gjc 0.18.6, darwin arm64

Summary

When an SDK session host exits on its own (detached-idle shutdown after SESSION_HOST_DETACHED_IDLE_GRACE_MS = 30 min), it leaves <cwd>/.gjc/state/sdk/<id>.lifecycle.ready.json (and <id>.lifecycle.json) on disk. A later session.resume for the same id spawns a new host. That host's writeSessionLifecycleReady creates the ready path with O_CREAT | O_EXCL, gets EEXIST, and dies during startup with an uncaught exception. The resume caller sees gjc sdk request failed: unavailable.

After the first failure the session can never be resumed again: the failed launch rewrote <id>.lifecycle.json with its own pid and effect marker. The leftover ready file still holds the old host's marker, so reapDeadLifecycleMarkers skips the pair forever.

Reproduction

  1. Create an SDK session in a workspace (session.create, cwd = ~/workspace), then let it go idle until the host exits (detached_idle).
  2. Confirm that gjc sdk session inspect <id> reports live: false, deleted: false, and that ~/workspace/.gjc/state/sdk/<id>.lifecycle.ready.json still exists with a dead pid.
  3. Resume it:
    gjc sdk session raw global --op session.resume --idempotency-key <key> \
      --json-input '{"sessionId":"<id>","sessionPath":"<saved transcript .jsonl>","cwd":"~/workspace","readinessTimeoutMs":60000}'
    
  4. The new host crashes with EEXIST, and the session stays closed. Repeating step 3 fails the same way.

Deleting only the stale <id>.lifecycle.ready.json (its pid is gone) before step 3 makes the resume succeed (live: true). An older <id>.lifecycle.failure.<marker>.json from the failed attempt does not block it.

Cause (line numbers from 0.18.6 src/)

  • commands/sdk.ts:1270-1299 (exitAfterSessionDisposal, reached from stop("detached_idle") at 1301): disposes the session, checks that the endpoint <id>.json is gone, and exits. It never removes <id>.lifecycle.ready.json. The ready file is only removed on the publication-failure path (sdk/broker/lifecycle.ts:2001) or by broker-driven retirement and cleanup.
  • sdk/broker/lifecycle.ts:1921-1925 (writeSessionLifecycleReady): fs.open(readyPath, O_CREAT | O_EXCL | O_RDWR) for the placeholder. A leftover file from the previous host makes this throw EEXIST.
  • sdk/broker/lifecycle.ts:6653: the launch path calls reapDeadLifecycleMarkers(launch.root) before spawning. The reaper rarely clears the leftover:
    • 1669-1670: inspected counts every directory entry before the name filter, so only the first BROKER_DEAD_REGISTRATION_SWEEP_LIMIT (64, lifecycle.ts:8433) entries in readdir order are considered. In a state dir with 196 entries, only 1 of 24 *.lifecycle.json markers fell in that window.
    • 1679: a marker younger than DEAD_LIFECYCLE_MARKER_EXPIRY_MS (60 min, lifecycle.ts:124) is skipped, so a resume within an hour of the last launch always hits the leftover.
    • 1686: the ready file must carry the same effect marker as <id>.lifecycle.json. After one failed resume, lifecycle.ts:6813 has rewritten <id>.lifecycle.json for the crashed child. The markers then differ permanently, and the pair is never reaped.

Observed state: 24 ready files in one workspace state dir, 21 of them with dead pids. The 3 sessions that had already crashed once all show mismatched markers and one failure record each.

Crash log excerpt (~/.gjc/agent/gjc-crash.log)

2026-10-02T23:30:09.932Z pid=49293 [Uncaught Exception] Error: EEXIST: file already exists, open '~/workspace/.gjc/state/sdk/<id>.lifecycle.ready.json'
Error: EEXIST: file already exists, open '~/workspace/.gjc/state/sdk/<id>.lifecycle.ready.json'
    at async <anonymous> (node:fs/promises:179:56)
    at async rh (/$bunfs/root/cli-banstjrh.js:15:17336)
    at async <anonymous> (/$bunfs/root/sdk-kv5zafr2.js:20:1076)
    at async <anonymous> (/$bunfs/root/sdk-kv5zafr2.js:19:8924)
    at async Bp (/$bunfs/root/sdk-kv5zafr2.js:20:1047)
    at async run (/$bunfs/root/sdk-kv5zafr2.js:24:191)
    at async Ne (/$bunfs/root/cli-g6592wm8.js:8:225)
{"code":"EEXIST","path":"~/workspace/.gjc/state/sdk/<id>.lifecycle.ready.json","syscall":"open","errno":-17}
gjc-crash-record.v1 fp:a8ea25840233bd33f371d6a278a397b6 fpv:1

(9 such records over 3 days, same fingerprint.)

Suggested fixes (any one would do)

  1. On graceful host exit (exitAfterSessionDisposal), remove the host's own <id>.lifecycle.ready.json through the same exact-identity unlink used on the failure path.
  2. Before spawning for session.resume/create of <id>, retire that id's own ready/marker pair when its recorded owner is proven exited (observeProcess(...) === "exited"). Do this regardless of age, the sweep window, or a marker mismatch caused by a crashed launch of the same id.
  3. In reapDeadLifecycleMarkers, count only *.lifecycle.json entries toward the inspection limit.

Activity

  1. probepark commented on Oct 3, 2026

    @probepark
    Collaborator

    Triage: A (actionable, fix being prepared) · bug · P1

    Why: A detached-idle host exits without removing its own <id>.lifecycle.ready.json. The next host for that id fails on the O_CREAT | O_EXCL open in writeSessionLifecycleReady. After one failed resume the effect markers no longer match, so the session can't be resumed again unless someone deletes the file by hand.

    Checked against current dev (42eb13a); the line numbers in the report still match:

    • packages/coding-agent/src/commands/sdk.ts:1270-1299: exitAfterSessionDisposal checks the endpoint file but never removes the ready file.
    • packages/coding-agent/src/sdk/broker/lifecycle.ts:1921-1925: the exclusive open of the placeholder at readyPath.
    • lifecycle.ts:1667-1670: inspected is incremented before the .lifecycle.json name filter, so the 64-entry limit is used up by unrelated files.
    • lifecycle.ts:1679 and lifecycle.ts:1686: the 60-minute age gate and the exact-marker match. Together they keep the self-same-id leftover from being reaped.

    Planned fix scope: suggested fixes 1 and 2 together. Fix 1 stops new leftovers. Fix 2 retires a same-id ready/marker pair whose recorded owner is proven exited before spawning, which also repairs state dirs that are already stuck. Fix 3 (count only *.lifecycle.json toward the limit) is a one-liner and goes in the same change. Each fix gets a failing regression test first.

    Note: open PR #6148 also touches sdk.ts and lifecycle.ts (not this code path). Whichever lands second will need a rebase.

    Thanks for the detailed root-cause report.

  2. added
    bugSomething isn't working
    P1High priority
    on Oct 3, 2026
  3. gimso2x commented on Oct 3, 2026

    @gimso2x
    Contributor

    Additional independent reproduction on gjc 0.18.5, Linux aarch64 using an otherwise empty workspace. This is not limited to detached-idle exit:

    1. sdk session raw global --op session.create --idempotency-key <create-key> --json-input '{"cwd":"<fixture-workspace>","readinessTimeoutMs":60000}' succeeds. Exact inspect reports the original ID and live:true; create ledger reaches terminal_ok.
    2. Official global session.close with {sessionId:<same-id>,cwd:<fixture-workspace>} succeeds and its ledger reaches terminal_ok.
    3. Official global session.resume with the same ID, workspace and saved transcript path, separate resume key and readinessTimeoutMs:60000 returns exit 1 / envelope error.code=unavailable.
    4. The exact launch failure marker reports phase=startup, reason=failed, EEXIST: file already exists, open '<fixture-workspace>/.gjc/state/sdk/<same-id>.lifecycle.ready.json'. Its rollback evidence has fenced, hostStopped, runtimeRemoved, brokerRegistrationReleased all true. Subsequent exact inspect reports live:false, deleted:false.

    Important evidence distinction: the resume ledger observed here remains at awaiting_ready; it does not have terminal_error.response.error.code=EEXIST. EEXIST is from the correlated failure artifact, whereas the CLI exposes unavailable. This may also warrant a regression for failure-artifact-to-ledger settlement.

    Shared broker PID and model-profile fingerprints were unchanged. No prompt was submitted, no existing production operation was replayed, no SDK marker was removed, and no shared daemon was stopped. Gateway rollout remains blocked rather than masking this with a replacement session or permissive readiness. Please include the explicit close → same-ID resume sequence in the regression coverage, alongside detached-idle exit and mismatched ready/primary markers.

  4. added a commit that references this issue on Oct 3, 2026
  5. Yeachan-Heo commented on Oct 3, 2026

    @Yeachan-Heo
    Owner

    Fixed on dev by #6265 (merge e29e39e2): stale lifecycle ready markers are retired, so session.resume no longer trips over them. Closing by hand because closing keywords don't fire on dev. The fix ships with the next release.

    —
    [repo owner's gaebal-gajae (clawdbot) 🦞]

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P1High prioritybugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions