Skip to content

fix(cli): refuse cascade sync while a server holds the memory root - #458

Merged
gloryfromca merged 4 commits into
mainfrom
fix/cascade-sync-single-writer
Sep 24, 2026
Merged

gloryfromca merged 4 commits into
mainfrom
fix/cascade-sync-single-writer

Conversation

@gloryfromca

@gloryfromca gloryfromca commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

Summary

Every cascade command that writes the index — sync, fix --apply, rebuild — now acquires the OME jobstore lock (the same file and flags everos server uses) and holds it for its whole run. While a server or another exclusive CLI phase holds the root, the command exits 3 before opening anything; a server that starts mid-run fails at its own lock instead of becoming a second writer. status and fix (listing) are read-only and keep working next to a server.

Why: sync was deliberately exempted from the lock in #384 on the belief that draining from a second process is safe. The 10-hour Windows soak (PR #454) ran two cascade sync processes next to the server and ended with 1 158 / 1 219 / 1 241 duplicate global ids in episode / atomic_fact / foresight (4.3–4.6 % extra rows): same md_path, same content, updated_at ~14 s apart. Mechanism: the server's LanceDB connection keeps its own snapshot (read_consistency_seconds = None), so when another process inserts a row the server's merge_insert match phase still sees "no row" and inserts it again; two insert-only transactions do not conflict in Lance and no unique-key constraint exists. No consistency interval closes that window — a single writer does.

Behaviour change

  • everos cascade sync and everos cascade fix --apply refuse (exit 3) while a server holds the memory root. Forcing a path to re-index while a server runs is no longer possible from the CLI; the watcher covers edits, and an API re-enqueue endpoint would be the way to bring that back.
  • The refusal names the lock holder and the two cases where the server is not syncing (EVEROS_DISABLE_CASCADE=1, quiesced) so the operator knows to stop it first.
  • The soak harness's "CLI storm" (two cascade sync processes next to the server, .work_context/lancedb_soak/harness) is retired by this change: every iteration exits 3. Its purpose — exercising cross-process commit races — is what this PR removes on purpose.

Verification

  • Integration tests hold the real lock file the way a server does: sync and fix --apply exit 3 with get_engine patched to fail if reached (refusal precedes any DB open); fix and status still run; sync observes the lock held during its own drain (ome_lock_is_free() is False inside sync_once) and released after. Red without the held lock (3 failed), green with it.
  • tests/integration/test_cascade_cli_integration.py + tests/unit/test_entrypoints/test_cli + tests/unit/test_memory/test_cascade: 295 passed. make lint green.
  • Duplicate evidence: .work_context/lancedb_soak/results/windows_run1/summary.md (local), probe dup_probe.py.

Docs

docs/cascade_runbook.md: the sync section no longer says "safe to run in parallel with a live server"; the rebuild note no longer calls rebuild "the one command not safe next to a server"; the reference-file workaround says to re-save SKILL.md (or stop the server); docs/how-memory-works.md no longer suggests forcing the queue with sync for read-your-write.

Review round

An adversarial review of the first commit found: probe-and-release instead of a held lock, fix --apply ungated, the refusal claiming the server "already syncs" (false for EVEROS_DISABLE_CASCADE=1 / quiesced), a test that could not tell where the guard sat, and five doc passages still prescribing sync next to a server. All addressed in 2660d43.

🤖 Generated with Claude Code

zhanghui and others added 3 commits September 24, 2026 12:28
`cascade rebuild` already refuses to run next to a live server (the OME
jobstore lock, exit code 3); `cascade sync` was deliberately left out in
#384 because draining the queue from a second process was believed safe.
The 10-hour Windows soak (two `cascade sync` processes alongside the
server) shows it is not: the server reads its own LanceDB snapshot
(`read_consistency_seconds` defaults to None) and cannot see what the CLI
process just committed, so both decide a row is new and both insert it.
After 10 h every table carried 4-5 % duplicate global ids, same md_path,
same content, `updated_at` about 14 s apart.

With a server running the CLI is redundant anyway — the daemon's watcher
already syncs every change — so `sync` now takes the same lock probe and
exits 3 with a message that says so. Without a server nothing changes. The
runbook's "unlike cascade sync" sentence and the "run sync after batch
edits" advice are corrected to match.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Adversarial review of the first commit found the guard too narrow to
deliver a single writer: it probed the lock once and released it, so a
server started mid-drain became a second writer anyway; `cascade fix
--apply` drains through the same orchestrator and was not gated; and the
refusal text claimed the server "already syncs", which is false for a
server started with EVEROS_DISABLE_CASCADE=1 or one that has been
quiesced.

`_runtime(exclusive=True)` now acquires the OME jobstore lock — the same
file and flags the engine uses — and holds it until the command's runtime
tears down, so a late server fails at its own lock instead of joining in.
`sync`, `fix --apply` and `rebuild` pass it; `status` and `fix` (listing)
stay read-only and keep working next to a server. The refusal names the
lock holder and the two exceptions, and exits 3 before anything is opened.

Tests hold the real lock file the way a server does: `sync` and
`fix --apply` exit 3 with `get_engine` patched to fail if reached; `fix`
and `status` still run; `sync` observes the lock held during its own
drain and released after. The runbook's "safe to run in parallel with a
live server" paragraph, the rebuild note, the reference-file workaround
and how-memory-works.md's "force the queue" advice are corrected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@arelchan
arelchan self-requested a review September 24, 2026 08:44
@gloryfromca
gloryfromca merged commit 243bddd into main Sep 24, 2026
10 checks passed
@gloryfromca
gloryfromca deleted the fix/cascade-sync-single-writer branch September 24, 2026 08:52
This was referenced Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants