You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[Bug]: Second process's close() discards the first process's committed writes (single-file storage, 0.5.42) #405
Two processes that open the same on-disk database both succeed, and the one that
closes last overwrites the whole store with its own in-memory view. Writes the
other process already committed are gone. Neither process raises an error, and
neither execute() nor close() returns anything that indicates a conflict.
Expected: either the writes from both processes are retained, or the second GrafeoDB(path) fails with a clear "already open" error.
Observed: the last writer to close silently wins.
This looks like the multi-process consequence of the design described in #327 —
single-file .grafeo storage checkpoints the in-memory LPG store into the
container at close(). With one process that is correct. With two, the second
checkpoint is built from a store that never saw the first process's commits, so
writing it discards them. #395 reports a related symptom of the same
close()-owns-persistence model from a single process.
The interleaving, with the barrier ordering from the reproduction below:
sequenceDiagram
participant P1 as Process 1
participant Disk as graph.db (single file)
participant P2 as Process 2
P1->>Disk: GrafeoDB(path)
Note over P1: in-memory store = {}
P2->>Disk: GrafeoDB(path) — succeeds, no error
Note over P2: in-memory store = {}
P1->>P1: INSERT A — commits, returns OK
P1->>Disk: close() → snapshot {A}
P2->>P2: INSERT B — commits, returns OK
P2->>Disk: close() → snapshot {B}
Note over Disk: file now holds {B}.<br/>A is gone: P2's snapshot was built<br/>from a store that never saw it.
Loading
Because the checkpoint writes the whole store rather than a delta, close() is
effectively "the file is now what I have in memory" rather than "add my work to
the file". With a single process that distinction never shows.
Two details that make this worse than a plain "don't do that":
It is silent. In a 4-writer race, all writers reported success and 100 of
200 writes were missing. Callers have no signal to retry or fail on.
The Python binding cannot avoid it.grafeo.GrafeoDB exposes only GrafeoDB() and GrafeoDB(path); there is no way to select WalDirectory
storage, and dir(grafeo) exports no storage-format or config type. Python
users are on the affected path by default with no documented opt-out.
I could not find a statement in README.md, docs/, or CONTRIBUTING.md that an
on-disk database is single-process-only, so I am reporting it as a bug rather
than as a documentation gap. If single-process is the intended contract, a
loud failure at open plus a line in the docs would be enough to close this.
Either resolution would work for my use case:
Make concurrent opens safe (shared WAL, or reload-before-checkpoint), or
Fail the second open with an explicit error, the way SQLite and other
embedded engines do, and document the constraint.
An error at open is the more valuable of the two, because silent loss is the
part that is hard to defend against from the outside.
AI Disclosure: This issue was prepared with the assistance of Claude
(Anthropic) running in Claude Code. The failure was found while I was
benchmarking Grafeo as an embedded store for a personal project; the AI
assistant ran the experiments, reduced the race to the deterministic barrier
version above, and drafted this report. I ran the reproduction myself and
verified the reported output before submitting. No fix or patch is proposed
here.
How to reproduce
Deterministic — reproduces 3 out of 3 runs. Two processes are interleaved with
file barriers so the ordering is fixed:
P1 open -> P2 open -> P1 INSERT A, close -> P2 INSERT B, close -> read back
exit [0, 0] expected ['A','B'] got ['B'] -> DATA LOSS
exit [0, 0] expected ['A','B'] got ['B'] -> DATA LOSS
exit [0, 0] expected ['A','B'] got ['B'] -> DATA LOSS
Under an unsynchronised race the loss is partial and varies, which is how I hit
it originally. Four processes each doing 50 distinct MERGEs against one
database, expecting 200 nodes:
expected 200 got 100 -> LOST 100 WRITES
expected 200 got 100 -> LOST 100 WRITES
expected 200 got 129 -> LOST 71 WRITES
expected 200 got 100 -> LOST 100 WRITES
expected 200 got 100 -> LOST 100 WRITES
Five runs, five failures, every child process exiting 0. The 129 run shows the
result can also be an interleaving of two snapshots rather than a clean
last-writer-wins.
With two writers instead of four it still occurs, just less often — 1 of 3 runs
in my testing — which is the awkward part: light concurrency looks fine in
testing and loses data later.
The unsynchronised race script is in the collapsed block below.
repro_race.py — four writers, no synchronisation
"""Minimal repro: two processes writing to one on-disk Grafeo database. python repro.py # runs the whole thingExpected: 100 nodes (2 writers x 50 distinct MERGEs).Observed: fewer, with no error raised by either writer."""importshutil, subprocess, sys, tempfile, pathlibimportgrafeoDB=pathlib.Path(tempfile.gettempdir()) /"grafeo_mp_repro.db"N=50iflen(sys.argv) >1: # child: writer roletag=sys.argv[1]
db=grafeo.GrafeoDB(str(DB))
foriinrange(N):
db.execute(f"MERGE (n:Item {{id:'{tag}-{i}'}})")
db.close()
sys.exit(0)
shutil.rmtree(DB, ignore_errors=True)
DB.unlink(missing_ok=True)
kids= [subprocess.Popen([sys.executable, __file__, t]) fortin ("A", "B", "C", "D")]
codes= [k.wait() forkinkids]
db=grafeo.GrafeoDB(str(DB))
got=list(db.execute("MATCH (n:Item) RETURN count(n) AS n"))[0]["n"]
print(f"grafeo {grafeo.__version__ifhasattr(grafeo,'__version__') else'?'}"f" child exit codes {codes} expected {4*N} got {got}"f" -> {'OK'ifgot==4*Nelse'LOST %d WRITES'% (4*N-got)}")
Version
0.5.42, Python (grafeo 0.5.42 from PyPI, cp312-abi3 wheel)
Python 3.14.7, Linux x86_64, glibc 2.44
What happened?
Two processes that open the same on-disk database both succeed, and the one that
closes last overwrites the whole store with its own in-memory view. Writes the
other process already committed are gone. Neither process raises an error, and
neither
execute()norclose()returns anything that indicates a conflict.Expected: either the writes from both processes are retained, or the second
GrafeoDB(path)fails with a clear "already open" error.Observed: the last writer to close silently wins.
This looks like the multi-process consequence of the design described in #327 —
single-file
.grafeostorage checkpoints the in-memory LPG store into thecontainer at
close(). With one process that is correct. With two, the secondcheckpoint is built from a store that never saw the first process's commits, so
writing it discards them. #395 reports a related symptom of the same
close()-owns-persistence model from a single process.
The interleaving, with the barrier ordering from the reproduction below:
sequenceDiagram participant P1 as Process 1 participant Disk as graph.db (single file) participant P2 as Process 2 P1->>Disk: GrafeoDB(path) Note over P1: in-memory store = {} P2->>Disk: GrafeoDB(path) — succeeds, no error Note over P2: in-memory store = {} P1->>P1: INSERT A — commits, returns OK P1->>Disk: close() → snapshot {A} P2->>P2: INSERT B — commits, returns OK P2->>Disk: close() → snapshot {B} Note over Disk: file now holds {B}.<br/>A is gone: P2's snapshot was built<br/>from a store that never saw it.Because the checkpoint writes the whole store rather than a delta,
close()iseffectively "the file is now what I have in memory" rather than "add my work to
the file". With a single process that distinction never shows.
Two details that make this worse than a plain "don't do that":
200 writes were missing. Callers have no signal to retry or fail on.
grafeo.GrafeoDBexposes onlyGrafeoDB()andGrafeoDB(path); there is no way to selectWalDirectorystorage, and
dir(grafeo)exports no storage-format or config type. Pythonusers are on the affected path by default with no documented opt-out.
I could not find a statement in README.md, docs/, or CONTRIBUTING.md that an
on-disk database is single-process-only, so I am reporting it as a bug rather
than as a documentation gap. If single-process is the intended contract, a
loud failure at open plus a line in the docs would be enough to close this.
Either resolution would work for my use case:
embedded engines do, and document the constraint.
An error at open is the more valuable of the two, because silent loss is the
part that is hard to defend against from the outside.
AI Disclosure: This issue was prepared with the assistance of Claude
(Anthropic) running in Claude Code. The failure was found while I was
benchmarking Grafeo as an embedded store for a personal project; the AI
assistant ran the experiments, reduced the race to the deterministic barrier
version above, and drafted this report. I ran the reproduction myself and
verified the reported output before submitting. No fix or patch is proposed
here.
How to reproduce
Deterministic — reproduces 3 out of 3 runs. Two processes are interleaved with
file barriers so the ordering is fixed:
Output, three consecutive runs:
Under an unsynchronised race the loss is partial and varies, which is how I hit
it originally. Four processes each doing 50 distinct
MERGEs against onedatabase, expecting 200 nodes:
Five runs, five failures, every child process exiting 0. The 129 run shows the
result can also be an interleaving of two snapshots rather than a clean
last-writer-wins.
With two writers instead of four it still occurs, just less often — 1 of 3 runs
in my testing — which is the awkward part: light concurrency looks fine in
testing and loses data later.
The unsynchronised race script is in the collapsed block below.
repro_race.py — four writers, no synchronisation
Version