Skip to content

fix(lancedb): treat a lance spill failure as retryable - #463

Merged
gloryfromca merged 1 commit into
mainfrom
fix/lance-transient-read-errors
Sep 24, 2026
Merged

gloryfromca merged 1 commit into
mainfrom
fix/lance-transient-read-errors

Conversation

@gloryfromca

Copy link
Copy Markdown
Member

Summary

The cascade worker retries ExternalServiceError with backoff and files any other exception as unrecoverable (mark_failed(retryable=False), surfaces in cascade fix). lancedb raises LanceError(IO): Execution error: Spill has sent an error — DataFusion's sort/merge spill to the OS temp dir failing mid-query — as a bare RuntimeError, so every occurrence marked an md file permanently failed.

On the Windows soak box (ThinkPad, nine concurrent clients) it fired 187 / 180 / 34 times across the three soak runs, never on an idle box, and the same rows project fine on the next attempt; roughly 200 files per run were left needing a manual cascade fix. The run-1 report attributed the unrecoverable count to fuzzed malformed markdown only; this class was hiding under it.

Fix: translate that one phrase into VectorStoreBusyError inside the read deadline wrapper (LanceRepoBase._deadline) — the observed path is worker → find_by_md_path / delete_by_md_path → repository → lancedb query. Other IO errors (No space left, missing file) propagate unchanged and stay permanent.

Area

  • architecture (storage)

Verification

  • tests/unit/test_core/test_persistence/test_lancedb/test_transient_execution_errors.py: the phrase becomes VectorStoreBusyError; an unrelated IO error does not; the marker predicate. Emptying the marker tuple fails the translation test.
  • test_core/test_persistence/test_lancedb + test_cascade: green. ruff green.
  • Not reproduced locally: the spill failure only appears on the loaded Windows box; the soak on the integration branch with this change is the live check (its unrecoverable_total should drop to the fuzz-only baseline).

Checklist

  • tests added, mutation checked

Notes for Reviewers

Why not a broader LanceError(IO) match: disk-full and vanished-file errors carry the same prefix and must stay permanent. Why not in _locked too: the write path was not where the error was observed; widening is a one-line follow-up if a soak shows it there.

🤖 Generated with Claude Code

The cascade worker retries ExternalServiceError with backoff and files any
other exception as unrecoverable. lancedb raises 'LanceError(IO):
Execution error: Spill has sent an error' -- DataFusion's sort/merge spill
to the OS temp dir failing mid-query -- as a bare RuntimeError, so every
occurrence marked an md file permanently failed. On the Windows soak box
it fired 187 / 180 / 34 times across three runs under nine concurrent
clients, never on an idle box, and the same rows project fine on retry;
about 200 files per run were left for a manual 'cascade fix'.

Translate that one phrase into VectorStoreBusyError inside the read
deadline wrapper (the observed path: worker -> find/delete by md_path ->
repository -> lancedb query). Other IO errors keep propagating unchanged.

Tests: the phrase is translated, an unrelated IO error is not; removing the
marker fails the translation test.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@gloryfromca
gloryfromca enabled auto-merge (squash) September 24, 2026 08:54
@arelchan
arelchan self-requested a review September 24, 2026 08:59
@gloryfromca
gloryfromca merged commit 4fd3472 into main Sep 24, 2026
10 checks passed
@gloryfromca
gloryfromca deleted the fix/lance-transient-read-errors branch September 24, 2026 09:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants