Skip to content

Commit 8024fe5

Browse files
gloryfromcazhanghuiclaude
authored
fix: harden the stall defect class found auditing the 1.2.2 fix (#392)
* fix(lancedb): bound reads and unhook the husk sweep from the write lock Two instances of the same defect class the write-lock deadline work left behind: an await in a scheduler-gated path with nothing bounding it. Reads (count / get_by_id / find_where / find_where_paginated / search) were skipped last round on the reasoning that a read takes no lock and so blocks no writer. True, but incomplete: the cascade drain loop reads on every batch and advances strictly one batch at a time, so a read that never returns stops the whole md -> LanceDB projection. Claimed rows stay `processing` forever (claim_pending_batch only takes `pending`, orphan recovery runs once at startup), and /health keeps reporting healthy because a hang raises nothing. Budget 60s, ~1000x the measured 62ms flat scan over 117k rows. The empty-index-dir sweep ran inside the prune critical section under a docstring contract requiring the write lock. That contract could not hold: the sweep runs via asyncio.to_thread, and a deadline cancels the future, not the thread, so an orphan sweep outlives the lock -- and Path.iterdir is a lazy os.scandir, so it can yield a dir created after the scan began. It could therefore rmdir a directory a concurrent create_index had just made, leaving the table with no FTS index and every search on that kind 500ing. Safety now comes from an age filter (skip dirs younger than 300s), which holds regardless of lock ownership; the sweep moved out of the critical section so a slow filesystem walk can no longer overrun the prune budget. Mutation-verified: moving the table handle back outside the read deadline hangs the new test; moving the sweep back inside the lock fails it; setting the age floor to 0 fails the fresh-dir test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(cascade): supervise background loops and unmask the optimize alert The drain / heartbeat / rebuild loops were plain create_task coroutines. One uncaught exception ended that loop permanently: nothing restarted it, and because the worker holds a strong reference to the task the interpreter never printed "Task exception was never retrieved" either (that fires on GC). The loop's job just stopped happening with zero output. _run_loop had an inner try; its two siblings did not. Each loop now runs under _supervise: log, wait, restart with escalating backoff (5s / 15s / 45s), then request process exit via SIGTERM so a restarting supervisor (systemd Restart=always, Docker restart: unless-stopped, a k8s Deployment) can recover it. SIGTERM rather than os._exit so the ASGI server runs its graceful-shutdown path. A done-callback is the last-resort observer for the supervisor itself ending unexpectedly. Separately, the fallback rebuild reset the same counter the health verdict reads, so the optimize-failure threshold was effectively unreachable: a table failing 100% of the time cycled 1..5 -> 0 -> 1.. and the threshold value existed only during the sub-second rebuild, ~1% observable against a 30s scrape. cascade.healthy stayed green while the table never reclaimed a version. The rate limiter moves to failures_since_fallback; only a successful optimize clears the alert streak. Same shape as the run7 cross-kind max() masking bug -- a remediation path refreshing the signal meant to report it. Mutation-verified: dropping the restart budget, removing the exit request, and restoring the counter reset each turn the corresponding test red. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(persistence): bound and announce the memory-root lock wait Acquisition polls with LOCK_NB instead of blocking inside a worker thread. A blocking flock could be neither bounded nor cancelled: cancelling the awaiting coroutine leaves the thread to acquire the lock later with nobody left to release it, which is strictly worse than waiting. The wait itself is correct by design -- the second process is supposed to wait, then find the migration already done -- and flock is released by the kernel on process exit, so a dead holder never wedges it. What was wrong is that it had no upper bound and emitted nothing: a server startup landing on a held lock looked like a hang whose last log line was lifespan_provider_startup name=lancedb. It now logs memory_root_lock_waiting on first contention, reports how long it waited on success, and gives up after timeout_seconds (default 300s, generous because the legitimate holder is an O(rows) migration). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(cascade): wait for projection quiescence in the lap-race scenario test_lap_append_during_handler_no_loss asserted no loss after _wait_path_done, whose settle window is 0.1s. That is a bet that the filesystem event for the appends which landed *during* a handler invocation has already been delivered — a terminal row does not mean the file is fully projected, because the handler read the md at whatever length it had then, marked the row done, and the rest arrive on a later event. The bet holds on macOS/fsevents and lost on a loaded Linux runner (md=30 lance=17), failing the assertion for a reason unrelated to the behaviour under test. Waits for quiescence instead: terminal row + empty pending queue + a projected count unchanged across three consecutive polls. Strictly stronger than the old condition, and real loss still fails — the count just converges below the md entry count and stays there. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(lancedb): replace indexes in place and retry a lost rebuild race rebuild_indexes dropped every index and recreated it. The docstring justified that with "LanceDB transparently falls back to brute-force scan", which is true for vector search and false for FTS: with no inverted index a BM25 query raises "Cannot perform full text search unless an INVERTED index has been created". The recall legs are gathered without return_exceptions, so one failing leg fails the whole search request -- every keyword search landing in the window returned 500. Measured: 55 failures across 3 rebuilds with drop+create, 0 with create_index(replace=True). Replacing also collapses the live fragment set identically (7 index files back to 4 after 25 optimize beats), so nothing the rebuild existed for is lost. Only indexes on columns that are no longer indexed at all are still dropped -- nothing queries those. A rebuild that loses the manifest race is now retried rather than deferred to the next 12h sweep. Lance marks the conflict Retryable and means it: another process committed first. Retries are recorded as a deadline on the kind (10min / 30min / 3h) and picked up by the rebuild loop, not slept through -- the loop walks kinds sequentially, so sleeping would park every later kind behind the backoff (7 kinds x 3h outlasts the cadence itself). The loop tick is min(60s, cadence) so a shorter configured interval is not quantised. Removes the empty-index-dir sweep. cleanup_older_than deletes the files under a superseded _indices/<uuid>/ but leaves the directory, and everos was removing those with its own rmdir. No LanceDB contract says an empty index dir is garbage, so this is being raised upstream instead. Note it is a separate gap from index *files* not being reclaimed under delete_unverified=False (260MB retained on a 19k-row soak table) -- solving that still leaves the empty dirs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(changelog): put only the sweep under Removed The rebuild-bound and lock-wait entries were swallowed into the Removed section when the husk-sweep entry was inserted above them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(lancedb): bound the husk sweep by lance's own unverified threshold Restores the empty-index-dir sweep with a safety argument instead of a self-chosen number. Upstream context, read from lance's cleanup.rs: it unlinks a superseded index's files but never the directory, and contains no rmdir at all. That is structural, not an oversight -- lance targets object stores, where paths are flat keys and an empty directory does not exist. Only a local filesystem materialises them, where they accumulate as inodes (a soak run reached 13061 dirs, 98% empty) and slow every directory scan. Three independent guarantees, in order of strength: 1. rmdir cannot delete data. The kernel refuses it on a non-empty directory, so no file can be lost whatever the rest of the logic decides -- and because the check *is* the operation, there is no check-then-act window to race. 2. Live indexes are excluded by UUID, read from list_indices(). 3. Anything else must outlive UNVERIFIED_THRESHOLD_DAYS = 7, which is lance's own bound for deciding an unreferenced index UUID is dead rather than an index build in progress. Matching it means the sweep can never be more aggressive than lance itself; the previous 300s was our invention, and that is what made it indefensible. Each guarantee is pinned by the same test and mutation-verified: dropping the live-UUID check, zeroing the age gate, and swapping rmdir for a recursive delete each turn it red. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(cascade): make the maintenance cadences configurable The four cadences were already constructor arguments on CascadeWorker, but CascadeConfig did not carry them and none of the three production construction paths passed a config, so the module defaults were unreachable from outside the code. That is why no soak run shorter than half a day could exercise the 12h rebuild sweep: not a missing parameter, a config layer that dropped it. Adds CascadeSettings ([cascade] in default.toml) and CascadeConfig.from_settings, which the orchestrator now uses when no config is passed -- so the CLI, backfill and server paths all pick settings up at once. Deliberately not exposed: the read / write / prune / rebuild deadlines. Those are hang-catchers sized from measured durations, and both directions are worse -- too low manufactures failures on a healthy table, too high leaves a wedged one invisible for longer. Cadences depend on write volume and are a real tuning axis; deadlines are not. A test pins the exposed field set so a later change has to state its intent. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(cascade): cover the AgentCase handler agent_case was the one business kind with no handler test, and the storage soak never writes it either -- four of the seven tables stay at zero rows there -- so its md -> row contract was unexercised from both directions. Covers what makes this kind different from its daily-log siblings: it lives on the agent track, and it embeds task_intent only while approach is BM25-indexed but deliberately never sent to the embedder. Plus the branches every handler shares: soft-dependency embedding (no provider -> vector=None, row still written for keyword-only deployments), optional KeyInsight, the content_sha256 short-circuit that stops the 30s scanner re-embedding untouched files, edit detection, and delete-by-path. Mutation-verified: routing approach into the embedder, and dropping section:TaskIntent from content_change_keys, each turn the relevant test red. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(lancedb): record the empty-dir cost and what caps it Two constants now carry the arithmetic instead of leaving it in a chat log. _HUSK_MIN_AGE_SECONDS states what the 7-day gate costs: nothing reclaims an empty index dir before then, each is an inode plus a 4KB block, and at ceiling load that is ~890k dirs / ~3.6GB / 14% of a default 98GB ext4's inodes at the 7-day steady state. Also that this is the worst case and needs sustained saturation -- a single-user deployment sits four orders of magnitude below it -- and that only ext4 has a fixed inode budget (APFS and xfs allocate dynamically, Windows is out of scope). DEFAULT_OPTIMIZE_MIN_INTERVAL_SECONDS gets two things it never said. First, it is not a visibility delay: a row is searchable as soon as its upsert commits, because LanceDB flat-scans the unindexed tail -- verified to cover BM25, not just vector and scalar, which was the leg worth doubting given a missing FTS index hard-fails rather than degrading. Sparse writes do not wait at all, since the scheduler uses max(0, interval - elapsed). Second, it is the ceiling on index-directory growth: past roughly one write per table per interval the beats coalesce, so the accrual rate is capped by this interval rather than by write volume, and raising it lowers the cost proportionally. That makes it the knob to reach for if the empty dirs ever bite -- not the husk threshold. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(cascade): reset the loop-restart budget after a stable run The supervisor's restart budget was spent per process lifetime: three strikes ever, regardless of how long the loop ran healthily between them. A loop hitting one recoverable transient every few days — each cleared by a single restart — would still pool those strikes and SIGTERM a healthy server weeks in, on the 4th, which punishes exactly the case supervision exists to absorb. The budget now counts consecutive quick crashes: a body that ran at least 60s before raising starts a fresh incident with the full ladder. A deterministic crash-on-entry still exhausts the budget in ~65s. Same windowed counting as systemd StartLimitIntervalSec / Erlang max_restarts-per-max_seconds. Also corrects the husk accrual numbers in the optimize-cooldown docstring to the ~14-day effective reclaim horizon (see the sibling lancedb commit for why the age gate doubles). Mutation-verified: with the reset removed, the new test exits after run 3 instead of surviving to run 5. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(lancedb): keep a husk-sweep timeout from failing the prune By the time the sweep runs, the cleanup commit — the thing prune exists for — has already succeeded, and the sweep is best-effort by contract. Letting its deadline escape prune() billed the failure to the wrong account: the optimize scheduler counted a prune failure (feeding the fallback-rebuild threshold) and the prune-staleness clock stopped advancing, so both alarms reported a cleanup stall that did not happen. Same defect shape as the alert counter the fallback rebuild used to zero: an auxiliary path corrupting the main signal's ledger. Reachable, not theoretical: sweep time is proportional to dir count (~35us/dir measured) and the ceiling-load steady state sits right at the 60s budget. The timeout is tolerable exactly because it is now swallowed — and the orphaned worker thread finishes the walk anyway, so the reclamation still happens. Also corrects the age-gate docstrings: the gate reads st_mtime, which POSIX bumps when lance's cleanup empties the husk, so the effective reclaim horizon is file wait + 7 days (~14 days total) and the ceiling-load steady state is ~1.8M dirs / ~7GB, twice the previously recorded figure. The "never more aggressive than lance" property is unaffected (it is strictly more conservative). Mutation-verified: with the try/except removed, the new test fails on the escaping VectorStoreBusyError. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(persistence): move the lock timeout an order above the legit hold 300s sat at the edge of what the docstring itself calls a legitimate hold (a large migration is minutes of O(rows) work), so the worst honest migration turned every waiting process's startup into a LockError crash. Now 1800s: the wait has been visible since the first poll (memory_root_lock_waiting), and against the one case the bound exists for — a holder alive but wedged — giving up at 5 minutes buys nothing over 30, because the timeout's job is diagnosis, not recovery. The timeout message now says which way to look: the kernel releases a dead holder's flock automatically, so reaching the timeout means the holder is alive — inspect that process instead of retrying this one. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): fold the review fixes into the unreleased notes Supervisor bullet gains the per-incident budget, the sweep bullet gains the ~14-day effective horizon and the swallowed timeout, and the lock bullet records the 30min default with its sizing rationale. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(lancedb): pin that only a sweep timeout is absorbed by prune The sibling test covers one side -- a deadline miss must not bill prune's ledger. Nothing covered the other: widening the catch to `except Exception` passes every other test in the file, and would turn a genuine fault in _remove_empty_index_dirs (a TypeError after a signature change, a permission error on the index dir) into a silent removed = 0 with no signal anywhere. That is the failure shape this module keeps being audited for, so the narrowness of the catch needs its own guard. Found by mutating the catch rather than by reading it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(cascade): bound the optimize runner's wait on a rebuild The two maintenance jobs park on each other -- whichever arrives second waits on the first -- so the two waits are one hazard seen from opposite ends. The rebuild side was bounded; this side was left open on the argument that rebuild_indexes carries its own 300s deadline. That deadline covers its critical section, not the task's dispatch and teardown around it, so the transitive bound was never real. While the runner waits, its per-kind task slot stays occupied, every _schedule_optimize call short-circuits on it, and that table silently stops being pruned -- the same shape as the stall this branch has been chasing. Bounded at 180s. On expiry the beat is skipped rather than run: compacting under a live rebuild is the interleaving the wait exists to prevent, and both commit on the same manifest. Writes keep the dirty flag set, so the next beat retries. Mutation-verified: replacing the timeout with a plain await hangs the new test until its own guard fires. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: zhanghui <zhanghui@shanda.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
1 parent d5668dc commit 8024fe5

15 files changed

Lines changed: 1950 additions & 168 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 127 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -7,6 +7,133 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
77

88
## [Unreleased]
99

10+
### Added
11+
12+
- **`[cascade]` settings section** — the four maintenance cadences
13+
(`optimize_heartbeat_seconds`, `optimize_prune_interval_seconds`,
14+
`optimize_prune_retention_seconds`, `optimize_rebuild_interval_seconds`) are
15+
now configurable. They were already constructor arguments on `CascadeWorker`,
16+
but `CascadeConfig` did not carry them and no production path passed one, so
17+
the defaults were unreachable — which is why the 12h rebuild sweep could not
18+
be exercised by any soak run shorter than half a day. The deadlines that
19+
bound a hung call are deliberately **not** exposed: they are hang-catchers
20+
sized from measured durations, where too low manufactures failures on a
21+
healthy table and too high leaves a wedged one invisible for longer. Note
22+
`optimize_prune_retention_seconds` has a second effect worth reading before
23+
tuning — it also decides how long index files keep a manifest naming them,
24+
and below LanceDB's 7-day unverified window they then wait out the full 7
25+
days.
26+
27+
### Fixed
28+
29+
- **Reads now carry a deadline** (`count` / `get_by_id` / `find_where` /
30+
`find_where_paginated` / `search`). The write-side deadline work skipped them
31+
on the reasoning that a read takes no lock and so blocks no writer — true, but
32+
the cascade drain loop reads on every batch and advances strictly one batch at
33+
a time, so a read that never returns stops the **whole md → LanceDB
34+
projection**: claimed rows stay `processing` forever, nothing new is indexed,
35+
and `/health` still reports healthy because a hang raises nothing. Budget 60s
36+
(~1000x the measured 62ms flat scan over 117k rows); expiry raises the
37+
retryable `VectorStoreBusyError`.
38+
- **Background loops are supervised.** The drain / heartbeat / rebuild loops were
39+
plain `create_task` coroutines: one uncaught exception ended that loop
40+
permanently, and because the worker holds a strong reference to the task the
41+
interpreter never printed "Task exception was never retrieved" either — the
42+
loop's job simply stopped happening with zero output. Each now runs under a
43+
supervisor that logs and restarts with escalating backoff (5s / 15s / 45s),
44+
then asks the process to exit via `SIGTERM` so a restarting supervisor
45+
(systemd, Docker, k8s) can recover it. The restart budget counts consecutive
46+
*quick* crashes, not crashes over the process lifetime — a body that ran 60s+
47+
before raising starts a fresh incident, so independent transients days apart
48+
cannot pool into a process exit (same windowed counting as systemd's
49+
`StartLimitIntervalSec`). A done-callback covers the case the supervisor
50+
itself ends unexpectedly.
51+
- **The optimize-failure alert is reachable again.** The fallback rebuild reset
52+
the same counter the health verdict reads, so a table failing 100% of the time
53+
cycled `1..5 → 0 → 1..` and the threshold value existed only during the
54+
sub-second rebuild — roughly 1% observable against a 30s scrape, so
55+
`cascade.healthy` stayed green while the table never reclaimed a version. The
56+
rate limiter now lives in its own counter (`failures_since_fallback`); only a
57+
successful optimize clears the alert streak. Same shape as the cross-kind
58+
`max()` masking bug: a remediation path refreshing the signal meant to report
59+
it.
60+
- **A rebuild no longer leaves the column without an FTS index.** `rebuild_indexes`
61+
dropped every index and recreated it, on the assumption that LanceDB falls
62+
back to a brute-force scan meanwhile. That holds for vector search and **not**
63+
for FTS: with no inverted index a BM25 query raises `Cannot perform full text
64+
search unless an INVERTED index has been created`, and since the recall legs
65+
are gathered without `return_exceptions`, one failing leg 500s the whole
66+
search request. Now uses `create_index(replace=True)`, which swaps atomically
67+
— measured 0 failures across 49 queries spanning 3 replaces, versus 55
68+
failures for the same test against drop-then-create — and collapses the live
69+
index fragment set exactly as before (7 index files back to 4).
70+
- **The empty-index-dir sweep is bounded by lance's own threshold.** lance's
71+
`cleanup.rs` unlinks a superseded index's files but never its directory — it
72+
contains no `rmdir` at all, which is structural rather than an oversight: it
73+
targets object stores, where paths are flat keys and an empty directory does
74+
not exist. Only a local filesystem materialises them, and a soak run reached
75+
13061 dirs, 98% empty. everos sweeps them, now with three independent
76+
guarantees instead of a self-chosen age: `rmdir` cannot delete a non-empty
77+
directory (the kernel refuses it, so no file can be lost and there is no
78+
check-then-act window), live index UUIDs are excluded via `list_indices()`,
79+
and anything else must outlive `UNVERIFIED_THRESHOLD_DAYS = 7` — lance's own
80+
bound for deciding an unreferenced index UUID is dead rather than mid-build.
81+
The previous 300s was our invention, which is what made it indefensible.
82+
Two consequences worth knowing: the age gate reads the dir's mtime, which
83+
POSIX bumps when lance's cleanup empties it, so the effective reclaim horizon
84+
is up to ~14 days (file wait + age gate) and the ceiling-load steady state is
85+
~1.8M dirs / ~7GB; and a sweep that blows its 60s deadline is swallowed
86+
inside `prune()` — the cleanup commit already succeeded, so escaping would
87+
bill a prune "failure" (feeding the fallback-rebuild threshold) and stall the
88+
prune-staleness clock for a cleanup stall that did not happen.
89+
- **The optimize runner's wait on an in-flight rebuild is bounded too.** The two
90+
maintenance jobs park on each other — whichever arrives second waits — so an
91+
unbounded wait on this side is the same hazard as the one already fixed on
92+
the rebuild side, just seen from the other end: the kind's task slot stays
93+
occupied, `_schedule_optimize` keeps short-circuiting on it, and that table
94+
quietly stops being pruned. It was left open on the argument that
95+
`rebuild_indexes` carries its own 300s deadline, which covers its critical
96+
section but not the task's dispatch and teardown, so the transitive bound was
97+
never real. Now bounded at 180s, logging
98+
`cascade_lancedb_optimize_skipped_rebuild_unfinished` and skipping the beat
99+
rather than compacting under a live rebuild — the two commit on the same
100+
manifest, which is what the wait exists to prevent.
101+
- **A rebuild that loses a commit race is retried** instead of waiting out the
102+
full 12h cadence. Lance labels the conflict `Retryable` and it is: a
103+
concurrent writer in another process won the manifest, nothing is wrong with
104+
the table. Retries are scheduled on the kind (10min / 30min / 3h) rather than
105+
slept through, so the other kinds in the sweep are not parked behind the
106+
backoff. A soak run at a 600s cadence hit 3 conflicts in 119 attempts, all
107+
while a concurrent CLI storm was running.
108+
109+
- **The index-rebuild sweep can no longer park forever** waiting on the
110+
optimize runner. That wait had no deadline, and the runner's loop condition is
111+
"keep going while there is unindexed data" — which under sustained writes is
112+
never, since the drain loop re-raises the flag every second against a 10s
113+
cooldown. Now bounded at 180s, logging
114+
`cascade_lancedb_rebuild_skipped_optimize_unfinished` and skipping the sweep
115+
rather than dropping indices under a live optimize. **This makes the stall
116+
visible, not absent**: under sustained ingest every sweep still times out, so
117+
active index-UUID / FTS `part_N` growth stays unbounded there. The functional
118+
fix requires the optimize runner to yield when a rebuild is pending, which
119+
changes the optimize/rebuild mutual-exclusion contract and needs its own
120+
validation — the rebuild cadence is 12h, longer than any soak run so far, so
121+
the periodic sweep has never been exercised under load.
122+
- **The memory-root lock wait is bounded and visible.** Acquisition polls with
123+
`LOCK_NB` instead of blocking inside a worker thread: a blocking `flock` could
124+
not be bounded or cancelled — cancelling the awaiting coroutine left the
125+
thread to acquire the lock later with nobody to release it. The wait itself is
126+
by design (the second process is supposed to wait, then find the migration
127+
already done), but it now logs `memory_root_lock_waiting` and gives up after
128+
`timeout_seconds` (default 30min) instead of leaving a server startup looking
129+
like a hang whose last message is `lifespan_provider_startup name=lancedb`.
130+
The default sits an order of magnitude above the worst legitimate hold (a
131+
large migration is minutes) on purpose: the wait is already visible from the
132+
first poll, and against the one case the bound exists for — a holder that is
133+
alive but wedged — giving up at 5 minutes buys nothing over 30, while a bound
134+
near the legitimate hold turns a slow migration into startup crashes for
135+
every waiting process.
136+
10137
### Changed
11138

12139
- **`cascade_lancedb_optimize_conflict` now records `pruned`** — which

‎src/everos/config/__init__.py‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@
44
from everos.config import (
55
Settings, MemorySettings, SqliteSettings, LanceDBSettings,
66
LLMSettings, EmbeddingSettings, RerankSettings,
7-
BoundaryDetectionSettings,
7+
BoundaryDetectionSettings, CascadeSettings,
88
load_settings, resolve_root,
99
)
1010
@@ -13,6 +13,7 @@
1313
"""
1414

1515
from .settings import BoundaryDetectionSettings as BoundaryDetectionSettings
16+
from .settings import CascadeSettings as CascadeSettings
1617
from .settings import EmbeddingSettings as EmbeddingSettings
1718
from .settings import LanceDBSettings as LanceDBSettings
1819
from .settings import LLMSettings as LLMSettings

‎src/everos/config/default.toml‎

Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -146,6 +146,18 @@ threshold = 0.65
146146
time_window_days = 7.0
147147

148148

149+
[cascade]
150+
# Maintenance cadences (seconds). Deadlines that bound a hung call are not
151+
# here on purpose — they live next to the code they guard and are sized from
152+
# measurement, not preference.
153+
optimize_heartbeat_seconds = 60.0
154+
optimize_prune_interval_seconds = 300.0
155+
# Passed to LanceDB as cleanup_older_than. Only has to outlive an in-flight
156+
# read, but it also decides how long index files keep a manifest naming them:
157+
# under LanceDB's 7-day unverified window they then wait out the full 7 days.
158+
optimize_prune_retention_seconds = 60.0
159+
optimize_rebuild_interval_seconds = 43200.0
160+
149161
[observability]
150162
# OpenTelemetry tracing export. Off by default; pure OTLP/HTTP, vendor-neutral
151163
# (Langfuse, an OTel Collector, or any OTLP backend). EverOS ships no vendor SDK.

‎src/everos/config/settings.py‎

Lines changed: 38 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -338,6 +338,43 @@ class LanceDBSettings(BaseModel):
338338
index_cache_size_bytes: int = 16 * 1024 * 1024
339339

340340

341+
class CascadeSettings(BaseModel):
342+
"""Cascade maintenance cadences.
343+
344+
These are *how often* each background job runs, not how long it is allowed
345+
to take — the deadlines that bound a hung call stay as constants next to the
346+
code they guard, sized from measurement, because a wrong value there either
347+
masks a hang or manufactures failures.
348+
349+
``optimize_heartbeat_seconds``:
350+
Idle sweep that offers every kind to the optimizer, so an unindexed tail
351+
left by a crash is merged even without new writes.
352+
353+
``optimize_prune_interval_seconds``:
354+
How often the heavy beat runs: reclaim the files of superseded dataset
355+
versions. Raise it if the write-lock hold is disruptive, lower it if disk
356+
transients are.
357+
358+
``optimize_prune_retention_seconds``:
359+
Passed straight to LanceDB as ``cleanup_older_than`` — versions replaced
360+
longer ago than this become eligible for deletion. It only has to outlive
361+
an in-flight read (sub-second). Shorter shrinks the transient footprint of
362+
superseded data fragments, but note it also decides how long index files
363+
keep a manifest that names them: below LanceDB's 7-day unverified window,
364+
index files lose that reference and wait out the full 7 days.
365+
366+
``optimize_rebuild_interval_seconds``:
367+
Full index rebuild per kind, which collapses the active index fragment
368+
count that every ``optimize()`` grows. Bounded by rebuild cost, not
369+
correctness — a missed sweep only defers cleanup.
370+
"""
371+
372+
optimize_heartbeat_seconds: float = 60.0
373+
optimize_prune_interval_seconds: float = 300.0
374+
optimize_prune_retention_seconds: float = 60.0
375+
optimize_rebuild_interval_seconds: float = 12 * 60 * 60.0
376+
377+
341378
class KnowledgeSearchSettings(BaseModel):
342379
"""``[knowledge.search]`` — retrieval tuning for the knowledge module."""
343380

@@ -417,6 +454,7 @@ class Settings(BaseSettings):
417454
boundary_detection: BoundaryDetectionSettings = BoundaryDetectionSettings()
418455
memorize: MemorizeSettings = MemorizeSettings()
419456
clustering: ClusteringSettings = ClusteringSettings()
457+
cascade: CascadeSettings = CascadeSettings()
420458
multimodal: MultimodalSettings = MultimodalSettings()
421459
knowledge: KnowledgeSettings = KnowledgeSettings()
422460
observability: ObservabilitySettings = ObservabilitySettings()

‎src/everos/core/persistence/lancedb/base.py‎

Lines changed: 21 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -104,11 +104,28 @@ def to_arrow_schema(cls) -> pa.Schema:
104104
)
105105

106106
@classmethod
107-
async def ensure_fts_indexes(cls, table: AsyncTable) -> None:
107+
async def ensure_fts_indexes(
108+
cls, table: AsyncTable, *, replace: bool = False
109+
) -> None:
108110
"""Create FTS indexes on every column in :attr:`BM25_FIELDS`.
109111
110112
Idempotent: columns that already have an index are skipped, so
111-
this is safe to call on every startup. The FTS config is fixed
113+
this is safe to call on every startup.
114+
115+
``replace=True`` rebuilds each column's index in place instead of
116+
skipping it — used by :meth:`LanceRepoBase.rebuild_indexes`, which
117+
needs a fresh index but must never leave the column *without* one.
118+
Dropping first would do that, and a BM25 query in that window does not
119+
degrade — it raises ``Cannot perform full text search unless an
120+
INVERTED index has been created`` (measured; vector search does fall
121+
back to a flat scan, FTS does not). Since the recall legs are gathered
122+
without ``return_exceptions``, that window turns into a 500 on the
123+
whole search request. ``create_index(replace=True)`` is atomic: 49
124+
concurrent queries across 3 replaces saw 0 failures, and it collapses
125+
the live fragment set exactly as drop+create does (7 index files back
126+
to 4, measured).
127+
128+
The FTS config is fixed
112129
to the app-layer pre-tokenisation + LanceDB normalisation
113130
convention (designed for **multilingual mixed content**):
114131
@@ -145,10 +162,11 @@ async def ensure_fts_indexes(cls, table: AsyncTable) -> None:
145162
indices = await table.list_indices()
146163
indexed_cols = {col for idx in indices for col in (idx.columns or [])}
147164
for field in cls.BM25_FIELDS:
148-
if field in indexed_cols:
165+
if field in indexed_cols and not replace:
149166
continue
150167
await table.create_index(
151168
column=field,
169+
replace=replace,
152170
config=FTS(
153171
with_position=False,
154172
base_tokenizer="whitespace",

0 commit comments

Comments
 (0)