Repository navigation
fix(*): give bulk memory writes the extraction budget and halve the import batch - #563
gloryfromca wants to merge 3 commits into
Conversation
…mport batch Every real cold-start import source failed against a live EverOS: the importer's non-final appends ran under the per-turn store budget (10s plus 0.5s per message on main), while EverOS extracts on the add itself. Measured against a real service, a 15-message batch took 12s, a 52-message batch 24s, and a batch of 100 ran past the six-minute extraction budget. Of 18 memory-file sources on a real machine, 14 failed at the client while the server kept writing, so a retry would have duplicated what had landed. Two changes, one per layer: - The importer marks its appends with metadata["bulk"]; the EverOS backend gives a bulk write the extraction budget outright instead of the per-message estimate. Nothing waits on an import write, so the turn budget's reason does not apply to it. Chat turns carry no flag and keep the short budget; the hosted backends (Mem0, Zep, MemOS) do not read the incoming metadata, so the key is invisible to them. - _BATCH_MSG_LIMIT drops from 100 to 50, the top of the range where EverOS extraction cost is still linear in the message count. A pin test records the measurement so the ceiling is not raised back casually. Batch boundaries feed EverOS message ids, so content imported under the old limit is not deduplicated against a re-import; no install had completed an import at either limit. Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
The measurement behind the batch ceiling now lives in one place, the orchestrator constant and its pin test; the backend comment only says what the flag does. Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
gloryfromca
left a comment
There was a problem hiding this comment.
Blocking: update the import integration test to match the new 50-message batching contract.
I reviewed the full diff, the MemoryBackend metadata contract, importer and EverOS call paths, affected callers and integration coverage, both branch commits and surrounding history, backward compatibility, test changes, and the repository rules in AGENTS.md plus CONTEXT-MAP.md/CONTEXT.md. The production behavior is internally consistent: ordinary turn writes retain the short size-based timeout, import writes get the extraction timeout, final-flush semantics remain unchanged, and unknown metadata remains optional for other backends. I found one blocking test-suite regression inline.
Verification:
uv run --extra dev pytest tests/test_everos_backend.py -x: 171 passed.uv run --extra dev pytest tests/test_importer_orchestrator.py tests/integration/test_import_e2e.py: 28 passed, 1 failed (test_batching_large_conversation).git diff --check github/main...HEAD: passed.
No tests were weakened to conceal behavior, and I found no dependency, asset, source-language, test-naming, or architecture-boundary violation beyond the stale integration expectation called out inline.
…ssage contract The integration test still asserted two batches of 100 and 60 for a 160-message conversation; under the new ceiling the run makes four calls of 50, 50, 50 and 10. The total-message and final-batch coverage stays, and the test now also checks that every import write carries the bulk flag. Co-authored-by: Claude (claude-fable-5-1) <noreply@anthropic.com>
|
Addressed the one blocker in 7bde3b5 (integration batching expectation moved to the 50-message contract). Verification on the branch head: the two files the review reproduced with plus the four unit files that exercise the importer and the EverOS backend all pass locally; ruff, commitlint, check_commit_messages and check-source-language pass on github/main..HEAD. No production code changed in this commit. |
gloryfromca
left a comment
There was a problem hiding this comment.
No blockers; this can merge as far as I am concerned.
The new commit updates only tests/integration/test_import_e2e.py and directly fixes the prior failure without weakening coverage: it asserts the exact 50, 50, 50, 10 batch sequence, the final flag only on the last batch, the bulk flag on every batch, and the unchanged 160-message total. I rechecked the full diff and the earlier review areas: repository rules and context vocabulary, importer and EverOS callers, commit history, compatibility of optional metadata with other backends, architecture boundaries, and test integrity. No new issue emerged.
Verification: uv run pytest tests/test_importer_orchestrator.py tests/integration/test_import_e2e.py tests/test_everos_backend.py -q -> 200 passed in 9.22s; git diff --check github/main...HEAD passed. The blocker thread was replied to and resolved.
|
Not a blocker. Nothing here holds the PR; two of the three points are about 1.
|
|
No blockers; suggestions only, and they are marked inline. The cancellation-latency calculation holds: The 15- and 52-message samples show that The identity point is already recorded in the original fix commit body: it states that batch boundaries feed EverOS message IDs, re-import under the old limit will not deduplicate, and no install had completed an import at either limit. The source comment also preserves the boundary/identity coupling, so no additional merge condition follows from that point. Verification on the unchanged head: |
|
Closed in favour of #570 at the maintainer's request: the same three commits are folded onto the web import branch so the whole import work lands on refactor/ui_web_architecture in one PR, and nothing goes to main. |
Summary
Every real cold-start import source failed against a live EverOS, and the failure was on the client side: the importer's non-final appends ran under the per-turn store budget (10s plus 0.5s per message on main), while EverOS extracts on the add itself. Measured against a real service, a 15-message batch took 12s, a 52-message batch 24s, and a batch of 100 ran past the six-minute extraction budget. On a real machine 14 of 18 memory-file sources failed at the client while the server kept writing, so a retry would have duplicated what had already landed.
Two changes, one per layer:
metadata["bulk"], and the EverOS backend gives a bulk write the extraction budget (_MEMORIZE_TIMEOUT_S) outright instead of the per-message estimate. Nothing waits on an import write, so the turn budget's reason does not apply. Chat turns carry no flag and keep the short budget. The hosted backends (Mem0, Zep, MemOS) do not read the incomingmetadata, so the key is invisible to them._BATCH_MSG_LIMITdrops from 100 to 50, the top of the range where EverOS extraction cost is still linear in the message count. A pin test records the measurement so the ceiling is not raised back casually.Not in this PR: the web wizard's sync step still hardcodes the
fulltier and shows no progress; the CLI tier labels (minutes/hours) predate these measurements. Both are follow-ups once the numbers below are agreed.Type
Verification
Unit tests (the four files that exercise the importer and the EverOS backend):
Reverse check, with the two source files reverted and the tests kept:
TestBatching::test_msg_count_limit,TestBatching::test_a_batch_stays_inside_the_zone_everos_extracts_linearly,TestMetadata::test_metadata_marks_the_write_bulk_and_names_no_ownerandTestWriteBudgetFollowsTheCaller::test_a_bulk_write_gets_the_extraction_budget_however_smallfail; restored, 585 pass again.Lint and gates:
End to end, against a real EverOS 1.2.3 with real LLM and embedding credentials, in an isolated
RAVEN_HOMEwith an empty EverOS root on port 18893 (the normal install untouched). The source is a real Claude Code project memory directory, 36 files / 91 KB / 249 messages, driven throughraven.cli.import_commands._build_and_run, the same pathraven import runtakes:main (before): the first append (100 messages) timed out at the client after 10s;
storereturned False, the source was marked failed, and the server kept extracting for another three minutes.this branch (after): the same source completed,
submitted=1 failed=0, in 348.6s. Five bulk appends of 50 messages took 41s, 58s, 43s, 97s and 58s, the final flush 47s; the server logged 11 episodes, 9 atomic-fact batches and 5 user-profile updates for it. Oneepisode_extract_retry(EverOS failing to parse its own model's JSON) appeared and was absorbed by the budget; that retry is EverOS-side and not touched here.Relevant tests pass locally
Relevant lint / type checks pass locally
User-facing docs or screenshots are updated when needed (no user-facing copy changes; the CLI tier labels are a named follow-up)
Risk
User-visible change: cold-start imports that used to fail on every real source now complete; each import write may hold the EverOS client for up to the extraction budget (360s) instead of the per-message estimate, which only affects
raven importand the web wizard's sync step, never a chat turn. Batch boundaries feed EverOS message ids, so content imported under the old limit is not deduplicated against a re-import; no install had completed an import at either limit, so there is nothing to migrate.Rollback: revert the commit. The
bulkkey is additive and ignored by every other backend.Related Issues
N/A