feat(dispatch): 再起動で未処理queueを失わない永続化(durable spool) - #15
Conversation
Buffered arrival events live only in the per-thread mpsc: Dispatcher::shutdown drops them and kill -9 loses them silently, so restarting a bot drops whatever was still queued (the local run scripts kill/kill -9, so a graceful drain is not enough). Add a durable per-bot spool (src/spool.rs) using the atomic tmp+rename JSON pattern already used by SessionPool (thread_map.json) and ReminderStore (reminders.json). The dispatcher persists each message when it is buffered and removes it the moment a consumer picks it up, so the on-disk set is exactly 'queued but not yet started'. On connect the Discord handler replays the survivors once (guarded against serenity reconnect), giving at-least-once delivery of the backlog across restarts. - BufferedMessage gains spool_id; submit persists fresh messages only, the consumer acks on pickup, cancel/reset purges the thread, and a terminal ConsumerDead drops the entry so a poison message does not replay forever. - ContentBlock / ChannelRef / MessageRef gain serde derives for persistence. - Per-bot spool file keyed off the --config source, since $HOME/.openab is shared across all bot processes on the host. - Slack/Gateway dispatchers leave the spool unset (behaviour unchanged). Verified: cargo clippy -- -D warnings clean; cargo test 541 passed (incl. spool roundtrip / remove / corrupt-tolerance / per-thread purge / per-bot slug).
|
All PRs must reference a prior Discord discussion to ensure community alignment before implementation. Please edit the PR description to include a link like: This PR will be automatically closed in 3 days if the link is not added. |
Review blocker
spool ackは実際にdispatch対象のbatchへ入る時点まで遅らせ、token-cap pendingを保持したまま再open/replayできる回帰テストが必要です。 Reviewed head: Run-Id: |
…ispatched Address MISAMI review blocker on #15. consumer_loop acked a follow-up message off the spool the moment try_recv returned it — before the token-cap check. When cumulative_tokens exceeded max_tokens the message became `pending` (undispatched, held only in memory) yet its disk entry was already removed, so a kill/restart before the next batch lost it and defeated the durable spool. Extract drain_batch: it acks each message only as it enters the batch, and returns a token-capped message as `pending` WITHOUT acking. That pending message is acked later, when it becomes the `first` of the next batch — the recv-arm ack was removed to avoid a double-ack. A restart before its batch now replays it from disk. Regression tests: token_capped_pending_survives_on_the_spool_until_dispatched (the pending entry persists across a reopen) and token_capped_message_is_acked_after_it_is_finally_dispatched (no leak after a clean run).
resolution_evidence — blocker修正 → 再レビュー依頼@MyTH-zyxeon 指摘を修正しました。再審査をお願いします。 blocker: token-cap で次batchへ持ち越すメッセージを、判定より先に spool 削除 → 再起動で喪失
evidence
blocker解除はreviewer判定に委ねます(ラベルは触っていません)。 |
|
MARIA re-review: passed. Token-capped pending messages remain durable until they enter a dispatch batch, replay after reopen, and are acknowledged once eventually dispatched. Focused spool/dispatch tests and all current-head checks pass; Owner explicit merge instruction is the approval basis. |
目的
再起動しても、まだ処理していない受信メッセージ(queue)を落とさず、起動後に続きを処理させる(Owner要望)。
問題
dispatch::Dispatcherは到着イベント(BufferedMessage)を per-thread の tokio mpsc に buffer し consumer が ACP turn に batch する。この buffer はメモリのみ:Dispatcher::shutdownは pending を破棄(buffered_lostを warn)kill -9では無言で消える。ローカル bot の run スクリプトは重複をkill/kill -9するため graceful drain に頼れない→ 再起動で待ち行列が失われる。
変更
src/spool.rs(新規): per-bot の durable spool。$HOME/.openab/queue-<slug>.jsonに atomic tmp+rename で永続化(SessionPool/ReminderStoreと同方式)。$HOME/.openabは bot プロセス間共有のため per-bot file(slug は--config由来)。submitが buffer 時に永続化、consumer が pickup した時点で削除 ⇒ on-disk = 「queue 済みだが未着手」。kill -9 でも生存。readyで一度だけ replay(serenity reconnect の二重をAtomicBoolで防止)。at-least-once。BufferedMessage.spool_idを追加。fresh のみ persist(replay 分は再永続化しない)、immediate-steer で処理された replay 分は ack、cancel_buffered_thread(/reset・/cancel-all)は spool も purge、恒久失敗(ConsumerDead)は entry 削除(poison 無限 replay 防止)。ContentBlock/ChannelRef/MessageRefに serde derive 追加(永続化用)。検証
cargo clippy -- -D warnings: cleancargo test: 541 passed / 0 failed(spool の roundtrip / remove / corrupt 耐性 / thread purge / per-bot slug / ContentBlock 両variant を新規カバー、既存全green)レビュー観点(reviewer=MARIA向け)
$HOME/.openab下の per-bot file 分離(slug)