What happened
Observed: During the 11 September 2026 incident, a 10% cleanup of a 193 GB RocksDB ledger was followed within roughly 15 seconds by all JSON-RPC methods on port 8899 becoming unresponsive, including getVersion and sendTransaction. Slots continued advancing and Prometheus remained up. Truncation never logged its compaction step; shutdown hung joining RPC and systemd eventually killed the process. After restart, the same truncation completed in 8 seconds.
Likely mechanism: A long history scan held lowest_cleanup_slot.read() while truncation queued for the write lock. Later synchronous ledger reads blocked behind that writer and exhausted the RPC runtime workers. This matches the reported 30 August devnet-asia incidents, but the exact lock convoy has not been confirmed by captured thread stacks.
Expected: Ledger cleanup and slow history reads must not starve unrelated RPC methods.
Completion criteria:
- Replace the cleanup
RwLock with the existing monotonic atomic retention boundary shared with compaction.
- Bound concurrent ledger-backed RPC operations to
(num_cpus::get() / 4).max(1), await capacity asynchronously, and keep synchronous RocksDB work from pinning RPC runtime workers.
- Validate retention at logical read boundaries rather than repeatedly within each row; reject cleanup-invalidated compound results and preserve the latest persisted slot.
- Keep non-ledger RPC methods outside the ledger concurrency gate.
Commit / version
Deployed master; 0.15.2
What happened
Observed: During the 11 September 2026 incident, a 10% cleanup of a 193 GB RocksDB ledger was followed within roughly 15 seconds by all JSON-RPC methods on port 8899 becoming unresponsive, including
getVersionandsendTransaction. Slots continued advancing and Prometheus remained up. Truncation never logged its compaction step; shutdown hung joining RPC and systemd eventually killed the process. After restart, the same truncation completed in 8 seconds.Likely mechanism: A long history scan held
lowest_cleanup_slot.read()while truncation queued for the write lock. Later synchronous ledger reads blocked behind that writer and exhausted the RPC runtime workers. This matches the reported 30 August devnet-asia incidents, but the exact lock convoy has not been confirmed by captured thread stacks.Expected: Ledger cleanup and slow history reads must not starve unrelated RPC methods.
Completion criteria:
RwLockwith the existing monotonic atomic retention boundary shared with compaction.(num_cpus::get() / 4).max(1), await capacity asynchronously, and keep synchronous RocksDB work from pinning RPC runtime workers.Commit / version
Deployed
master; 0.15.2