Summary
bb's event cleanup (thread pruning) only runs when the entire workspace is idle: no thread is running or starting, no interaction is waiting on the user, and no project deletion or environment provisioning is in progress. On a workspace where some thread is almost always busy or waiting on a question, the cleanup rarely gets to run, so rows it would normally delete pile up in events. The expected behavior is that cleanup keeps making progress while you work and skips only the threads that are busy.
Versions and environment
- bb server
0.45.1-nightly.37211842106.1 (built from a2bbfaadaa), self-hosted on a 4-core / 8 GB Linux VM, reached through bb Connect.
- The database is 11.6 GB.
events takes 8.7 GB, plus about 1.5 GB of indexes on it, across about 4,600 threads.
- Code checked at
main df913a79dc.
Steps to reproduce
- Keep at least one thread running, or one thread waiting on a permission or question, for an extended period, which is typical with several agents.
- Let
thread-event-pruning run on its 10-second interval.
- At debug log level, each run logs
Thread pruning skipped while app work is active and advances nothing.
Expected vs actual
- Expected: cleanup keeps working through resolved deltas, old usage and rate-limit snapshots, and turn diffs while other threads are busy, skipping only those threads.
- Actual: any active thread or pending interaction anywhere stops the whole sweep (
reason: "busy").
Evidence
- The gate:
isDatabaseMaintenanceIdle requires zero active threads, zero pending interactions, zero provisioning, and zero project deletions.
- Where it stops the sweep:
thread-pruning-sweep.ts.
- Per-run budget: 50 ms and 64 advances (
THREAD_PRUNING_SWEEP_LIMITS). Space comes back at most 20,000 pages (about 78 MiB) per hourly incremental vacuum (DATABASE_INCREMENTAL_VACUUM_MAX_PAGES).
- Backlog on a live workspace: I ran a one-off script against that workspace's database that calls the same
advanceThreadPruning/getNextThreadPruningPolicy code, throttled (one advance every 250 ms, backing off on SQLITE_BUSY). In the first 5.5 minutes it removed 7,538 rows that cleanup had never reached. That's only about 3 MB. These policies don't touch item/started and item/completed payloads, which make up most of the size.
- Why the gate was added: it came with
1cfa17779e ("add retention maintenance sweeps"). The commit doesn't say it's needed for correctness. advanceThreadPruning already runs with busy_timeout = 0 and batches of 500, so it won't block app writes.
Proposed change
- Replace the global idle gate in the pruning sweep with a per-thread exclusion. Skip events for threads that are active, starting, or have a pending interaction, and keep advancing everything else. Turn completion and archiving already do per-thread live pruning (
LIVE_BATCH_SIZE), so this is consistent with what's there.
- Keep the global gate for full
VACUUM, which really does block, but consider raising DATABASE_INCREMENTAL_VACUUM_MAX_PAGES so freed pages come back faster than 78 MiB per hour.
- Add a test showing the sweep advances while an unrelated thread is active, and doesn't touch the active thread's events.
What I ruled out
Suggested priority and effort
Medium priority, small effort. It affects any workspace with frequently busy threads. Nothing is lost; the database just grows and search competes for page cache on small hosts. The workaround is to leave the workspace fully idle for long stretches.
BB-Thread: Quick switcher in palette
AGENT GENERATED
Summary
bb's event cleanup (thread pruning) only runs when the entire workspace is idle: no thread is running or starting, no interaction is waiting on the user, and no project deletion or environment provisioning is in progress. On a workspace where some thread is almost always busy or waiting on a question, the cleanup rarely gets to run, so rows it would normally delete pile up in
events. The expected behavior is that cleanup keeps making progress while you work and skips only the threads that are busy.Versions and environment
0.45.1-nightly.37211842106.1(built froma2bbfaadaa), self-hosted on a 4-core / 8 GB Linux VM, reached through bb Connect.eventstakes 8.7 GB, plus about 1.5 GB of indexes on it, across about 4,600 threads.maindf913a79dc.Steps to reproduce
thread-event-pruningrun on its 10-second interval.Thread pruning skipped while app work is activeand advances nothing.Expected vs actual
reason: "busy").Evidence
isDatabaseMaintenanceIdlerequires zero active threads, zero pending interactions, zero provisioning, and zero project deletions.thread-pruning-sweep.ts.THREAD_PRUNING_SWEEP_LIMITS). Space comes back at most 20,000 pages (about 78 MiB) per hourly incremental vacuum (DATABASE_INCREMENTAL_VACUUM_MAX_PAGES).advanceThreadPruning/getNextThreadPruningPolicycode, throttled (one advance every 250 ms, backing off onSQLITE_BUSY). In the first 5.5 minutes it removed 7,538 rows that cleanup had never reached. That's only about 3 MB. These policies don't touchitem/startedanditem/completedpayloads, which make up most of the size.1cfa17779e("add retention maintenance sweeps"). The commit doesn't say it's needed for correctness.advanceThreadPruningalready runs withbusy_timeout = 0and batches of 500, so it won't block app writes.Proposed change
LIVE_BATCH_SIZE), so this is consistent with what's there.VACUUM, which really does block, but consider raisingDATABASE_INCREMENTAL_VACUUM_MAX_PAGESso freed pages come back faster than 78 MiB per hour.What I ruled out
item/*payloads.auto_vacuumis2(incremental), andfreelist_countwas 4, so incremental vacuum is reclaiming space. The backlog is rows that were never deleted.Suggested priority and effort
Medium priority, small effort. It affects any workspace with frequently busy threads. Nothing is lost; the database just grows and search competes for page cache on small hosts. The workaround is to leave the workspace fully idle for long stretches.
BB-Thread: Quick switcher in palette