Skip to content

Add opt-in Storage & retention with durable cleanup - #4476

Open
ymichael wants to merge 65 commits into
mainfrom
bb/create-customizable-settings-page-thr_yvi4iy98d6
Open

ymichael wants to merge 65 commits into
mainfrom
bb/create-customizable-settings-page-thr_yvi4iy98d6

Conversation

@ymichael

@ymichael ymichael commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Human comments

What was wrong

BB had no unified way to see storage use across machines, reclaim thread files or abandoned development instances, or opt into automatic retention. Cleanup tied only to live notifications could also lose work while a plugin was disabled, reloading, or its machine was offline.

What changed

  • Add the experimental, default-disabled Storage & retention plugin with a machine overview, cached disk measurements, thread/worktree/development breakdowns, explicit scans, and matching bb storage commands. Reports include hidden threads and link development data to its source checkout where possible.
  • Offer file clearing, archived-file and large-file cleanup, orphan removal, failed worktree cleanup retry, and development-instance removal. Bulk work runs in tracked background jobs with progress and retryable errors. Different thread clears can run concurrently; conflicting maintenance is serialized per machine. Host workers validate paths and recheck eligibility before deleting files.
  • Add hourly thread archiving/deletion, both defaulting to Never. Pinned members exempt their lifecycle group. Two separate switches default off: delete thread storage after future archives, and delete development data whose checkout is missing. Settings save immediately and quietly; failed saves restore the previous value. Destructive bulk actions keep inline confirmations, and running development instances require stop-and-remove confirmation.
  • Archive storage cleanup persists pending work, waits for the shared 30-second undo grace, and retries once the thread stops and its machine reconnects. Unarchiving cancels it. File clearing preserves conversations and uploaded attachment originals. Development cleanup checks filesystem state on plugin startup, machine reconnect, SDK reconnect, and hourly when enabled; it can clean existing missing-checkout data. A live removal notification re-measures only that machine's ~/.bb-dev, queued behind any scan or cleanup already running there; later scans retry offline/busy machines or failed cleanup. Disabling the plugin stops its schedules and cancels active workers while retaining saved state.
  • Announce successful provider environment removal after its lifecycle transition commits. Notifications are ephemeral; startup, reconnect, and scheduled filesystem scans recover missed notifications. Provider removal success does not guarantee the path was deleted. No core database tables or migration are added.
  • Keep policy, report assembly, scan caching, and cleanup scheduling in the plugin/server. Host workers own filesystem measurement and deletion. Save development checkout metadata at launch. Canonicalize deleted checkout paths correctly and sweep processes for a whole batch of directories. Protocol version is 230 and Plugin SDK version is 0.6.27. Update CLI guides, plugin skills, configuration docs, and the Plugin Guide.

Public Plugin SDK changes

  1. New event: bb.events.on("experimental_environment.removed", handler) receives { removal } after successful provider removal commits. The payload contains environmentId, removedAt, hostId, path, and providerOwnedPath; delivery is ephemeral. PluginThreadEventPayloads and the fake-host harness support it.
  2. Changed host helper: experimental_killProcessesWithCwdUnder({ directories, graceMs? }) replaces the singular directory input with a directory list. Existing callers must pass directories: [directory]. Each sweep enumerates processes for the whole batch, including deleted paths resolved through their nearest existing ancestor. It uses SIGTERM then SIGKILL; Windows cwd enumeration remains a no-op.
  3. New host helper: experimental_readProcessIdentity(pid) from @get-bb/plugin-sdk/host returns { command, startedAt } or null, the same probe bb's launcher uses before stopping a recorded PID. Storage & retention uses it to confirm a development instance's recorded launcher (entry path and start time) before signalling it.

Plugin-owned RPC and data contracts

  • Inspection: state, hosts, host, scanHost, scanAll, and preview.
  • Configuration: configure accepts archive/delete thresholds plus deleteStorageOnArchive and deleteDevDataOnCheckoutRemoval.
  • Cleanup: clearThread, clearArchivedFiles, clearLargeFiles, removeOrphans, retryWorktreeCleanup, and removeDevInstances. The latter accepts all missing-checkout entries or selected names with explicit stopRunning.
  • Background jobs: startClearLargeFiles, startClearArchivedFiles, and startCleanup (orphans, development, or worktrees). Machine responses expose cleanup progress/completion/failure.
  • The plugin owns SQLite scan snapshots and pending archive cleanup; policy and the latest retention run use plugin KV storage. Development cleanup uses filesystem reconciliation and needs no removal-history tables, cursor, or persistent scan queue.

How you verified

  • Turbo lint/typecheck pass for storage, server, CLI, SDK, database, Plugin SDK, and Plugin Guide.
  • Tests pass: storage 69, focused server lifecycle/guide tests 164, database 588, SDK 109, CLI environment commands 35, Plugin SDK 365, and Plugin Guide 75.
  • Development-cleanup regression exercises real host-worker filesystem cleanup after reload without removal history, machine and SDK reconnect, live removal notification, retry after failure, and disabled/offline behavior. Existing and unidentified source checkouts remain intact. Temporarily removing startup reconciliation makes the test fail because orphaned data remains; restoring it passes.
  • Review hardening, each with a regression test that fails without the fix: the retention sweep re-selects every group right before archiving or deleting it (a thread pinned or unarchived mid-sweep is kept); stop-and-remove verifies the recorded launcher before signalling a PID; pending archive cleanups are dropped once their thread or machine returns 404; environment removal no longer rescans every machine.
  • After rebasing onto main, turbo run lint typecheck test --filter='...[origin/main]' passes except builtin-plugins.test.ts > surfaces builtin app build failures…, a file-watch rebuild test unrelated to this plugin that times out only under full parallel load and passes 3/3 in isolation.
  • Drizzle generation reports no schema changes and nothing to migrate. git diff --check passes.
  • Earlier branch verification covered process sweeps, development path recovery, desktop/compact layouts, disposable cleanup, and app tests. No new live-browser or simulator pass is claimed for this backend simplification.
  • CI fixes on this branch also align marketplace defaults and make Codex fixtures/lifecycle assertions tolerate Windows process-ID reuse.

Known limits: automation targets have no special retention exemption; pin threads that must be kept. Reports are cached snapshots, and filesystem eligibility is rechecked before cleanup. Windows does not enumerate working directories for process sweeps.

AGENT GENERATED

@ymichael
ymichael force-pushed the bb/create-customizable-settings-page-thr_yvi4iy98d6 branch 2 times, most recently from 16986b0 to 83b14e1 Compare October 2, 2026 00:09
@patleeman

Copy link
Copy Markdown
Contributor

Since this adds automatic thread deletion, I want to raise a need it touches.

What I'm trying to achieve: delete old threads to keep bb.db in check (by hand today, and with retention policies like these once this lands) without permanently losing their transcripts. When a thread is deleted, I'd like its full history to still exist somewhere on disk, in a plain file format like JSONL, so I can grep it, read it, or bring it back into context later.

Why it matters for me:

  • My bb.db is 2.1 GB. The events table is about 1.1 GB of data plus about 400 MB of indexes. 83% of the event data (910 MB) belongs to 813 archived threads, 519 MB of it in threads archived more than 30 days ago.
  • Deleting old archived threads is the only meaningful way to get that space back. Since deletion is permanent, I've been holding off.
  • An automatic delete policy makes this sharper: threads would disappear without anyone looking at them first.

Why I can't cover it with a plugin today: thread.deleted fires after the soft delete, and the server doesn't wait for handlers. For a thread with no environment, the hard delete (and the cascade through events) runs in the same request, so the events are already gone by the time a handler could read them. For threads with an environment, the window lasts only until the daemon confirms storage deletion. The only reliable option left is to mirror every thread's events to disk continuously, just in case.

Related: #3598 (undelete).

AGENT GENERATED

@ymichael

ymichael commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator Author

Since this adds automatic thread deletion, I want to raise a need it touches.

What I'm trying to achieve: delete old threads to keep bb.db in check (by hand today, and with retention policies like these once this lands) without permanently losing their transcripts. When a thread is deleted, I'd like its full history to still exist somewhere on disk, in a plain file format like JSONL, so I can grep it, read it, or bring it back into context later.

Why it matters for me:

  • My bb.db is 2.1 GB. The events table is about 1.1 GB of data plus about 400 MB of indexes. 83% of the event data (910 MB) belongs to 813 archived threads, 519 MB of it in threads archived more than 30 days ago.
  • Deleting old archived threads is the only meaningful way to get that space back. Since deletion is permanent, I've been holding off.
  • An automatic delete policy makes this sharper: threads would disappear without anyone looking at them first.

Why I can't cover it with a plugin today: thread.deleted fires after the soft delete, and the server doesn't wait for handlers. For a thread with no environment, the hard delete (and the cascade through events) runs in the same request, so the events are already gone by the time a handler could read them. For threads with an environment, the window lasts only until the daemon confirms storage deletion. The only reliable option left is to mirror every thread's events to disk continuously, just in case.

Related: #3598 (undelete).

AGENT GENERATED

@patleeman tracking your feedback here: #4714 - I'm curious what you'd find useful in the snapshot / feel free to comment on the issue to clarify anything I missed

@ymichael
ymichael force-pushed the bb/create-customizable-settings-page-thr_yvi4iy98d6 branch 2 times, most recently from 2922da2 to b9e4b93 Compare October 5, 2026 20:24
@ymichael
ymichael marked this pull request as ready for review October 5, 2026 21:12
@ymichael
ymichael force-pushed the bb/create-customizable-settings-page-thr_yvi4iy98d6 branch from 9806ea0 to 859ddc9 Compare October 6, 2026 17:57
@ymichael ymichael changed the title Add opt-in Storage & retention plugin Add opt-in Storage & retention with durable cleanup Oct 6, 2026
@ymichael
ymichael force-pushed the bb/create-customizable-settings-page-thr_yvi4iy98d6 branch from f690eef to 7088e28 Compare October 6, 2026 21:33
ymichael and others added 18 commits October 6, 2026 19:39
Drop the duplicated page header, show machine online status and a usage
summary per machine, replace day inputs with preset selects, and surface
pinned-thread protection as a banner. Scans now record the free and total
bytes of the thread-storage volume.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Largest threads render as single-line rows with archived and running
pills. Machine rows follow Settings → Machines: server pill, status,
and per-machine thread counts that need no scan. Drop the retention
footer when there is nothing to report.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Thread counts come from the scan, the only source that can attribute
storage to a machine: most archived threads no longer reference an
environment, so environment-based counts undercount them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The panel owns the data and machine pages read their report from the
host list, so moving between the list and a machine renders instantly.
The first load shows nothing for 200 ms, then a page-shaped skeleton.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Add Scan all and per-machine scan buttons to the machine list, a banner
that clears storage of archived threads across scanned online machines,
and an archived row in each machine's cleanup list. Pinned and running
threads are skipped. Host discard now takes batches of entries. The CLI
gains clear-archived and usage --rescan without --machine. Use the new
DatabaseRestore icon for the plugin.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Clear-archived now targets archived threads holding 100 MB or more, which
on real machines covers about 99% of archived bytes while touching a few
dozen threads instead of thousands. Reports carry the exact clearable
count and size, and the banner moves to the top of the page with copy
that says only thread-storage files go and conversations stay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scans now find files of 10 MB or more in thread storage. When archived,
unpinned threads hold 1 GB or more of them, the Storage page suggests
deleting just those files, keeping smaller files like reports and every
conversation. Replaces whole-folder clearing of archived threads; the
CLI command becomes clear-large-files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Both passes stat every file in thread storage and are kernel-bound, so
running them side by side cuts a 5.5M-file scan from about 3.5 to about
2 minutes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The scan now runs du -a once and streams its output, recording files
at or above the threshold while summing directory sizes, instead of a
separate find walk over the same files. The walker fallback records
them too, and deleting large files reuses the same measurement. A scan
of 5.5M files costs one du pass again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Bundled plugins added after #4576 must use a bb-- id. Rename the package
to bb-plugin-bb--storage-retention so the plugin id is
bb--storage-retention; the directory and builtin:storage-retention
source stay. The plugin has not shipped, so no migration is needed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ymichael and others added 28 commits October 6, 2026 19:39
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ervers

Dev servers keep running after their checkout is deleted (for example when a
thread's storage is cleared), so removal first stops every process whose
command line runs from the missing checkout. The host re-verifies each checkout
is missing before deleting and skips entries whose processes survive.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Clearing a thread's files renamed the folder to trash and deleted it while dev
servers started inside it kept running from the trash path. Discard now stops
processes whose working directory is inside each folder first, using the same
SDK helper as worktree removal.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sweep

Removing a development instance now uses experimental_killProcessesWithCwdUnder
for each missing checkout instead of a plugin-owned ps/PowerShell command-line
matcher and kill loop. The core sweep now resolves the nearest existing ancestor
of a deleted directory, so a removed checkout under a symlinked path such as
/tmp still matches the canonical working directory the OS reports.

The Windows dev-instance test timed out listing processes through PowerShell;
the cwd sweep is a no-op on Windows, so the test is skipped there like the
thread-storage process test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The largest-threads toggle sat outside its card, and the development storage
toggle expanded every section at once from below the last one. Each list now
previews five rows and ends with its own Show N more / Show fewer control inside
the card, so each development section expands independently.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rows whose source checkout is missing offer Remove instance… in their menu,
confirmed inline under the row. removeMissingDevInstances takes names (null for
every missing instance) and rejects names the last scan did not mark missing;
the host still re-verifies each checkout before stopping servers and removing.
bb storage remove-dev-instances gains --instance NAME. Rows whose checkout
exists or whose source is unidentified get no remove action.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Pills sat in a flex column beside the title, so a wrapped title left them
floating at the right edge. Thread titles in the largest-threads list and the
development card now share one ThreadTitle that renders pills inline after the
title text, and the development card uses the same archived pill instead of
plain text.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Clean up rows share one CleanupRow and a SectionHeading, and the bulk
  development-instance removal moves there from a banner inside the dev card.
- Largest-thread rows match development rows: title with inline pills, size
  and a three-dot menu (Open thread, Clear files) on the right at every width.
  Clearing still starts immediately; progress and errors show under the title.
- Linked threads of one development entry stack one per line, group counts
  drop the double gap, and pills or size badges that wrap no longer indent.
- Scan times read "Scanned 3d ago" everywhere with the exact time on hover,
  and machine rows say "26.4 GB in 222 dev instances" like the thread total.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Some ~/.bb-dev entries are symlinks into thread storage or temp folders. Removal
rejected any non-directory entry, so one symlink aborted the whole batch with
"Development instance must be a directory, not a symbolic link". A symlinked
entry is now unlinked, keeping the folder it points to, and other non-directory
entries are skipped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each killProcessesWithCwdUnder call lists every process working directory
(about 120 ms with lsof on macOS), and storage cleanup called it once per entry:
roughly 16 s for 133 development instances and minutes for thousands of
archived thread folders, so the remove buttons appeared to hang.

The helper now takes directories and matches processes under any of them from
one listing. Storage & retention passes each whole batch; the daemon,
worktree and personal-workspace callers pass their single directory. 133
directories now take one ~120 ms sweep.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ances

Thread rows and missing-checkout development rows now have an inline trash
button that acts immediately, with a spinner on the button and progress or
error text under the title. The ⋯ menu on development rows keeps navigation
and path copying; thread rows drop their menu since the title opens the thread.

Single-instance removal takes a per-row lock like thread clearing, so different
rows can run at once while the bulk Clean up actions stay exclusive, and each
removal re-reads the cached scan before updating it so concurrent removals
don't overwrite each other.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The right-hand cluster's alignment placeholder was taller than the title line,
pushing sizes below their titles and spacing rows apart; both row clusters now
sit on the title's line height. An instance linked to several threads shows its
first thread with an "Also linked to N other threads" line instead of stacking
titles that read like separate rows. Group headings carry their own short
description, replacing the footer that only explained "Other instances".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Only instances whose checkout was gone could be removed, so once those were
cleared no row had an action. Each row now has the trash button.

The host reports an instance as running when its daemon lock
(daemon.lock.lock, refreshed every few seconds by proper-lockfile) changed in
the last 15 seconds. Removing a single instance (removeDevInstances with names,
or --instance) stops servers and removes it when its checkout is gone, removes
it when it is not running, and otherwise refuses with "Stop the development
server first"; BB never stops servers whose checkout still exists. The bulk
Clean up action still removes only missing-checkout instances. The RPC is
renamed from removeMissingDevInstances since it is no longer limited to them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Removing one development instance never refuses. When its dev server is
running (fresh daemon lock) and its checkout still exists, a plain removal
leaves it in place and reports it in `running`; the row then asks "Stop and
remove?" and confirming sends SIGTERM to the launcher PID recorded in that
instance's bb-app-runtime.json (SIGKILL after 15 s) before removing it. The CLI
treats --yes as that confirmation. Only the instance's own launcher is stopped,
never other processes in the checkout.

removeDevInstances takes {hostId, names: null} for the bulk missing-checkout
cleanup or {hostId, names, stopRunning} for chosen instances, and the host
contract carries the matching mode instead of a refusal condition.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Re-select each retention group right before archiving or deleting it, so a
  thread pinned or unarchived during a sweep is kept.
- Confirm a development instance's recorded launcher (entry path and start
  time) before signalling its PID; expose experimental_readProcessIdentity
  from the Plugin SDK host entry for it.
- Drop pending archive cleanups whose thread or machine no longer exists.
- Clean development data on environment removal by re-measuring only that
  machine's ~/.bb-dev, queued behind running maintenance instead of dropped.
- Replace a toast-placement UI test with the stop-and-remove request flow.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ymichael
ymichael force-pushed the bb/create-customizable-settings-page-thr_yvi4iy98d6 branch from 86e8023 to 6b2bdf3 Compare October 7, 2026 02:43
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants