Split out of aios#2347 (remedy 5) now that the uv-cache leak is closed by #2348. This is the remaining structural disk defect on server-b.
Observation (server-b, 2026-09-03 22:50Z)
| session |
docker image inspect .Size (view) |
Σ docker history layer sizes (on disk) |
layers |
seat sess_01KVBPGT… |
6.6 GB |
20.1 GB (docker images "virtual" 26.8) |
16 |
davenant sess_01KZVHVS… |
9.0 GB |
25.5 GB |
53 |
ultron sess_01KQ8NQ3… |
7.6 GB |
21.5 GB |
96 |
The seat's writable view from inside the container was 1.5 GB (du -x /). Overlay keeps every superseded byte in the interior layers; a commit-per-idle-exit chain grows monotonically and only a flatten reclaims it.
Why nothing flattens them (read from the code, not inferred)
src/aios/sandbox/backends/docker.py:700-712:
base_size = await self._image_size_or_zero(base_ref)
parent_fields = await self._inspect_image_fields(parent_image) # (.Id, .Size, len(.RootFS.Layers), ...)
parent_size = parent_fields[1] if parent_fields else 0
projected_unique = max(0, parent_size - base_size) + rw
over_budget = flatten_if_unique_bytes_over is not None and projected_unique > flatten_if_unique_bytes_over
flattening = parent_depth + 1 >= _FLATTEN_DEPTH_CEILING or over_budget # ceiling = 200
.Size is the current filesystem view. The worker runs with AIOS_SANDBOX_SNAPSHOT_BUDGET_BYTES=12884901888 (12 GiB), so 6.6/9.0/7.6 GB never trip it; depth 16/53/96 never reaches 200. Every idle exit is committed (0 flattened events in 24 h of worker logs).
The pool reclaimer has the identical blind spot — registry.py:3537-3550 _unique_bytes_for_image = tag.Size − base.Size — so the 60 GB pool budget reads 28.6 GB used while tagged sandbox images occupy ~127 GB virtual / ~38 GB unique layer bytes; snapshot_pool_reclaim logs reclaimed_bytes: 0 every tick in enforce mode. Both controls are correct by their own metric and never fire. (Falsifier for the whole claim: a kind=flattened event for any of these three sessions in the worker log. There is none.)
Remedy
- Measure chain cost, not view: Σ layer sizes from
docker history --no-trunc --format '{{.Size}}' (or the image manifest / docker image inspect --format '{{json .RootFS.Layers}}' + docker system df -v), cached per image id.
- Flatten when
chain_bytes > budget or chain_bytes > K × view_bytes (K≈2 — the point where more than half the chain is dead history).
- Use the same chain-cost figure in
_unique_bytes_for_image so the pool budget is enforceable.
- Keep the existing flatten headroom gate (
estimate × 1.75 + flatten_disk_floor); with ~27 GB free now the seat's chain (est ~1.5 GB) can flatten immediately, the 7–9 GB chains after it.
Expected recovery on this host: ~40 GB. Tests: a fake image graph where view=6 GB, chain=20 GB, budget=12 GiB must choose flatten; view=6, chain=7 must commit; the pool pass with the same graph must report chain bytes, and the LRU must pick the archived candidate first.
Interim lever (operator, not code)
Lowering AIOS_SANDBOX_SNAPSHOT_BUDGET_BYTES to the 4 GiB default trips the view-based trigger for all three (6.6 > 4). Not pulled by the seat: it is the chairman's setting and the flatten order matters under the headroom gate. Available if disk pressure returns before this lands.
Split out of aios#2347 (remedy 5) now that the uv-cache leak is closed by #2348. This is the remaining structural disk defect on server-b.
Observation (server-b, 2026-09-03 22:50Z)
docker image inspect .Size(view)docker historylayer sizes (on disk)sess_01KVBPGT…docker images"virtual" 26.8)sess_01KZVHVS…sess_01KQ8NQ3…The seat's writable view from inside the container was 1.5 GB (
du -x /). Overlay keeps every superseded byte in the interior layers; a commit-per-idle-exit chain grows monotonically and only a flatten reclaims it.Why nothing flattens them (read from the code, not inferred)
src/aios/sandbox/backends/docker.py:700-712:.Sizeis the current filesystem view. The worker runs withAIOS_SANDBOX_SNAPSHOT_BUDGET_BYTES=12884901888(12 GiB), so 6.6/9.0/7.6 GB never trip it; depth 16/53/96 never reaches 200. Every idle exit iscommitted(0flattenedevents in 24 h of worker logs).The pool reclaimer has the identical blind spot —
registry.py:3537-3550 _unique_bytes_for_image = tag.Size − base.Size— so the 60 GB pool budget reads 28.6 GB used while tagged sandbox images occupy ~127 GB virtual / ~38 GB unique layer bytes;snapshot_pool_reclaimlogsreclaimed_bytes: 0every tick in enforce mode. Both controls are correct by their own metric and never fire. (Falsifier for the whole claim: akind=flattenedevent for any of these three sessions in the worker log. There is none.)Remedy
docker history --no-trunc --format '{{.Size}}'(or the image manifest /docker image inspect --format '{{json .RootFS.Layers}}'+docker system df -v), cached per image id.chain_bytes > budgetorchain_bytes > K × view_bytes(K≈2 — the point where more than half the chain is dead history)._unique_bytes_for_imageso the pool budget is enforceable.estimate × 1.75 + flatten_disk_floor); with ~27 GB free now the seat's chain (est ~1.5 GB) can flatten immediately, the 7–9 GB chains after it.Expected recovery on this host: ~40 GB. Tests: a fake image graph where view=6 GB, chain=20 GB, budget=12 GiB must choose
flatten; view=6, chain=7 mustcommit; the pool pass with the same graph must report chain bytes, and the LRU must pick the archived candidate first.Interim lever (operator, not code)
Lowering
AIOS_SANDBOX_SNAPSHOT_BUDGET_BYTESto the 4 GiB default trips the view-based trigger for all three (6.6 > 4). Not pulled by the seat: it is the chairman's setting and the flatten order matters under the headroom gate. Available if disk pressure returns before this lands.