Skip to content

perf(compaction): compact altimate-base earlier to cut prefill cost - #1403

Closed
anandgupta42 wants to merge 1 commit into
mainfrom
perf/altimate-base-compaction
Closed

anandgupta42 wants to merge 1 commit into
mainfrom
perf/altimate-base-compaction

Conversation

@anandgupta42

@anandgupta42 anandgupta42 commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Why

altimate-base production contexts are input p50 52k / p90 82k / p99 96k tokens and latency is prefill-dominated, but compaction essentially never fires — its overflow threshold is ~99k (base 131k − headroom 32k) and 0% of sessions exceed it. So the large, prefill-heavy contexts are re-sent and re-prefilled every turn with no size reduction.

What changed (packages/opencode/src/session/compaction.ts)

  • ALTIMATE_BASE_COMPACTION_THRESHOLD = 70_000 — the p75+ tail (between input p50 52k and p90 82k). Compacts the largest ~quarter of sessions, deliberately not the median, to limit the summarization/quality cost.
  • overflowThreshold() takes an optional model; only for FreeTier.PROVIDER_ID + MODEL_ID it caps the trigger at the 70k effective limit (Math.min(threshold, effectiveContextLimit(70k, fraction)), so it never raises a stricter model/config limit). Threaded through every caller (isOverflow, retention-budget, pin sizing).

Before → after

  • altimate-base session at ~70k tokens: before → no compaction (99k threshold) → full ~70k prefill every turn; after → compaction triggers → subsequent turns prefill a compacted context.
  • Any other model/provider: unchanged (returns the normal threshold).

Tenant/user impact

altimate-base users on long sessions get faster turns (less prefill) at the cost of earlier summarization on the largest ~25% of sessions (some context fidelity traded for speed). No change for any other model. Tunable via the one constant.

How tested

  • bun test test/session/compaction.test.ts: 41 pass, 0 fail — includes a non-altimate-base isolation test (another model's threshold unchanged) and just-below/just-above boundary tests at 70k. Mutation-verified the altimate-base branch.
  • Not run: live latency impact — needs deploy + telemetry; estimate ~40% fewer prefilled tokens on the affected tail → proportional TTFT drop on those turns.

Notes

altimate_change markers present per the upstream-marker convention; change is isolated to the free-tier model.


Note

Cursor Bugbot is generating a summary for commit 23f2a51. Configure here.


Summary by cubic

Lowers the altimate-base compaction trigger from ~99k to 70k tokens so the large, prefill-heavy tail of sessions actually compacts, cutting prefill cost on repeated turns.

  • The 70k cap targets the p75+ tail (input p50 52k / p90 82k), so most sessions are unaffected and only the largest ~quarter summarize earlier.
  • The cap applies only to FreeTier.PROVIDER_ID + MODEL_ID; all other models and providers keep their normal threshold unchanged.
  • The cap is threaded through overflow detection, retention budgets, ledger budget, and pin sizing so downstream budgets stay consistent with the earlier trigger.
  • Stricter configured headroom or model limits are never raised; they still take precedence.

Written for commit 23f2a51. Summary will update on new commits.

Review in cubic

Summary by CodeRabbit

  • Bug Fixes
    • Improved automatic conversation compaction for the Altimate Free Base model, applying a 70,000-token threshold when model limits don’t provide a stricter boundary.
    • Compaction continues to account for configured safety margins and model limits, and other models retain their existing thresholds.

altimate-base production contexts are input p50 52k / p90 82k / p99 96k, and
latency is prefill-dominated, but compaction essentially never fires (its
overflow threshold is ~99k and 0% of sessions exceed it). Add an
altimate-base/altimate-free-scoped overflow cap so the large-context tail
compacts:

- ALTIMATE_BASE_COMPACTION_THRESHOLD = 70_000 (the p75+ tail between input
  p50=52k and p90=82k) -- compacts the largest ~quarter of sessions, not the
  median, to limit the summarization/quality cost.
- overflowThreshold() takes an optional model and, ONLY for
  FreeTier.PROVIDER_ID/MODEL_ID, caps the trigger at the 70k effective limit;
  all other models/providers are byte-for-byte unchanged. Threaded through every
  caller (isOverflow, retention-budget, pin sizing).

bun test test/session/compaction.test.ts: 41 pass, 0 fail (incl. a non-altimate
isolation test and just-below/above boundary tests).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review for a one-time review, or @claude review always to subscribe this PR to a review on every future push.

Tip: disable this comment in your organization's Code Review settings.

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@cursor

cursor Bot commented Oct 2, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_e3949670-ab62-4a57-8752-0a2291d48349)

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

Hey! Your PR title perf(compaction): compact altimate-base earlier to cut prefill cost doesn't follow conventional commit format.

Please update it to start with one of:

  • feat: or feat(scope): new feature
  • fix: or fix(scope): bug fix
  • docs: or docs(scope): documentation changes
  • chore: or chore(scope): maintenance tasks
  • refactor: or refactor(scope): code refactoring
  • test: or test(scope): adding or updating tests

Where scope is the package name (e.g., app, desktop, opencode).

See CONTRIBUTING.md for details.

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

This PR doesn't fully meet our contributing guidelines and PR template.

What needs to be fixed:

  • PR description is missing required template sections. Please use the PR template.

Please edit this PR description to address the above within 2 hours, or it will be automatically closed.

If you believe this was flagged incorrectly, please let a maintainer know.

@coderabbitai

coderabbitai Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

🧰 Additional context used
📚 Code guidelines (2)
packages/opencode/specs/effect/migration.md — configured
packages/opencode/AGENTS.md — auto-discovered

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Advanced

Run ID: 374adf1d-345a-4066-bda6-66343d94a8be

📥 Commits

Reviewing files that changed from the base of the PR and between 3ab191c and 23f2a51.

📒 Files selected for processing (2)
  • packages/opencode/src/session/compaction.ts
  • packages/opencode/test/session/compaction.test.ts

Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 1 remain after this review.


📝 Walkthrough

Walkthrough

Compaction now applies a model-aware 70,000-token base threshold to altimate-free/altimate-base, while retaining stricter existing limits and existing thresholds for other models. Overflow checks and ledger, retained-tail, and pin budgets pass their model to the shared threshold calculation. Tests cover model matching, usage boundaries, configuration, and budget effects.

Changes

Model-Aware Compaction Threshold

Layer / File(s) Summary
Define and verify the Altimate threshold
packages/opencode/src/session/compaction.ts, packages/opencode/test/session/compaction.test.ts
The shared threshold calculation caps the Altimate base model at a fraction-adjusted 70,000 tokens, subject to stricter existing limits. Tests cover model matching, overflow boundaries, and configuration cases.
Apply the threshold to overflow and budgets
packages/opencode/src/session/compaction.ts, packages/opencode/test/session/compaction.test.ts
Overflow detection and ledger, retained-tail, and pin budget calculations pass their model to the shared threshold function. Tests check the resulting threshold and budget values.

Priority: ➖ Normal

Estimated code review effort: 2 (Simple) | ~12 minutes

Change: Feature

Merge Risk: ⚪ Minimal · up to 23f2a

The change compacts the altimate-base model earlier at 70k tokens and leaves other models unchanged. The tests cover the boundaries, and no outstanding defects were found.

Security Architecture Review

Security architecture risk: 🔵 Low · up to 23f2a

Earlier summarization remains scoped to the originating conversation and does not add permissions. No introduced security defect was established, but recovery from overlapping cancellation or interrupted writes remains incompletely demonstrated.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The demonstrated exposure change is earlier summarization and smaller retained-context budgets for sessions using the exact targeted identity pair. Inspected summary, replay, and continuation writes remain keyed to the originating session. This establishes local scope, not comprehensive tenant-isolation assurance.

Security Findings and Attack Paths

  • inferred — User and tool content can contribute to the usage that initiates compaction, so the lower threshold increases opportunities to enter the existing summarization path. The inspected change does not add a tool-execution sink or broaden message destinations. No introduced authorization bypass or cross-session disclosure was established along this path.

Trust Boundaries and Controls

  • observed — The ledger records completed or errored tool events and renders bounded advisory context; it is not the inspected permission store. User tool restrictions are persisted as session permissions, and tool permission requests merge agent and session rules. This is counterevidence to treating reduced ledger fidelity as an authorization-control loss.

Resilience and Maintainability Implications

  • observed — Existing containment includes a per-session attempt limit, one identical-input retry for an empty summary, error-based stopping, and history filtering that recognizes only finished summaries without errors. Ordinary concurrent prompt requests queue on the active generation. Cancellation can replace a closing generation, and persistence exposes separate per-record upserts; complete atomicity and stale-generation write exclusion across interruption remain unproven.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 42.86% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 7 functions across 2 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly describes the main change: triggering earlier compaction for altimate-base to reduce prefill cost.
Description check ✅ Passed The description is detailed and relevant. It explains the problem, implementation, impact, testing results, and known limitations. It omits explicit template sections for the issue number, change type…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 2
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

A rabbit checks the token line,
Seventy thousand marks the sign.
Model names guide the threshold’s place,
Stricter limits keep their pace.
Budgets follow, tests confirm,
Then carrots wait beyond the term.

Comment @coderabbitai help to get the list of available commands.

@kilo-code-bot

kilo-code-bot Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Code Review Summary

Status: No Issues Found | Recommendation: Merge

Files Reviewed (2 files)
  • packages/opencode/src/session/compaction.ts
  • packages/opencode/test/session/compaction.test.ts

Reviewed by gpt-sol-latest · Input: 0 · Output: 0 · Cached: 0

Review guidance: REVIEW.md from base branch main

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 2 files

Re-trigger cubic

@anandgupta42

Copy link
Copy Markdown
Contributor Author

Closing: production data shows this change would be net-negative, not a win.

The premise ("cut prefill cost") doesn't hold because altimate-base is served by SGLang with RadixAttention prefix caching, and the production evidence is decisive:

  • Prefix-cache hit rate ≈ 88% on the live WARM server (sglang:prefill_effective_tokens_total): 340.9M tokens served from the device prefix cache vs only 47.9M actually prefilled (cache miss). Prefill is overwhelmingly cache-served.
  • 83% of altimate-base sessions are multi-turn (502 sessions/7d: p50 = 5 turns, p90 = 22, max 310; only 17% single-turn) — exactly the regime where the prefix cache pays off and where compaction would hurt most.
  • Compaction rewrites the conversation prefix, which invalidates the radix cache and forces a full re-prefill of the post-compaction context — converting free cache hits into expensive L4 prefill compute, on the deepest/most-cached sessions.
  • KV is not the bottleneck: kv_cache_memory_usage_gb ≈ 4.09 GB/GPU out of ~10 GB free per L4 (kv_usage 0.16–0.66), so there is no footprint pressure for earlier compaction to relieve. It would only add the context-fidelity cost of earlier summarization.

Note: client-side tokens_cache_read reads 0 for altimate-base (the gateway doesn't surface SGLang's cache to the client), which is why the server-side radix cache is easy to overlook — but the server is caching 88% of prefill.

The real performance lever for altimate-base remains the H100 BURST autoscaler work (altimate-gateway #77/#78, already merged), which targets slow L4 prefill on cache misses and premature escalation — not compaction.

Reopen if altimate-base ever moves to a backend without prefix caching.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant