Repository navigation
perf(compaction): compact altimate-base earlier to cut prefill cost - #1403
anandgupta42 wants to merge 1 commit into
Conversation
altimate-base production contexts are input p50 52k / p90 82k / p99 96k, and latency is prefill-dominated, but compaction essentially never fires (its overflow threshold is ~99k and 0% of sessions exceed it). Add an altimate-base/altimate-free-scoped overflow cap so the large-context tail compacts: - ALTIMATE_BASE_COMPACTION_THRESHOLD = 70_000 (the p75+ tail between input p50=52k and p90=82k) -- compacts the largest ~quarter of sessions, not the median, to limit the summarization/quality cost. - overflowThreshold() takes an optional model and, ONLY for FreeTier.PROVIDER_ID/MODEL_ID, caps the trigger at the 70k effective limit; all other models/providers are byte-for-byte unchanged. Threaded through every caller (isOverflow, retention-budget, pin sizing). bun test test/session/compaction.test.ts: 41 pass, 0 fail (incl. a non-altimate isolation test and just-below/above boundary tests). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
There was a problem hiding this comment.
Claude Code Review
This repository is configured for manual code reviews. Comment @claude review for a one-time review, or @claude review always to subscribe this PR to a review on every future push.
Tip: disable this comment in your organization's Code Review settings.
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_e3949670-ab62-4a57-8752-0a2291d48349) |
|
Hey! Your PR title Please update it to start with one of:
Where See CONTRIBUTING.md for details. |
|
This PR doesn't fully meet our contributing guidelines and PR template. What needs to be fixed:
Please edit this PR description to address the above within 2 hours, or it will be automatically closed. If you believe this was flagged incorrectly, please let a maintainer know. |
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. 🧰 Additional context used📚 Code guidelines (2)No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (2)
Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 1 remain after this review. 📝 WalkthroughWalkthroughCompaction now applies a model-aware 70,000-token base threshold to ChangesModel-Aware Compaction Threshold
Priority: ➖ Normal Estimated code review effort: 2 (Simple) | ~12 minutes Change: Feature Merge Risk: ⚪ Minimal · up to The change compacts the altimate-base model earlier at 70k tokens and leaves other models unchanged. The tests cover the boundaries, and no outstanding defects were found. Security Architecture ReviewSecurity architecture risk: 🔵 Low · up to Earlier summarization remains scoped to the originating conversation and does not add permissions. No introduced security defect was established, but recovery from overlapping cancellation or interrupted writes remains incompletely demonstrated. Retained concerns Security review detailsSecurity Blast Radius
Security Findings and Attack Paths
Trust Boundaries and Controls
Resilience and Maintainability Implications
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches 💡 2📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
🛠️ Fix failing CI checks 💡
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. A rabbit checks the token line, Comment |
Code Review SummaryStatus: No Issues Found | Recommendation: Merge Files Reviewed (2 files)
Reviewed by gpt-sol-latest · Input: 0 · Output: 0 · Cached: 0 Review guidance: REVIEW.md from base branch |
|
Closing: production data shows this change would be net-negative, not a win. The premise ("cut prefill cost") doesn't hold because altimate-base is served by SGLang with RadixAttention prefix caching, and the production evidence is decisive:
Note: client-side The real performance lever for altimate-base remains the H100 BURST autoscaler work (altimate-gateway #77/#78, already merged), which targets slow L4 prefill on cache misses and premature escalation — not compaction. Reopen if altimate-base ever moves to a backend without prefix caching. |
Why
altimate-base production contexts are input p50 52k / p90 82k / p99 96k tokens and latency is prefill-dominated, but compaction essentially never fires — its overflow threshold is ~99k (
base 131k − headroom 32k) and 0% of sessions exceed it. So the large, prefill-heavy contexts are re-sent and re-prefilled every turn with no size reduction.What changed (
packages/opencode/src/session/compaction.ts)ALTIMATE_BASE_COMPACTION_THRESHOLD = 70_000— the p75+ tail (between input p50 52k and p90 82k). Compacts the largest ~quarter of sessions, deliberately not the median, to limit the summarization/quality cost.overflowThreshold()takes an optionalmodel; only forFreeTier.PROVIDER_ID+MODEL_IDit caps the trigger at the 70k effective limit (Math.min(threshold, effectiveContextLimit(70k, fraction)), so it never raises a stricter model/config limit). Threaded through every caller (isOverflow, retention-budget, pin sizing).Before → after
Tenant/user impact
altimate-base users on long sessions get faster turns (less prefill) at the cost of earlier summarization on the largest ~25% of sessions (some context fidelity traded for speed). No change for any other model. Tunable via the one constant.
How tested
bun test test/session/compaction.test.ts: 41 pass, 0 fail — includes a non-altimate-base isolation test (another model's threshold unchanged) and just-below/just-above boundary tests at 70k. Mutation-verified the altimate-base branch.Notes
altimate_changemarkers present per the upstream-marker convention; change is isolated to the free-tier model.Note
Cursor Bugbot is generating a summary for commit 23f2a51. Configure here.
Summary by cubic
Lowers the altimate-base compaction trigger from ~99k to 70k tokens so the large, prefill-heavy tail of sessions actually compacts, cutting prefill cost on repeated turns.
FreeTier.PROVIDER_ID+MODEL_ID; all other models and providers keep their normal threshold unchanged.Written for commit 23f2a51. Summary will update on new commits.
Summary by CodeRabbit