fix(agent-core-v2): omit completion token cap unless explicitly configured - #4091
Conversation
🦋 Changeset detectedLatest commit: e059783 The changes in this PR will be included in the next version bump. This PR includes changesets to release 1 package
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
commit: |
8e1f67e to
e059783
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8e1f67e97b
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| return Math.max(1, window - input.usedContextTokens); | ||
| } | ||
|
|
||
| export const kimiResponsesTrait: OpenAIResponsesTrait = { |
There was a problem hiding this comment.
Include static prompt tokens in the Kimi remaining-window cap
For a Kimi OpenAI/Responses request with no configured output limit, this computes the cap from usedContextTokens, but the new caller obtains that value from the context-message counter. On a session's first request that counter estimates only the messages; it does not include the system prompt or tool schemas, which are sent separately in the same request. Consequently, a strict Kimi-compatible server receives max_completion_tokens equal to nearly the full window plus the static prompt/tool cost and rejects the initial request for exceeding its context limit. Derive this from the complete outbound request size (or omit the cap until such a measurement exists).
Useful? React with 👍 / 👎.
Requirement or Bug
Supersedes #3978 (same change set, rebased onto latest main, now pushed from an origin branch). Internal incident (2026-09-20): third-party models served on strict OpenAI-compatible stacks got a 400 on every request and fell into an overflow→compaction death spiral.
Bug Reproduction Steps
prompt_tokens + max_tokens > max_model_lenbefore generation, no gateway clamp), serving a model with nomax_output_sizeconfigured.max_tokens= the whole context window, so prompt + cap exceedsmax_model_len); later turns sendwindow − last measured usage, still overshooting by the unmeasured delta, so every turn 400s. Each 400 is misclassified as a genuine context overflow, so compaction fires and retries in a loop — users see compaction trigger on every message, earlier and earlier.Root Cause
When no explicit output limit is configured, the completion-budget chain fell back to capping at the entire context window (
hardCap ?? window ?? reservedContextSize ?? 32000) and always encoded that guess asmax_tokens/max_completion_tokens. No client-side default can be right for every serving stack; the correct default is to omit the field and let the server allocate the actual remaining window. This is a fundamental fix on the client side (a server-side clamp is out of scope).Code Changes
KIMI_MODEL_MAX_COMPLETION_TOKENSor modelmax_output_size); otherwise the cap isundefinedand the openai / openai-responses / google-genai egresses omitmax_tokens/max_completion_tokens, letting the server decide output headroom. The whole-context-window and 32000 fallback chain (reservedContextSize,DEFAULT_UNKNOWN_CONTEXT_FALLBACK,computeCompletionBudgetCap,CompletionBudgetConfig.fallback) is removed.max(1, model.maxContextSize − usedContextTokens)at the protocol layer (kimiUnsetCompletionTokens), so kimi requests still fill the remaining context window. Nothing is sent when the window or the usage figure is unknown, and an explicit opt-out (<= 0) is honored.max_tokensfrom the built-in per-model ceiling table; the fallback for unrecognized models is lowered from 128000 to 64000 (unknown models may be old ones whose server ceiling is below 128000).size(measured anchor + estimated tail) instead ofmeasured, covering messages added between turns and aligning with the human domain.maxTokensfield, the compaction unknown-window case asserts omission, trait tests cover the kimi egress value. No new test files; test count not increased. Patch changeset included.Behavior Changes and Affected Users
max_tokens/max_completion_tokens= whole context window (or 32000 fallback)max_output_sizeorKIMI_MODEL_MAX_COMPLETION_TOKENSmax(1, maxContextSize − usedContextTokens); omit when window/usage unknown<= 0opts outmax_tokens: 128000max_tokens: 64000max_output_size/ compaction override)llm.request.maxTokenstelemetryCoverage by module:
llm-adapter/model/completion-budget.ts(completionBudget.test.ts),agent/llmRequester/llmRequesterService.ts(loop/tool/plan/config wire snapshots),human/llm/protocol/format.ts+ openai/openai-responses/kimi traits (trait.test.ts),human/llm/requester/bases/anthropic/profile.ts(trait.test.ts), compaction override (fullCompaction.test.ts).Checklist
gen-changesetsskill —.changeset/omit-default-completion-token-cap.mdincluded.gen-docsskill — provider docs mention of the old default will be updated in a follow-up (carried over from fix(agent-core-v2): omit completion token cap unless explicitly configured #3978).