Skip to content

fix(agent-core-v2): omit completion token cap unless explicitly configured - #4091

Merged
sailist merged 1 commit into
mainfrom
bug-264-09-22-overflow-budget-chain
Sep 30, 2026
Merged

sailist merged 1 commit into
mainfrom
bug-264-09-22-overflow-budget-chain

Conversation

@sailist

@sailist sailist commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

Requirement or Bug

Supersedes #3978 (same change set, rebased onto latest main, now pushed from an origin branch). Internal incident (2026-09-20): third-party models served on strict OpenAI-compatible stacks got a 400 on every request and fell into an overflow→compaction death spiral.

Bug Reproduction Steps

  1. Point an OpenAI-compatible provider at a bare vLLM deployment (validates prompt_tokens + max_tokens > max_model_len before generation, no gateway clamp), serving a model with no max_output_size configured.
  2. Send any prompt.
  3. The first request 400s (client declares max_tokens = the whole context window, so prompt + cap exceeds max_model_len); later turns send window − last measured usage, still overshooting by the unmeasured delta, so every turn 400s. Each 400 is misclassified as a genuine context overflow, so compaction fires and retries in a loop — users see compaction trigger on every message, earlier and earlier.

Root Cause

When no explicit output limit is configured, the completion-budget chain fell back to capping at the entire context window (hardCap ?? window ?? reservedContextSize ?? 32000) and always encoded that guess as max_tokens/max_completion_tokens. No client-side default can be right for every serving stack; the correct default is to omit the field and let the server allocate the actual remaining window. This is a fundamental fix on the client side (a server-side clamp is out of scope).

Code Changes

  • Budget chain: a completion cap is now produced only from explicit configuration (env KIMI_MODEL_MAX_COMPLETION_TOKENS or model max_output_size); otherwise the cap is undefined and the openai / openai-responses / google-genai egresses omit max_tokens/max_completion_tokens, letting the server decide output headroom. The whole-context-window and 32000 fallback chain (reservedContextSize, DEFAULT_UNKNOWN_CONTEXT_FALLBACK, computeCompletionBudgetCap, CompletionBudgetConfig.fallback) is removed.
  • kimi provider exception: when no cap is configured, kimi's openai / openai-responses egress computes max(1, model.maxContextSize − usedContextTokens) at the protocol layer (kimiUnsetCompletionTokens), so kimi requests still fill the remaining context window. Nothing is sent when the window or the usage figure is unknown, and an explicit opt-out (<= 0) is honored.
  • Anthropic keeps always sending max_tokens from the built-in per-model ceiling table; the fallback for unrecognized models is lowered from 128000 to 64000 (unknown models may be old ones whose server ceiling is below 128000).
  • The remaining-window subtraction now counts size (measured anchor + estimated tail) instead of measured, covering messages added between turns and aligning with the human domain.
  • Tests adapted in place: budget unit tests reshaped to the omit-by-default contract, wire snapshots drop the default maxTokens field, the compaction unknown-window case asserts omission, trait tests cover the kimi egress value. No new test files; test count not increased. Patch changeset included.

Behavior Changes and Affected Users

Behavior Before After Who relies on the old behavior Escape hatch
openai / openai-responses / google-genai requests without an explicit output limit send max_tokens/max_completion_tokens = whole context window (or 32000 fallback) field omitted; server decides permissive hosted APIs see no change; strict vLLM stacks previously 400'd set max_output_size or KIMI_MODEL_MAX_COMPLETION_TOKENS
kimi provider requests without an explicit output limit send whole-context-window cap send max(1, maxContextSize − usedContextTokens); omit when window/usage unknown kimi served models explicit cap via config; <= 0 opts out
Anthropic requests for models missing from the built-in table max_tokens: 128000 max_tokens: 64000 unknown/older Anthropic models with lower server ceilings explicit cap via config
Requests with an explicit cap (env / max_output_size / compaction override) cap reduced by last measured usage cap reduced by measured + estimated tail none (same budget, more accurate accounting) —
wire llm.request.maxTokens telemetry always present absent when no cap is configured consumers expecting the field on every request configure an explicit cap

Coverage by module: llm-adapter/model/completion-budget.ts (completionBudget.test.ts), agent/llmRequester/llmRequesterService.ts (loop/tool/plan/config wire snapshots), human/llm/protocol/format.ts + openai/openai-responses/kimi traits (trait.test.ts), human/llm/requester/bases/anthropic/profile.ts (trait.test.ts), compaction override (fullCompaction.test.ts).

Checklist

  • I have read the CONTRIBUTING document.
  • I have linked a related issue — N/A: internal incident (2026-09-20), reproduction and root cause above; supersedes fix(agent-core-v2): omit completion token cap unless explicitly configured #3978.
  • I have added tests that prove my feature works — existing budget / trait / compaction / wire-snapshot suites reshaped to the new contract.
  • The behavior-change table above is complete, and every removed behavior or flipped default is named in the changeset and either has an escape hatch or was explicitly approved by a maintainer in this PR — escape hatch for each row is the explicit cap config.
  • Ran gen-changesets skill — .changeset/omit-default-completion-token-cap.md included.
  • Ran gen-docs skill — provider docs mention of the old default will be updated in a follow-up (carried over from fix(agent-core-v2): omit completion token cap unless explicitly configured #3978).

@changeset-bot

changeset-bot Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: e059783

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 1 package
Name Type
@moonshot-ai/kimi-code Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@pkg-pr-new

pkg-pr-new Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
pnpm dlx https://pkg.pr.new/@moonshot-ai/kimi-code@e059783
npx https://pkg.pr.new/@moonshot-ai/kimi-code@e059783

commit: e059783

@sailist
sailist force-pushed the bug-264-09-22-overflow-budget-chain branch from 8e1f67e to e059783 Compare September 30, 2026 03:59

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8e1f67e97b

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +75 to +78
return Math.max(1, window - input.usedContextTokens);
}

export const kimiResponsesTrait: OpenAIResponsesTrait = {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Include static prompt tokens in the Kimi remaining-window cap

For a Kimi OpenAI/Responses request with no configured output limit, this computes the cap from usedContextTokens, but the new caller obtains that value from the context-message counter. On a session's first request that counter estimates only the messages; it does not include the system prompt or tool schemas, which are sent separately in the same request. Consequently, a strict Kimi-compatible server receives max_completion_tokens equal to nearly the full window plus the static prompt/tool cost and rejects the initial request for exceeding its context limit. Derive this from the complete outbound request size (or omit the cap until such a measurement exists).

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已知行为,符合预期

@sailist
sailist merged commit 21406fb into main Sep 30, 2026
15 checks passed
@sailist
sailist deleted the bug-264-09-22-overflow-budget-chain branch September 30, 2026 07:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants