You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
This repository was archived by the owner on Sep 23, 2026. It is now read-only.
When using kimi-cli with an OpenAI-compatible backend (e.g. SGLang serving Kimi-K2.5), the server occasionally returns streaming responses that complete without any content or tool calls. After exhausting 5 retries (tenacity), the error surfaces to the user as:
LLM provider error: The API returned an empty response.
The current turn is lost and the user must re-enter their prompt.
Root Cause
In packages/kosong/src/kosong/_generate.py, after the stream finishes, if the assembled message has neither content nor tool_calls, an APIEmptyResponseError is raised. The retry logic in kimisoul.py (_kosong_step_with_retry) correctly treats this as retryable, but with only 5 attempts and short backoff (0.3s initial, 5s max), all retries can exhaust quickly when the backend is transiently returning empty responses.
The shell UI (src/kimi_cli/ui/shell/__init__.py) has no specific handler for APIEmptyResponseError — it falls through to the generic ChatProviderError case with a terse red error message, which can be confusing for users since the server is reachable.
Proposed Fix
Better diagnostics in _generate.py: Add a logger.warning() before raising APIEmptyResponseError that captures the stream ID and the number of chunks received. Include these details in the error message for easier troubleshooting.
Dedicated UI handler in shell/__init__.py: Add an APIEmptyResponseError-specific branch in the error handler that displays a yellow (warning-level) message like: "Server returned an empty response after retries. This is typically a transient issue — please try again." This makes the error less alarming than a generic red "LLM provider error" and signals that it's retryable.
Environment
Backend: SGLang v0.4.x serving Kimi-K2.5 via OpenAI-compatible API
This discussion was converted from issue #1172 on February 20, 2026 00:47.
Heading
Bold
Italic
Quote
Code
Link
Numbered list
Unordered list
Task list
Attach files
Mention
Reference
Menu
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Problem
When using kimi-cli with an OpenAI-compatible backend (e.g. SGLang serving Kimi-K2.5), the server occasionally returns streaming responses that complete without any content or tool calls. After exhausting 5 retries (tenacity), the error surfaces to the user as:
The current turn is lost and the user must re-enter their prompt.
Root Cause
In
packages/kosong/src/kosong/_generate.py, after the stream finishes, if the assembled message has neithercontentnortool_calls, anAPIEmptyResponseErroris raised. The retry logic inkimisoul.py(_kosong_step_with_retry) correctly treats this as retryable, but with only 5 attempts and short backoff (0.3s initial, 5s max), all retries can exhaust quickly when the backend is transiently returning empty responses.The shell UI (
src/kimi_cli/ui/shell/__init__.py) has no specific handler forAPIEmptyResponseError— it falls through to the genericChatProviderErrorcase with a terse red error message, which can be confusing for users since the server is reachable.Proposed Fix
Better diagnostics in
_generate.py: Add alogger.warning()before raisingAPIEmptyResponseErrorthat captures the stream ID and the number of chunks received. Include these details in the error message for easier troubleshooting.Dedicated UI handler in
shell/__init__.py: Add anAPIEmptyResponseError-specific branch in the error handler that displays a yellow (warning-level) message like: "Server returned an empty response after retries. This is typically a transient issue — please try again." This makes the error less alarming than a generic red "LLM provider error" and signals that it's retryable.Environment
openai_legacywithstream=TrueReference
Fix implemented in fork: https://github.com/jhinpan/kimi-cli-fork/tree/fix/handle-empty-response-from-openai-compatible-backends
All reactions