Problem
When using kimi-cli with an OpenAI-compatible backend (e.g. SGLang serving Kimi-K2.5), the server occasionally returns streaming responses that complete without any content or tool calls. After exhausting 5 retries (tenacity), the error surfaces to the user as:
LLM provider error: The API returned an empty response.
The current turn is lost and the user must re-enter their prompt.
Root Cause
In packages/kosong/src/kosong/_generate.py, after the stream finishes, if the assembled message has neither content nor tool_calls, an APIEmptyResponseError is raised. The retry logic in kimisoul.py (_kosong_step_with_retry) correctly treats this as retryable, but with only 5 attempts and short backoff (0.3s initial, 5s max), all retries can exhaust quickly when the backend is transiently returning empty responses.
The shell UI (src/kimi_cli/ui/shell/__init__.py) has no specific handler for APIEmptyResponseError — it falls through to the generic ChatProviderError case with a terse red error message, which can be confusing for users since the server is reachable.
Proposed Fix
-
Better diagnostics in _generate.py: Add a logger.warning() before raising APIEmptyResponseError that captures the stream ID and the number of chunks received. Include these details in the error message for easier troubleshooting.
-
Dedicated UI handler in shell/__init__.py: Add an APIEmptyResponseError-specific branch in the error handler that displays a yellow (warning-level) message like: "Server returned an empty response after retries. This is typically a transient issue — please try again." This makes the error less alarming than a generic red "LLM provider error" and signals that it's retryable.
Environment
- Backend: SGLang v0.4.x serving Kimi-K2.5 via OpenAI-compatible API
- Provider:
openai_legacy with stream=True
- Hardware: AMD MI355X
- kimi-cli version: 1.12.0
Reference
Fix implemented in fork: https://github.com/jhinpan/kimi-cli-fork/tree/fix/handle-empty-response-from-openai-compatible-backends
Problem
When using kimi-cli with an OpenAI-compatible backend (e.g. SGLang serving Kimi-K2.5), the server occasionally returns streaming responses that complete without any content or tool calls. After exhausting 5 retries (tenacity), the error surfaces to the user as:
The current turn is lost and the user must re-enter their prompt.
Root Cause
In
packages/kosong/src/kosong/_generate.py, after the stream finishes, if the assembled message has neithercontentnortool_calls, anAPIEmptyResponseErroris raised. The retry logic inkimisoul.py(_kosong_step_with_retry) correctly treats this as retryable, but with only 5 attempts and short backoff (0.3s initial, 5s max), all retries can exhaust quickly when the backend is transiently returning empty responses.The shell UI (
src/kimi_cli/ui/shell/__init__.py) has no specific handler forAPIEmptyResponseError— it falls through to the genericChatProviderErrorcase with a terse red error message, which can be confusing for users since the server is reachable.Proposed Fix
Better diagnostics in
_generate.py: Add alogger.warning()before raisingAPIEmptyResponseErrorthat captures the stream ID and the number of chunks received. Include these details in the error message for easier troubleshooting.Dedicated UI handler in
shell/__init__.py: Add anAPIEmptyResponseError-specific branch in the error handler that displays a yellow (warning-level) message like: "Server returned an empty response after retries. This is typically a transient issue — please try again." This makes the error less alarming than a generic red "LLM provider error" and signals that it's retryable.Environment
openai_legacywithstream=TrueReference
Fix implemented in fork: https://github.com/jhinpan/kimi-cli-fork/tree/fix/handle-empty-response-from-openai-compatible-backends