Skip to content

Commit 2ea54e6

Browse files
docs(sdk): describe typed provider boundaries and Responses streaming (#891)
* docs(sdk): describe typed provider boundaries and response streaming Co-authored-by: openhands <openhands@all-hands.dev> * docs(sdk): include system message in Responses example Co-authored-by: openhands <openhands@all-hands.dev> --------- Co-authored-by: openhands <openhands@all-hands.dev>
1 parent c397fc6 commit 2ea54e6

2 files changed

Lines changed: 57 additions & 3 deletions

File tree

‎sdk/arch/llm.mdx‎

Lines changed: 15 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -74,6 +74,21 @@ flowchart TB
7474
| **Configuration Loaders** | Config hydration | `load_from_env()`, `load_from_json()` |
7575
| **Telemetry** | Usage tracking | Token counts, costs, latency |
7676

77+
### Provider Boundaries
78+
79+
Private adapters normalize provider data before the core LLM pipeline consumes
80+
it. Stream events use typed contracts; message conversion accepts typed provider
81+
objects, generic LiteLLM objects, and dictionaries. It preserves tool-call IDs,
82+
Responses item IDs, and reasoning metadata used to replay conversation history.
83+
84+
When `custom_tokenizer` is configured, the SDK attempts to load a chat-template
85+
tokenizer through the optional Transformers package. Token counts include the
86+
chat template and tool schemas when supported. Missing Transformers, tokenizer
87+
loading failures, or unsupported template results fall back to LiteLLM token
88+
counting. Transformers is not required to import or use the SDK.
89+
90+
See [LLM Streaming](/sdk/guides/llm-streaming) for callback and completion behavior.
91+
7792
## Configuration
7893

7994
See [`LLM` source](https://github.com/OpenHands/software-agent-sdk/blob/main/openhands-sdk/openhands/sdk/llm/llm.py) for complete list of supported fields.

‎sdk/guides/llm-streaming.mdx‎

Lines changed: 42 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -5,9 +5,8 @@ description: Stream LLM responses token-by-token for real-time display and inter
55

66
import RunExampleCode from "/sdk/shared-snippets/how-to-run-example.mdx";
77

8-
<Warning>
9-
This is currently only supported for the chat completion endpoint.
10-
</Warning>
8+
Streaming is supported by both Chat Completions (`completion()` / `acompletion()`)
9+
and Responses (`responses()` / `aresponses()`).
1110

1211
> A ready-to-run example is available [here](#ready-to-run-example)!
1312
@@ -76,6 +75,46 @@ complete response. This creates a more responsive user experience, especially fo
7675
</Step>
7776
</Steps>
7877

78+
## Responses Streaming
79+
80+
For direct LLM calls, pass `stream=True` and an `on_token` callback. The call
81+
returns a complete `LLMResponse` after the stream finishes; callbacks receive
82+
`ModelResponseStream` chunks while it is being read.
83+
84+
```python
85+
from openhands.sdk.llm import Message, TextContent
86+
from openhands.sdk.llm.streaming import ModelResponseStream
87+
88+
def show_text(chunk: ModelResponseStream) -> None:
89+
for choice in chunk.choices:
90+
if choice.delta.content:
91+
print(choice.delta.content, end="", flush=True)
92+
93+
response = llm.responses(
94+
[
95+
Message(role="system", content=[TextContent(text="You are a helpful assistant.")]),
96+
Message(role="user", content=[TextContent(text="Say hello")]),
97+
],
98+
stream=True,
99+
on_token=show_text,
100+
)
101+
```
102+
103+
Use `await llm.aresponses(...)` for the asynchronous equivalent. It accepts
104+
synchronous or asynchronous callbacks.
105+
106+
Without a callback, the SDK normally falls back to a non-streaming request.
107+
Endpoints that require streaming, such as ChatGPT subscription endpoints, still
108+
drain the stream and return the complete response without a callback.
109+
110+
Instrumentation wrappers do not need to inherit from LiteLLM's stream classes.
111+
The SDK accepts synchronous iterables and asynchronous iterables on the async
112+
path, including synchronous wrappers returned to an async caller. A completed
113+
Responses event yielded by the stream remains valid if the wrapper's
114+
`completed_response` attribute is absent or `None`. A non-null wrapper completion
115+
takes precedence after the stream has been read. A stream with no completion
116+
raises `LLMNoResponseError` and follows the configured retry policy.
117+
79118
## Ready-to-run Example
80119

81120
<Note>

0 commit comments

Comments
 (0)