Skip to content

Commit c91e368

Browse files
committed
done with llm and condenser
1 parent 41be81b commit c91e368

8 files changed

Lines changed: 386 additions & 337 deletions

File tree

‎docs.json‎

Lines changed: 9 additions & 14 deletions
Original file line numberDiff line numberDiff line change
@@ -181,20 +181,18 @@
181181
{
182182
"group": "Guides",
183183
"pages": [
184+
"sdk/guides/hello-world",
185+
"sdk/guides/custom-tools",
186+
"sdk/guides/mcp",
187+
"sdk/guides/activate-skill",
188+
"sdk/guides/context-condenser",
184189
{
185-
"group": "Getting Started",
186-
"pages": [
187-
"sdk/guides/hello-world",
188-
"sdk/guides/custom-tools",
189-
"sdk/guides/mcp"
190-
]
191-
},
192-
{
193-
"group": "Agent Configuration",
190+
"group": "LLM Configuration",
194191
"pages": [
195192
"sdk/guides/llm-registry",
196193
"sdk/guides/llm-routing",
197-
"sdk/guides/model-reasoning"
194+
"sdk/guides/llm-reasoning",
195+
"sdk/guides/llm-metrics"
198196
]
199197
},
200198
{
@@ -204,15 +202,12 @@
204202
"sdk/guides/pause-and-resume",
205203
"sdk/guides/confirmation-mode",
206204
"sdk/guides/send-message-while-processing",
207-
"sdk/guides/conversation-costs",
208-
"sdk/guides/llm-metrics",
209-
"sdk/guides/context-condenser"
205+
"sdk/guides/conversation-costs"
210206
]
211207
},
212208
{
213209
"group": "Agent Capabilities",
214210
"pages": [
215-
"sdk/guides/activate-skill",
216211
"sdk/guides/async",
217212
"sdk/guides/planning-agent-workflow",
218213
"sdk/guides/browser-use",

‎sdk/guides/activate-skill.mdx‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,5 @@
11
---
2-
title: Skills
2+
title: Agent Skills
33
description: Skills add specialized behaviors, domain knowledge, and context-aware triggers to your agent through structured prompts.
44
---
55

@@ -128,7 +128,7 @@ uv run python examples/01_standalone_sdk/03_activate_skill.py
128128

129129
Skills are defined with a name, content (the instructions), and an optional trigger:
130130

131-
```python highlight={2-5,8-10}
131+
```python highlight={3-14}
132132
agent_context = AgentContext(
133133
skills=[
134134
Skill(
@@ -151,7 +151,7 @@ agent_context = AgentContext(
151151

152152
Use `KeywordTrigger` to activate skills only when specific words appear:
153153

154-
```python highlight={3}
154+
```python highlight={4}
155155
Skill(
156156
name="magic-word",
157157
content="Special instructions when magic word is detected",
@@ -163,7 +163,7 @@ Skill(
163163

164164
Add consistent prefixes or suffixes to system and user messages:
165165

166-
```python highlight={2-3}
166+
```python highlight={3-4}
167167
agent_context = AgentContext(
168168
skills=[...],
169169
system_message_suffix="Always finish your response with the word 'yay!'",

‎sdk/guides/context-condenser.mdx‎

Lines changed: 59 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -3,13 +3,59 @@ title: Context Condenser
33
description: Manage agent memory by condensing conversation history to save tokens.
44
---
55

6+
## What is a Context Condenser?
7+
8+
A **context condenser** is a crucial component that addresses one of the most persistent challenges in AI agent development: managing growing conversation context efficiently. As conversations with AI agents grow longer, the cumulative history leads to:
9+
10+
- **💰 Increased API Costs**: More tokens in the context means higher costs per API call
11+
- **⏱️ Slower Response Times**: Larger contexts take longer to process
12+
- **📉 Reduced Effectiveness**: LLMs become less effective when dealing with excessive irrelevant information
13+
14+
The context condenser solves this by intelligently summarizing older parts of the conversation while preserving essential information needed for the agent to continue working effectively.
15+
16+
## Default Implementation: LLMSummarizingCondenser
17+
18+
OpenHands SDK provides `LLMSummarizingCondenser` as the default condenser implementation. This condenser uses an LLM to generate summaries of conversation history when it exceeds the configured size limit.
19+
20+
### How It Works
21+
22+
When conversation history exceeds a defined threshold, the LLM-based condenser:
23+
24+
1. **Keeps recent messages intact** - The most recent exchanges remain unchanged for immediate context
25+
2. **Preserves key information** - Important details like user goals, technical specifications, and critical files are retained
26+
3. **Summarizes older content** - Earlier parts of the conversation are condensed into concise summaries using LLM-generated summaries
27+
4. **Maintains continuity** - The agent retains awareness of past progress without processing every historical interaction
28+
29+
![Condenser Overview](https://openhands.dev/assets/blog/20250409-oh-condenser-release/condenser-overview.png)
30+
31+
32+
This approach achieves remarkable efficiency gains:
33+
- Up to **2x reduction** in per-turn API costs
34+
- **Consistent response times** even in long sessions
35+
- **Equivalent or better performance** on software engineering tasks
36+
37+
Learn more about the implementation and benchmarks in our [blog post on context condensation](https://openhands.dev/blog/openhands-context-condensensation-for-more-efficient-ai-agents).
38+
39+
### Extensibility
40+
41+
The `LLMSummarizingCondenser` extends the `RollingCondenser` base class, which provides a framework for condensers that work with rolling conversation history. You can create custom condensers by extending base classes ([source code](https://github.com/All-Hands-AI/agent-sdk/blob/main/openhands-sdk/openhands/sdk/context/condenser/base.py)):
42+
43+
- **`RollingCondenser`** - For condensers that apply condensation to rolling history
44+
- **`CondenserBase`** - For more specialized condensation strategies
45+
46+
This architecture allows you to implement custom condensation logic tailored to your specific needs while leveraging the SDK's conversation management infrastructure.
47+
48+
49+
### Example Usage
50+
651
<Note>
752
This example is available on GitHub: [examples/01_standalone_sdk/14_context_condenser.py](https://github.com/All-Hands-AI/agent-sdk/blob/main/examples/01_standalone_sdk/14_context_condenser.py)
853
</Note>
954

55+
1056
Automatically condense conversation history when context length exceeds limits, reducing token usage while preserving important information:
1157

12-
```python icon="python" examples/01_standalone_sdk/14_context_condenser.py
58+
```python icon="python" expandable examples/01_standalone_sdk/14_context_condenser.py
1359
"""
1460
To manage context in long-running conversations, the agent can use a context condenser
1561
that keeps the conversation history within a specified size limit. This example
@@ -158,17 +204,24 @@ uv run python examples/01_standalone_sdk/14_context_condenser.py
158204

159205
### Setting Up Condensing
160206

161-
Configure a condenser when creating the agent:
207+
Create a `LLMSummarizingCondenser` to manage the context.
208+
The condenser will automatically truncate conversation history when it exceeds max_size, and replaces the dropped events with an LLM-generated summary.
209+
210+
This condenser triggers when there are more than `max_context_length` events in
211+
the conversation history, and always keeps the first `keep_first` events (system prompts,
212+
initial user messages) to preserve important context.
162213

163214
```python highlight={3-4}
164-
from openhands.sdk.context import LLMCondenser
215+
from openhands.sdk.context import LLMSummarizingCondenser
165216

166-
condenser = LLMCondenser(llm=llm, max_context_length=100000)
217+
condenser = LLMSummarizingCondenser(
218+
llm=llm.model_copy(update={"usage_id": "condenser"}), max_size=10, keep_first=2
219+
)
220+
221+
# Agent with condenser
167222
agent = Agent(llm=llm, tools=tools, condenser=condenser)
168223
```
169224

170-
When context exceeds `max_context_length`, the condenser summarizes older messages to reduce token usage while maintaining important information.
171-
172225
## Next Steps
173226

174227
- **[LLM Metrics](/sdk/guides/llm-metrics)** - Track token usage reduction

‎sdk/guides/llm-metrics.mdx‎

Lines changed: 20 additions & 21 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,5 @@
11
---
2-
title: LLM Metrics
2+
title: Metrics Tracking
33
description: Track token usage, costs, and performance metrics for your agents.
44
---
55

@@ -99,31 +99,30 @@ uv run python examples/01_standalone_sdk/13_get_llm_metrics.py
9999

100100
### Getting Metrics
101101

102-
Access metrics after running the conversation:
102+
Access metrics directly from the LLM object after running the conversation:
103103

104-
```python highlight={3-6}
104+
```python highlight={3-4}
105105
conversation.run()
106106

107-
metrics = conversation.get_llm_metrics()
108-
print(f"Input tokens: {metrics.input_tokens}")
109-
print(f"Output tokens: {metrics.output_tokens}")
110-
print(f"Total cost: ${metrics.cost:.4f}")
111-
```
112-
113-
### Tracking Changes Over Time
114-
115-
Compare metrics between operations:
116-
117-
```python highlight={1,4}
118-
initial_metrics = conversation.get_llm_metrics()
119-
conversation.send_message("more work")
120-
conversation.run()
121-
final_metrics = conversation.get_llm_metrics()
122-
123-
tokens_used = final_metrics.input_tokens - initial_metrics.input_tokens
107+
assert llm.metrics is not None
108+
print(f"Final LLM metrics: {llm.metrics.model_dump()}")
124109
```
125110

126-
Metrics include: `input_tokens`, `output_tokens`, `cost`, `api_calls`, and `cache_reads` (if supported).
111+
The `llm.metrics` object is an instance of the [Metrics class](https://github.com/All-Hands-AI/agent-sdk/blob/main/openhands-sdk/openhands/sdk/llm/utils/metrics.py), which provides detailed information including:
112+
113+
- `accumulated_cost` - Total accumulated cost across all API calls
114+
- `accumulated_token_usage` - Aggregated token usage with fields like:
115+
- `prompt_tokens` - Number of input tokens processed
116+
- `completion_tokens` - Number of output tokens generated
117+
- `cache_read_tokens` - Cache hits (if supported by the model)
118+
- `cache_write_tokens` - Cache writes (if supported by the model)
119+
- `reasoning_tokens` - Reasoning tokens (for models that support extended thinking)
120+
- `context_window` - Context window size used
121+
- `costs` - List of individual cost records per API call
122+
- `token_usages` - List of detailed token usage records per API call
123+
- `response_latencies` - List of response latency metrics per API call
124+
125+
For more details on the available metrics and methods, refer to the [source code](https://github.com/All-Hands-AI/agent-sdk/blob/main/openhands-sdk/openhands/sdk/llm/utils/metrics.py).
127126

128127
## Next Steps
129128

0 commit comments

Comments
 (0)