Skip to content

Commit 048f6fe

Browse files
Merge main into docs/conversation-event-stream
Generated by OpenHands AI agent on behalf of Graham Neubig. Co-authored-by: openhands <openhands@all-hands.dev>
2 parents 064650e + c3833f2 commit 048f6fe

2 files changed

Lines changed: 49 additions & 90 deletions

File tree

‎docs.json‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -456,6 +456,7 @@
456456
"sdk/arch/agent",
457457
"sdk/arch/conversation",
458458
"sdk/arch/tool-system",
459+
"sdk/arch/mcp",
459460
"sdk/arch/events",
460461
"sdk/arch/workspace",
461462
"sdk/arch/llm",

‎sdk/arch/mcp.mdx‎

Lines changed: 48 additions & 90 deletions
Original file line numberDiff line numberDiff line change
@@ -32,7 +32,7 @@ flowchart TB
3232
end
3333
3434
subgraph Integration["Agent Integration"]
35-
Action["MCPToolAction<br><i>Dynamic model</i>"]
35+
Action["MCPToolAction<br><i>Argument wrapper</i>"]
3636
Obs["MCPToolObservation<br><i>Result wrapper</i>"]
3737
end
3838
@@ -67,9 +67,9 @@ flowchart TB
6767
| Component | Purpose | Design |
6868
|-----------|---------|--------|
6969
| **[`MCPClient`](https://github.com/OpenHands/software-agent-sdk/blob/main/openhands-sdk/openhands/sdk/mcp/client.py)** | Client wrapper | Extends FastMCP with sync/async bridge |
70-
| **[`MCPToolDefinition`](https://github.com/OpenHands/software-agent-sdk/blob/main/openhands-sdk/openhands/sdk/mcp/definition.py)** | Tool metadata | Converts MCP schemas to SDK format |
70+
| **[`MCPToolDefinition`](https://github.com/OpenHands/software-agent-sdk/blob/main/openhands-sdk/openhands/sdk/mcp/tool.py)** | Tool metadata | Converts MCP schemas to SDK format |
7171
| **[`MCPToolExecutor`](https://github.com/OpenHands/software-agent-sdk/blob/main/openhands-sdk/openhands/sdk/mcp/tool.py)** | Execution handler | Bridges agent actions to MCP calls |
72-
| **[`MCPToolAction`](https://github.com/OpenHands/software-agent-sdk/blob/main/openhands-sdk/openhands/sdk/mcp/definition.py)** | Dynamic action model | Runtime-generated Pydantic model |
72+
| **[`MCPToolAction`](https://github.com/OpenHands/software-agent-sdk/blob/main/openhands-sdk/openhands/sdk/mcp/definition.py)** | Action wrapper | Stores validated arguments in `data` |
7373
| **[`MCPToolObservation`](https://github.com/OpenHands/software-agent-sdk/blob/main/openhands-sdk/openhands/sdk/mcp/definition.py)** | Result wrapper | Wraps MCP tool results |
7474

7575
## MCP Client
@@ -109,7 +109,7 @@ flowchart TB
109109
- **Lifecycle Management:** `__enter__`/`__exit__` for context manager
110110
- **Timeout Support:** Configurable timeouts for MCP operations
111111
- **Error Handling:** Wraps MCP errors in observations
112-
- **Connection Pooling:** Reuses connections across tool calls
112+
- **Connection Reuse:** Tools share their connected MCP client
113113

114114
### MCP Server Configuration
115115

@@ -143,12 +143,12 @@ mcp_config = {
143143
%%{init: {"theme": "default", "flowchart": {"nodeSpacing": 30, "rankSpacing": 40}} }%%
144144
flowchart TB
145145
Config["MCP Config"]
146-
Spawn["Spawn Server"]
146+
Spawn["Connect to Server"]
147147
List["List Tools"]
148148
149149
subgraph Convert["Convert Each Tool"]
150150
Schema["MCP Schema"]
151-
Action["Generate Action Model"]
151+
Action["Generate Validation Model"]
152152
Def["Create ToolDefinition"]
153153
end
154154
@@ -169,67 +169,43 @@ flowchart TB
169169

170170
**Discovery Steps:**
171171

172-
1. **Spawn Server:** Launch MCP server via stdio
172+
1. **Connect:** Launch a stdio server or connect to a configured HTTP server
173173
2. **List Tools:** Call `tools/list` MCP endpoint
174174
3. **Parse Schemas:** Extract tool names, descriptions, parameters
175-
4. **Generate Models:** Dynamically create Pydantic models for actions
175+
4. **Generate Models:** Create Pydantic models from input schemas for argument validation
176176
5. **Create Definitions:** Wrap in `ToolDefinition` objects
177177
6. **Register:** Add to agent's tool registry
178178

179179
### Schema Conversion
180180

181-
MCP tool schemas are converted to SDK tool definitions:
181+
`MCPToolDefinition` keeps the original MCP tool metadata and input schema. The
182+
LLM-facing schema is built from that input schema, preserving nested properties.
183+
A separate Pydantic model derived from `Schema` validates the arguments.
182184

183-
```mermaid
184-
%%{init: {"theme": "default", "flowchart": {"nodeSpacing": 30}} }%%
185-
flowchart LR
186-
MCP["MCP Tool Schema<br><i>JSON Schema</i>"]
187-
Parse["Parse Parameters"]
188-
Model["Dynamic Pydantic Model<br><i>MCPToolAction</i>"]
189-
Def["ToolDefinition<br><i>SDK format</i>"]
190-
191-
MCP --> Parse
192-
Parse --> Model
193-
Model --> Def
194-
195-
style Parse fill:#f3e8ff,stroke:#7c3aed,stroke-width:2px
196-
style Model fill:#e8f3ff,stroke:#2b6cb0,stroke-width:2px
197-
```
185+
`MCPToolAction` is a wrapper with a `data` dictionary. Its fields do not change for
186+
each discovered tool. `action_from_arguments()` validates the arguments, removes
187+
null values and internal fields, and stores the sanitized result in `data`.
188+
The definition validates `action.data` again before execution.
198189

199-
**Conversion Rules:**
200-
201-
| MCP Schema | SDK Action Model |
202-
|------------|------------------|
203-
| **name** | Class name (camelCase) |
204-
| **description** | Docstring |
205-
| **inputSchema** | Pydantic fields |
206-
| **required** | Field(required=True) |
207-
| **type** | Python type hints |
208-
209-
**Example:**
190+
For a discovered `fetch_url` tool whose input schema accepts a string `url` and a
191+
numeric `timeout`, argument conversion looks like this:
210192

211193
```python
212-
# MCP Schema
213-
{
214-
"name": "fetch_url",
215-
"description": "Fetch content from URL",
216-
"inputSchema": {
217-
"type": "object",
218-
"properties": {
219-
"url": {"type": "string"},
220-
"timeout": {"type": "number"}
221-
},
222-
"required": ["url"]
223-
}
224-
}
194+
from openhands.sdk.mcp.tool import MCPToolDefinition
195+
225196

226-
# Generated Action Model
227-
class FetchUrl(MCPToolAction):
228-
"""Fetch content from URL"""
229-
url: str
230-
timeout: float | None = None
197+
def prepare_fetch(tool_definition: MCPToolDefinition):
198+
action = tool_definition.action_from_arguments(
199+
{"url": "https://example.com", "timeout": 10}
200+
)
201+
# action.data holds the validated arguments.
202+
# The executor forwards these arguments to the MCP server.
203+
return action.to_mcp_arguments()
231204
```
232205

206+
See [`MCPToolDefinition`](https://github.com/OpenHands/software-agent-sdk/blob/main/openhands-sdk/openhands/sdk/mcp/tool.py)
207+
for schema generation and argument validation.
208+
233209
## Tool Execution
234210

235211
### Execution Flow
@@ -267,9 +243,9 @@ flowchart TB
267243

268244
1. **Action Creation:** LLM generates tool call, parsed into `MCPToolAction`
269245
2. **Executor Lookup:** Find `MCPToolExecutor` for tool name
270-
3. **Format Conversion:** Convert action fields to MCP arguments
246+
3. **Format Conversion:** Read the argument dictionary using `action.to_mcp_arguments()`
271247
4. **MCP Call:** Execute `call_tool` via MCP client
272-
5. **Result Parsing:** Parse MCP result (text, images, resources)
248+
5. **Result Parsing:** Convert text and image blocks; log and skip unsupported blocks, including resources
273249
6. **Observation Creation:** Wrap in `MCPToolObservation`
274250
7. **Error Handling:** Catch exceptions, return error observations
275251

@@ -294,7 +270,7 @@ flowchart LR
294270
**Executor Responsibilities:**
295271
- **Client Management:** Hold reference to MCP client
296272
- **Tool Identification:** Know which MCP tool to call
297-
- **Argument Conversion:** Transform action fields to MCP format
273+
- **Argument Conversion:** Forward the action’s `data` dictionary as MCP arguments
298274
- **Result Handling:** Parse MCP responses
299275
- **Error Recovery:** Handle connection errors, timeouts, server failures
300276

@@ -352,41 +328,23 @@ flowchart TB
352328

353329
## MCP Annotations
354330

355-
MCP tools can include metadata hints for agents:
356-
357-
```mermaid
358-
%%{init: {"theme": "default", "flowchart": {"nodeSpacing": 30}} }%%
359-
flowchart LR
360-
Tool["MCP Tool"]
361-
362-
subgraph Annotations
363-
ReadOnly["readOnlyHint"]
364-
Destructive["destructiveHint"]
365-
Progress["progressEnabled"]
366-
end
367-
368-
Security["Security Analysis"]
369-
370-
Tool --> ReadOnly
371-
Tool --> Destructive
372-
Tool --> Progress
373-
374-
ReadOnly --> Security
375-
Destructive --> Security
376-
377-
style Destructive fill:#f3e8ff,stroke:#7c3aed,stroke-width:2px
378-
style Security fill:#fff4df,stroke:#b7791f,stroke-width:2px
379-
```
331+
MCP tool annotations are copied into the SDK's `ToolAnnotations` model:
380332

381-
**Annotation Types:**
333+
| Annotation | Meaning |
334+
|------------|---------|
335+
| **title** | Human-readable tool title |
336+
| **readOnlyHint** | Tool reports that it does not modify its environment |
337+
| **destructiveHint** | Tool may perform destructive updates |
338+
| **idempotentHint** | Repeated calls with the same arguments have no additional effect |
339+
| **openWorldHint** | Tool may interact with external entities |
382340

383-
| Annotation | Meaning | Use Case |
384-
|------------|---------|----------|
385-
| **readOnlyHint** | Tool doesn't modify state | Lower security risk |
386-
| **destructiveHint** | Tool modifies/deletes data | Require confirmation |
387-
| **progressEnabled** | Tool reports progress | Show progress UI |
341+
When `readOnlyHint` is true, the MCP schema adapter omits the additional
342+
`security_risk` prediction field from the LLM-facing schema. These annotations
343+
are hints, not enforcement guarantees. `destructiveHint` does not by itself
344+
require confirmation: confirmation depends on the configured security analyzer
345+
and confirmation policy. See [Security](/sdk/arch/security).
388346

389-
These annotations feed into the security analyzer for risk assessment.
347+
The SDK's `ToolAnnotations` model does not define `progressEnabled`.
390348

391349
## Component Relationships
392350

@@ -415,7 +373,7 @@ flowchart LR
415373
- **Skills → MCP**: Repository skills can embed MCP configurations
416374
- **MCP → Tools**: MCP tools registered alongside native tools
417375
- **Agent → Tools**: Agents use MCP tools like any other tool
418-
- **MCP → Security**: Annotations inform security risk assessment
376+
- **MCP → Security**: Read-only hints affect risk-prediction schema generation; the configured policy governs confirmation
419377
- **Transparent Integration**: Agent doesn't distinguish MCP from native tools
420378

421379
## Design Rationale
@@ -428,7 +386,7 @@ flowchart LR
428386

429387
**FastMCP Foundation:** Building on FastMCP (MCP SDK for Python) provides battle-tested client implementation, protocol compliance, and ongoing updates as MCP evolves.
430388

431-
**Annotation Support:** Exposing MCP hints (readOnly, destructive) enables intelligent security analysis and user confirmation flows based on tool characteristics.
389+
**Annotation Support:** MCP hints are preserved as tool metadata. Read-only hints affect risk-prediction schema generation, while confirmation is controlled by the configured policy.
432390

433391
**Lifecycle Management:** Automatic spawn/cleanup of MCP servers in conversation lifecycle ensures resources are properly managed without manual bookkeeping.
434392

0 commit comments

Comments
 (0)