Skip to content

Commit bbdf4df

Browse files
committed
add docs for agent
1 parent 92cb9a8 commit bbdf4df

1 file changed

Lines changed: 279 additions & 0 deletions

File tree

‎sdk/arch/agent.mdx‎

Lines changed: 279 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,279 @@
1+
---
2+
title: Agent
3+
description: High-level architecture of the reasoning-action loop
4+
---
5+
6+
The **Agent** component implements the core reasoning-action loop that drives autonomous task execution. It orchestrates LLM queries, tool execution, and context management through a stateless, event-driven architecture.
7+
8+
**Source:** [`openhands-sdk/openhands/sdk/agent/`](https://github.com/All-Hands-AI/agent-sdk/tree/main/openhands-sdk/openhands/sdk/agent)
9+
10+
## Core Responsibilities
11+
12+
The Agent system has four primary responsibilities:
13+
14+
1. **Reasoning-Action Loop** - Query LLM to generate next actions based on conversation history
15+
2. **Tool Orchestration** - Select and execute tools, handle results and errors
16+
3. **Context Management** - Apply [skills](/sdk/guides/skill), manage conversation history via [condensers](/sdk/guides/context-condenser)
17+
4. **Security Validation** - Analyze proposed actions for safety before execution via [security analyzer](/sdk/guides/security)
18+
19+
## Architecture
20+
21+
```mermaid
22+
%%{init: {"theme": "default", "flowchart": {"nodeSpacing": 20, "rankSpacing": 50}} }%%
23+
flowchart TB
24+
subgraph Input[" "]
25+
Events["Event History"]
26+
Context["Agent Context<br><i>Skills + Prompts</i>"]
27+
end
28+
29+
subgraph Core["Agent Core"]
30+
Condense["Condenser<br><i>History compression</i>"]
31+
Reason["LLM Query<br><i>Generate actions</i>"]
32+
Security["Security Analyzer<br><i>Risk assessment</i>"]
33+
end
34+
35+
subgraph Execution[" "]
36+
Tools["Tool Executor<br><i>Action → Observation</i>"]
37+
Results["Observation Events"]
38+
end
39+
40+
Events --> Condense
41+
Context -.->|Skills| Reason
42+
Condense --> Reason
43+
Reason --> Security
44+
Security --> Tools
45+
Tools --> Results
46+
Results -.->|Feedback| Events
47+
48+
classDef primary fill:#f3e8ff,stroke:#7c3aed,stroke-width:2px
49+
classDef secondary fill:#e8f3ff,stroke:#2b6cb0,stroke-width:2px
50+
classDef tertiary fill:#fff4df,stroke:#b7791f,stroke-width:2px
51+
52+
class Reason primary
53+
class Condense,Security secondary
54+
class Tools tertiary
55+
```
56+
57+
### Key Components
58+
59+
| Component | Purpose | Design |
60+
|-----------|---------|--------|
61+
| **[`Agent`](https://github.com/All-Hands-AI/agent-sdk/blob/main/openhands-sdk/openhands/sdk/agent/agent.py)** | Main implementation | Stateless reasoning-action loop executor |
62+
| **[`AgentBase`](https://github.com/All-Hands-AI/agent-sdk/blob/main/openhands-sdk/openhands/sdk/agent/base.py)** | Abstract base class | Defines agent interface and initialization |
63+
| **[`AgentContext`](https://github.com/All-Hands-AI/agent-sdk/blob/main/openhands-sdk/openhands/sdk/context/agent_context.py)** | Context container | Manages skills, prompts, and metadata |
64+
| **[`Condenser`](https://github.com/All-Hands-AI/agent-sdk/blob/main/openhands-sdk/openhands/sdk/context/condenser/)** | History compression | Reduces context when token limits approached |
65+
| **[`SecurityAnalyzer`](https://github.com/All-Hands-AI/agent-sdk/blob/main/openhands-sdk/openhands/sdk/security/)** | Safety validation | Evaluates action risk before execution |
66+
67+
## Reasoning-Action Loop
68+
69+
The agent operates through a **single-step execution model** where each `step()` call processes one reasoning cycle:
70+
71+
```mermaid
72+
%%{init: {"theme": "default", "flowchart": {"nodeSpacing": 10, "rankSpacing": 10}} }%%
73+
flowchart TB
74+
Start["step() called"]
75+
Pending{"Pending<br>actions?"}
76+
ExecutePending["Execute pending actions"]
77+
78+
HasCondenser{"Has<br>condenser?"}
79+
Condense["Call condenser.condense()"]
80+
CondenseResult{"Result<br>type?"}
81+
EmitCondensation["Emit Condensation event"]
82+
UseView["Use View events"]
83+
UseRaw["Use raw events"]
84+
85+
Query["Query LLM with messages"]
86+
ContextExceeded{"Context<br>window<br>exceeded?"}
87+
EmitRequest["Emit CondensationRequest"]
88+
89+
Parse{"Response<br>type?"}
90+
CreateActions["Create ActionEvents"]
91+
CreateMessage["Create MessageEvent"]
92+
93+
Confirmation{"Need<br>confirmation?"}
94+
SetWaiting["Set WAITING_FOR_CONFIRMATION"]
95+
96+
Execute["Execute actions"]
97+
Observe["Create ObservationEvents"]
98+
99+
Return["Return"]
100+
101+
Start --> Pending
102+
Pending -->|Yes| ExecutePending --> Return
103+
Pending -->|No| HasCondenser
104+
105+
HasCondenser -->|Yes| Condense
106+
HasCondenser -->|No| UseRaw
107+
Condense --> CondenseResult
108+
CondenseResult -->|Condensation| EmitCondensation --> Return
109+
CondenseResult -->|View| UseView --> Query
110+
UseRaw --> Query
111+
112+
Query --> ContextExceeded
113+
ContextExceeded -->|Yes| EmitRequest --> Return
114+
ContextExceeded -->|No| Parse
115+
116+
Parse -->|Tool calls| CreateActions
117+
Parse -->|Message| CreateMessage --> Return
118+
119+
CreateActions --> Confirmation
120+
Confirmation -->|Yes| SetWaiting --> Return
121+
Confirmation -->|No| Execute
122+
123+
Execute --> Observe
124+
Observe --> Return
125+
126+
style Query fill:#f3e8ff,stroke:#7c3aed,stroke-width:2px
127+
style Condense fill:#e8f3ff,stroke:#2b6cb0,stroke-width:2px
128+
style Confirmation fill:#fff4df,stroke:#b7791f,stroke-width:2px
129+
```
130+
131+
**Step Execution Flow:**
132+
133+
1. **Pending Actions:** If actions awaiting confirmation exist, execute them and return
134+
2. **Condensation:** If condenser exists:
135+
- Call `condenser.condense()` with current event view
136+
- If returns `View`: use condensed events for LLM query (continue in same step)
137+
- If returns `Condensation`: emit event and return (will be processed next step)
138+
3. **LLM Query:** Query LLM with messages from event history
139+
- If context window exceeded: emit `CondensationRequest` and return
140+
4. **Response Parsing:** Parse LLM response into events
141+
- Tool calls → create `ActionEvent`(s)
142+
- Text message → create `MessageEvent` and return
143+
5. **Confirmation Check:** If actions need user approval:
144+
- Set conversation status to `WAITING_FOR_CONFIRMATION` and return
145+
6. **Action Execution:** Execute tools and create `ObservationEvent`(s)
146+
147+
**Key Characteristics:**
148+
- **Stateless:** Agent holds no mutable state between steps
149+
- **Event-Driven:** Reads from event history, writes new events
150+
- **Interruptible:** Each step is atomic and can be paused/resumed
151+
152+
## Agent Context
153+
154+
The agent applies `AgentContext` which includes **skills** and **prompts** to shape LLM behavior:
155+
156+
```mermaid
157+
%%{init: {"theme": "default", "flowchart": {"nodeSpacing": 30}} }%%
158+
flowchart LR
159+
Context["AgentContext"]
160+
161+
subgraph Skills["Skills"]
162+
Repo["repo<br><i>Always active</i>"]
163+
Knowledge["knowledge<br><i>Trigger-based</i>"]
164+
end
165+
SystemAug["System prompt prefix/suffix<br><i>Per-conversation</i>"]
166+
System["Prompt template<br><i>Per-conversation</i>"]
167+
168+
subgraph Application["Applied to LLM"]
169+
SysPrompt["System Prompt"]
170+
UserMsg["User Messages"]
171+
end
172+
173+
Context --> Skills
174+
Context --> SystemAug
175+
Repo --> SysPrompt
176+
Knowledge -.->|When triggered| UserMsg
177+
System --> SysPrompt
178+
SystemAug --> SysPrompt
179+
180+
style Context fill:#f3e8ff,stroke:#7c3aed,stroke-width:2px
181+
style Repo fill:#e8f3ff,stroke:#2b6cb0,stroke-width:2px
182+
style Knowledge fill:#fff4df,stroke:#b7791f,stroke-width:2px
183+
```
184+
185+
| Skill Type | Activation | Use Case |
186+
|------------|------------|----------|
187+
| **repo** | Always included | Project-specific context, conventions |
188+
| **knowledge** | Trigger words/patterns | Domain knowledge, special behaviors |
189+
190+
Review [this guide](/sdk/guides/skill) for details on creating and applying agent context and skills.
191+
192+
193+
## Tool Execution
194+
195+
Tools follow a **strict action-observation pattern**:
196+
197+
```mermaid
198+
%%{init: {"theme": "default", "flowchart": {"nodeSpacing": 30, "rankSpacing": 40}} }%%
199+
flowchart TB
200+
LLM["LLM generates tool_call"]
201+
Convert["Convert to ActionEvent"]
202+
203+
Decision{"Confirmation<br>mode?"}
204+
Defer["Store as pending"]
205+
206+
Execute["Execute tool"]
207+
Success{"Success?"}
208+
209+
Obs["ObservationEvent<br><i>with result</i>"]
210+
Error["ObservationEvent<br><i>with error</i>"]
211+
212+
LLM --> Convert
213+
Convert --> Decision
214+
215+
Decision -->|Yes| Defer
216+
Decision -->|No| Execute
217+
218+
Execute --> Success
219+
Success -->|Yes| Obs
220+
Success -->|No| Error
221+
222+
style Convert fill:#f3e8ff,stroke:#7c3aed,stroke-width:2px
223+
style Execute fill:#e8f3ff,stroke:#2b6cb0,stroke-width:2px
224+
style Decision fill:#fff4df,stroke:#b7791f,stroke-width:2px
225+
```
226+
227+
**Execution Modes:**
228+
229+
| Mode | Behavior | Use Case |
230+
|------|----------|----------|
231+
| **Direct** | Execute immediately | Development, trusted environments |
232+
| **Confirmation** | Store as pending, wait for user approval | High-risk actions, production |
233+
234+
**Security Integration:**
235+
236+
Before execution, the security analyzer evaluates each action:
237+
- **Low Risk:** Execute immediately
238+
- **Medium Risk:** Log warning, execute with monitoring
239+
- **High Risk:** Block execution, request user confirmation
240+
241+
## Component Relationships
242+
243+
### How Agent Interacts
244+
245+
```mermaid
246+
%%{init: {"theme": "default", "flowchart": {"nodeSpacing": 30}} }%%
247+
flowchart LR
248+
Agent["Agent"]
249+
Conv["Conversation"]
250+
LLM["LLM"]
251+
Tools["Tools"]
252+
Context["AgentContext"]
253+
254+
Conv -->|.step calls| Agent
255+
Agent -->|Reads events| Conv
256+
Agent -->|Query| LLM
257+
Agent -->|Execute| Tools
258+
Context -.->|Skills and Context| Agent
259+
Agent -.->|New events| Conv
260+
261+
style Agent fill:#f3e8ff,stroke:#7c3aed,stroke-width:2px
262+
style Conv fill:#e8f3ff,stroke:#2b6cb0,stroke-width:2px
263+
style LLM fill:#fff4df,stroke:#b7791f,stroke-width:2px
264+
```
265+
266+
**Relationship Characteristics:**
267+
- **Conversation → Agent**: Orchestrates step execution, provides event history
268+
- **Agent → LLM**: Queries for next actions, receives tool calls or messages
269+
- **Agent → Tools**: Executes actions, receives observations
270+
- **AgentContext → Agent**: Injects skills and prompts into LLM queries
271+
272+
273+
## See Also
274+
275+
- **[Conversation Architecture](/sdk/arch/conversation)** - Agent orchestration and lifecycle
276+
- **[Tool System](/sdk/arch/tool-system)** - Tool definition and execution patterns
277+
- **[Events](/sdk/arch/events)** - Event types and structures
278+
- **[Skills & Context](/sdk/arch/context)** - Prompt engineering and context management
279+
- **[LLM](/sdk/arch/llm)** - Language model abstraction

0 commit comments

Comments
 (0)