|
| 1 | +--- |
| 2 | +title: Agent |
| 3 | +description: High-level architecture of the reasoning-action loop |
| 4 | +--- |
| 5 | + |
| 6 | +The **Agent** component implements the core reasoning-action loop that drives autonomous task execution. It orchestrates LLM queries, tool execution, and context management through a stateless, event-driven architecture. |
| 7 | + |
| 8 | +**Source:** [`openhands-sdk/openhands/sdk/agent/`](https://github.com/All-Hands-AI/agent-sdk/tree/main/openhands-sdk/openhands/sdk/agent) |
| 9 | + |
| 10 | +## Core Responsibilities |
| 11 | + |
| 12 | +The Agent system has four primary responsibilities: |
| 13 | + |
| 14 | +1. **Reasoning-Action Loop** - Query LLM to generate next actions based on conversation history |
| 15 | +2. **Tool Orchestration** - Select and execute tools, handle results and errors |
| 16 | +3. **Context Management** - Apply [skills](/sdk/guides/skill), manage conversation history via [condensers](/sdk/guides/context-condenser) |
| 17 | +4. **Security Validation** - Analyze proposed actions for safety before execution via [security analyzer](/sdk/guides/security) |
| 18 | + |
| 19 | +## Architecture |
| 20 | + |
| 21 | +```mermaid |
| 22 | +%%{init: {"theme": "default", "flowchart": {"nodeSpacing": 20, "rankSpacing": 50}} }%% |
| 23 | +flowchart TB |
| 24 | + subgraph Input[" "] |
| 25 | + Events["Event History"] |
| 26 | + Context["Agent Context<br><i>Skills + Prompts</i>"] |
| 27 | + end |
| 28 | + |
| 29 | + subgraph Core["Agent Core"] |
| 30 | + Condense["Condenser<br><i>History compression</i>"] |
| 31 | + Reason["LLM Query<br><i>Generate actions</i>"] |
| 32 | + Security["Security Analyzer<br><i>Risk assessment</i>"] |
| 33 | + end |
| 34 | + |
| 35 | + subgraph Execution[" "] |
| 36 | + Tools["Tool Executor<br><i>Action → Observation</i>"] |
| 37 | + Results["Observation Events"] |
| 38 | + end |
| 39 | + |
| 40 | + Events --> Condense |
| 41 | + Context -.->|Skills| Reason |
| 42 | + Condense --> Reason |
| 43 | + Reason --> Security |
| 44 | + Security --> Tools |
| 45 | + Tools --> Results |
| 46 | + Results -.->|Feedback| Events |
| 47 | + |
| 48 | + classDef primary fill:#f3e8ff,stroke:#7c3aed,stroke-width:2px |
| 49 | + classDef secondary fill:#e8f3ff,stroke:#2b6cb0,stroke-width:2px |
| 50 | + classDef tertiary fill:#fff4df,stroke:#b7791f,stroke-width:2px |
| 51 | + |
| 52 | + class Reason primary |
| 53 | + class Condense,Security secondary |
| 54 | + class Tools tertiary |
| 55 | +``` |
| 56 | + |
| 57 | +### Key Components |
| 58 | + |
| 59 | +| Component | Purpose | Design | |
| 60 | +|-----------|---------|--------| |
| 61 | +| **[`Agent`](https://github.com/All-Hands-AI/agent-sdk/blob/main/openhands-sdk/openhands/sdk/agent/agent.py)** | Main implementation | Stateless reasoning-action loop executor | |
| 62 | +| **[`AgentBase`](https://github.com/All-Hands-AI/agent-sdk/blob/main/openhands-sdk/openhands/sdk/agent/base.py)** | Abstract base class | Defines agent interface and initialization | |
| 63 | +| **[`AgentContext`](https://github.com/All-Hands-AI/agent-sdk/blob/main/openhands-sdk/openhands/sdk/context/agent_context.py)** | Context container | Manages skills, prompts, and metadata | |
| 64 | +| **[`Condenser`](https://github.com/All-Hands-AI/agent-sdk/blob/main/openhands-sdk/openhands/sdk/context/condenser/)** | History compression | Reduces context when token limits approached | |
| 65 | +| **[`SecurityAnalyzer`](https://github.com/All-Hands-AI/agent-sdk/blob/main/openhands-sdk/openhands/sdk/security/)** | Safety validation | Evaluates action risk before execution | |
| 66 | + |
| 67 | +## Reasoning-Action Loop |
| 68 | + |
| 69 | +The agent operates through a **single-step execution model** where each `step()` call processes one reasoning cycle: |
| 70 | + |
| 71 | +```mermaid |
| 72 | +%%{init: {"theme": "default", "flowchart": {"nodeSpacing": 10, "rankSpacing": 10}} }%% |
| 73 | +flowchart TB |
| 74 | + Start["step() called"] |
| 75 | + Pending{"Pending<br>actions?"} |
| 76 | + ExecutePending["Execute pending actions"] |
| 77 | + |
| 78 | + HasCondenser{"Has<br>condenser?"} |
| 79 | + Condense["Call condenser.condense()"] |
| 80 | + CondenseResult{"Result<br>type?"} |
| 81 | + EmitCondensation["Emit Condensation event"] |
| 82 | + UseView["Use View events"] |
| 83 | + UseRaw["Use raw events"] |
| 84 | + |
| 85 | + Query["Query LLM with messages"] |
| 86 | + ContextExceeded{"Context<br>window<br>exceeded?"} |
| 87 | + EmitRequest["Emit CondensationRequest"] |
| 88 | + |
| 89 | + Parse{"Response<br>type?"} |
| 90 | + CreateActions["Create ActionEvents"] |
| 91 | + CreateMessage["Create MessageEvent"] |
| 92 | + |
| 93 | + Confirmation{"Need<br>confirmation?"} |
| 94 | + SetWaiting["Set WAITING_FOR_CONFIRMATION"] |
| 95 | + |
| 96 | + Execute["Execute actions"] |
| 97 | + Observe["Create ObservationEvents"] |
| 98 | + |
| 99 | + Return["Return"] |
| 100 | + |
| 101 | + Start --> Pending |
| 102 | + Pending -->|Yes| ExecutePending --> Return |
| 103 | + Pending -->|No| HasCondenser |
| 104 | + |
| 105 | + HasCondenser -->|Yes| Condense |
| 106 | + HasCondenser -->|No| UseRaw |
| 107 | + Condense --> CondenseResult |
| 108 | + CondenseResult -->|Condensation| EmitCondensation --> Return |
| 109 | + CondenseResult -->|View| UseView --> Query |
| 110 | + UseRaw --> Query |
| 111 | + |
| 112 | + Query --> ContextExceeded |
| 113 | + ContextExceeded -->|Yes| EmitRequest --> Return |
| 114 | + ContextExceeded -->|No| Parse |
| 115 | + |
| 116 | + Parse -->|Tool calls| CreateActions |
| 117 | + Parse -->|Message| CreateMessage --> Return |
| 118 | + |
| 119 | + CreateActions --> Confirmation |
| 120 | + Confirmation -->|Yes| SetWaiting --> Return |
| 121 | + Confirmation -->|No| Execute |
| 122 | + |
| 123 | + Execute --> Observe |
| 124 | + Observe --> Return |
| 125 | + |
| 126 | + style Query fill:#f3e8ff,stroke:#7c3aed,stroke-width:2px |
| 127 | + style Condense fill:#e8f3ff,stroke:#2b6cb0,stroke-width:2px |
| 128 | + style Confirmation fill:#fff4df,stroke:#b7791f,stroke-width:2px |
| 129 | +``` |
| 130 | + |
| 131 | +**Step Execution Flow:** |
| 132 | + |
| 133 | +1. **Pending Actions:** If actions awaiting confirmation exist, execute them and return |
| 134 | +2. **Condensation:** If condenser exists: |
| 135 | + - Call `condenser.condense()` with current event view |
| 136 | + - If returns `View`: use condensed events for LLM query (continue in same step) |
| 137 | + - If returns `Condensation`: emit event and return (will be processed next step) |
| 138 | +3. **LLM Query:** Query LLM with messages from event history |
| 139 | + - If context window exceeded: emit `CondensationRequest` and return |
| 140 | +4. **Response Parsing:** Parse LLM response into events |
| 141 | + - Tool calls → create `ActionEvent`(s) |
| 142 | + - Text message → create `MessageEvent` and return |
| 143 | +5. **Confirmation Check:** If actions need user approval: |
| 144 | + - Set conversation status to `WAITING_FOR_CONFIRMATION` and return |
| 145 | +6. **Action Execution:** Execute tools and create `ObservationEvent`(s) |
| 146 | + |
| 147 | +**Key Characteristics:** |
| 148 | +- **Stateless:** Agent holds no mutable state between steps |
| 149 | +- **Event-Driven:** Reads from event history, writes new events |
| 150 | +- **Interruptible:** Each step is atomic and can be paused/resumed |
| 151 | + |
| 152 | +## Agent Context |
| 153 | + |
| 154 | +The agent applies `AgentContext` which includes **skills** and **prompts** to shape LLM behavior: |
| 155 | + |
| 156 | +```mermaid |
| 157 | +%%{init: {"theme": "default", "flowchart": {"nodeSpacing": 30}} }%% |
| 158 | +flowchart LR |
| 159 | + Context["AgentContext"] |
| 160 | + |
| 161 | + subgraph Skills["Skills"] |
| 162 | + Repo["repo<br><i>Always active</i>"] |
| 163 | + Knowledge["knowledge<br><i>Trigger-based</i>"] |
| 164 | + end |
| 165 | + SystemAug["System prompt prefix/suffix<br><i>Per-conversation</i>"] |
| 166 | + System["Prompt template<br><i>Per-conversation</i>"] |
| 167 | + |
| 168 | + subgraph Application["Applied to LLM"] |
| 169 | + SysPrompt["System Prompt"] |
| 170 | + UserMsg["User Messages"] |
| 171 | + end |
| 172 | + |
| 173 | + Context --> Skills |
| 174 | + Context --> SystemAug |
| 175 | + Repo --> SysPrompt |
| 176 | + Knowledge -.->|When triggered| UserMsg |
| 177 | + System --> SysPrompt |
| 178 | + SystemAug --> SysPrompt |
| 179 | + |
| 180 | + style Context fill:#f3e8ff,stroke:#7c3aed,stroke-width:2px |
| 181 | + style Repo fill:#e8f3ff,stroke:#2b6cb0,stroke-width:2px |
| 182 | + style Knowledge fill:#fff4df,stroke:#b7791f,stroke-width:2px |
| 183 | +``` |
| 184 | + |
| 185 | +| Skill Type | Activation | Use Case | |
| 186 | +|------------|------------|----------| |
| 187 | +| **repo** | Always included | Project-specific context, conventions | |
| 188 | +| **knowledge** | Trigger words/patterns | Domain knowledge, special behaviors | |
| 189 | + |
| 190 | +Review [this guide](/sdk/guides/skill) for details on creating and applying agent context and skills. |
| 191 | + |
| 192 | + |
| 193 | +## Tool Execution |
| 194 | + |
| 195 | +Tools follow a **strict action-observation pattern**: |
| 196 | + |
| 197 | +```mermaid |
| 198 | +%%{init: {"theme": "default", "flowchart": {"nodeSpacing": 30, "rankSpacing": 40}} }%% |
| 199 | +flowchart TB |
| 200 | + LLM["LLM generates tool_call"] |
| 201 | + Convert["Convert to ActionEvent"] |
| 202 | + |
| 203 | + Decision{"Confirmation<br>mode?"} |
| 204 | + Defer["Store as pending"] |
| 205 | + |
| 206 | + Execute["Execute tool"] |
| 207 | + Success{"Success?"} |
| 208 | + |
| 209 | + Obs["ObservationEvent<br><i>with result</i>"] |
| 210 | + Error["ObservationEvent<br><i>with error</i>"] |
| 211 | + |
| 212 | + LLM --> Convert |
| 213 | + Convert --> Decision |
| 214 | + |
| 215 | + Decision -->|Yes| Defer |
| 216 | + Decision -->|No| Execute |
| 217 | + |
| 218 | + Execute --> Success |
| 219 | + Success -->|Yes| Obs |
| 220 | + Success -->|No| Error |
| 221 | + |
| 222 | + style Convert fill:#f3e8ff,stroke:#7c3aed,stroke-width:2px |
| 223 | + style Execute fill:#e8f3ff,stroke:#2b6cb0,stroke-width:2px |
| 224 | + style Decision fill:#fff4df,stroke:#b7791f,stroke-width:2px |
| 225 | +``` |
| 226 | + |
| 227 | +**Execution Modes:** |
| 228 | + |
| 229 | +| Mode | Behavior | Use Case | |
| 230 | +|------|----------|----------| |
| 231 | +| **Direct** | Execute immediately | Development, trusted environments | |
| 232 | +| **Confirmation** | Store as pending, wait for user approval | High-risk actions, production | |
| 233 | + |
| 234 | +**Security Integration:** |
| 235 | + |
| 236 | +Before execution, the security analyzer evaluates each action: |
| 237 | +- **Low Risk:** Execute immediately |
| 238 | +- **Medium Risk:** Log warning, execute with monitoring |
| 239 | +- **High Risk:** Block execution, request user confirmation |
| 240 | + |
| 241 | +## Component Relationships |
| 242 | + |
| 243 | +### How Agent Interacts |
| 244 | + |
| 245 | +```mermaid |
| 246 | +%%{init: {"theme": "default", "flowchart": {"nodeSpacing": 30}} }%% |
| 247 | +flowchart LR |
| 248 | + Agent["Agent"] |
| 249 | + Conv["Conversation"] |
| 250 | + LLM["LLM"] |
| 251 | + Tools["Tools"] |
| 252 | + Context["AgentContext"] |
| 253 | + |
| 254 | + Conv -->|.step calls| Agent |
| 255 | + Agent -->|Reads events| Conv |
| 256 | + Agent -->|Query| LLM |
| 257 | + Agent -->|Execute| Tools |
| 258 | + Context -.->|Skills and Context| Agent |
| 259 | + Agent -.->|New events| Conv |
| 260 | + |
| 261 | + style Agent fill:#f3e8ff,stroke:#7c3aed,stroke-width:2px |
| 262 | + style Conv fill:#e8f3ff,stroke:#2b6cb0,stroke-width:2px |
| 263 | + style LLM fill:#fff4df,stroke:#b7791f,stroke-width:2px |
| 264 | +``` |
| 265 | + |
| 266 | +**Relationship Characteristics:** |
| 267 | +- **Conversation → Agent**: Orchestrates step execution, provides event history |
| 268 | +- **Agent → LLM**: Queries for next actions, receives tool calls or messages |
| 269 | +- **Agent → Tools**: Executes actions, receives observations |
| 270 | +- **AgentContext → Agent**: Injects skills and prompts into LLM queries |
| 271 | + |
| 272 | + |
| 273 | +## See Also |
| 274 | + |
| 275 | +- **[Conversation Architecture](/sdk/arch/conversation)** - Agent orchestration and lifecycle |
| 276 | +- **[Tool System](/sdk/arch/tool-system)** - Tool definition and execution patterns |
| 277 | +- **[Events](/sdk/arch/events)** - Event types and structures |
| 278 | +- **[Skills & Context](/sdk/arch/context)** - Prompt engineering and context management |
| 279 | +- **[LLM](/sdk/arch/llm)** - Language model abstraction |
0 commit comments