Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
56 changes: 56 additions & 0 deletions .pr/agent-reset-design.html

Large diffs are not rendered by default.

Binary file added .pr/evidence/canvas/01-notes-reset.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added .pr/evidence/canvas/02-history-recovered.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
56 changes: 56 additions & 0 deletions .pr/evidence/canvas/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,56 @@
# Canvas demonstration — scripted provider

This is a real Chrome browser using Agent Canvas and the B-source Agent Server.
The local HTTP provider uses `TestLLM` scripted responses. It verifies UI, tool,
event, and request integration; **it is not real-model behavior evidence**.

- SDK source: `20532dd3e5fb88e75e880ce07a6afbdff98df1b3` (1.53.0 development source).
- Canvas source: `a07364828c8f202e7745c6bce3dcef3915ae7ac1` (package 1.20.0).
- Final run: **1 passed**, 11.828 seconds, 2026-10-06 21:18 UTC.
- [Notes and reset](01-notes-reset.png), [history recovered](02-history-recovered.png),
[full recording](canvas-scripted-demo.webm) (242,159 bytes).
- [Verification summary](verification.json) records event IDs, forgotten IDs,
runtime versions, assertions, and artifact hashes.

The browser created a conversation from an agent profile configured with
`agent_reset`. A real `file_editor.create` wrote `retry-notes.md`; `new_context`
completed with its real result; a real `file_editor.view` reread the notes.
Both file observations were successful. The notes omitted the original
`Retry-After: 27` value. The next provider request retained the reset call/result
and omitted that original value. A browser follow-up then executed
`conversation_history.search` and `read`, retrieving the original event with
`in_active_view: false`; the final provider request and UI answer contained it.
The event log contained one voluntary reset and zero summary rescues.

The scripted sequence included a title response for Canvas's background title
request, followed by seven main-agent responses: create notes, reset, read notes,
first answer, history search, history read, final answer. Earlier trial runs
revealed this title request consuming a scripted turn; they are not the evidence
reported above. Final assertions checked actual file creation and successful
file-tool observations, rather than relying on the scripted answer's claims.

The isolated Canvas checkout ran `npm run dev:minimal` with
`OH_AGENT_SERVER_LOCAL_PATH` pointing to B, separate state/key paths, backend
port `18416`, frontend port `31416`, and `VITE_DO_NOT_TRACK=1`. The existing
`tests/e2e/mock-llm/scripts/mock-llm-server.py` ran with B's Python environment on
port `19416`. The browser command was:

```sh
MOCK_LLM_PORT=19416 \
MOCK_LLM_BACKEND_URL=http://localhost:31416 \
MOCK_LLM_SESSION_API_KEY=issue4916-local-demo \
npx playwright test --config .pr/issue4916-playwright.config.ts
```

The session key and provider key were demo-only dummy values. No real model
credentials, cookies, or external inference were used. The temporary Playwright
scenario/config and full raw request/event capture remain in the isolated Canvas
worktree under `.pr/` and `.agent_tmp/final/evidence/`; they are not proposed
repository changes. The original Canvas checkout and its existing `.DS_Store`
files were untouched.

This development stack omits automation and VS Code, so ancillary 404 toasts can
appear (visible at the edge of the screenshots). The notes/reset/
history tools themselves completed successfully. Canvas renders the new tools
through its existing generic `NEWCONTEXT` and `CONVERSATIONHISTORY` cards; this
does not claim a dedicated Canvas feature UI or full automation-stack validation.
Binary file added .pr/evidence/canvas/canvas-scripted-demo.webm
Binary file not shown.
78 changes: 78 additions & 0 deletions .pr/evidence/canvas/verification.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,78 @@
{
"scope": "Real Canvas UI + B-source Agent Server + scripted TestLLM HTTP provider. Not real-model behavior validation.",
"sdk_source_commit": "20532dd3e5fb88e75e880ce07a6afbdff98df1b3",
"sdk_source_version": "1.53.0 development source",
"canvas_source_commit": "a07364828c8f202e7745c6bce3dcef3915ae7ac1",
"canvas_package_version": "1.20.0",
"runtime": {
"os": "macOS-26.5.2-arm64-arm-64bit",
"node": "v24.21.0",
"chrome": "154.0.8037.98"
},
"test": {
"expected": 1,
"unexpected": 0,
"duration_seconds": 11.828,
"started_utc": "2026-10-06T21:18:19.929Z"
},
"conversation_id": "24df7261-4f97-4e29-8703-bf70adf31295",
"original_event_id": "c2c6396b-4be0-45d0-939b-dc0c3d533231",
"condensation_id": "e9061325-b6ff-4bf5-b81f-8b76e86057d2",
"forgotten_event_ids": [
"d10a2ab3-b76b-4e89-bef8-8c3376705d68",
"c2c6396b-4be0-45d0-939b-dc0c3d533231",
"899c7676-d020-4b4f-af1f-70f69341b7f1"
],
"voluntary_resets": 1,
"summary_rescues": 0,
"ordinary_notes": {
"filename": "retry-notes.md",
"create_succeeded": true,
"reread_succeeded": true,
"contains_original_value": false
},
"provider_requests": {
"total": 8,
"main_agent": 7,
"title_generation": 1,
"post_reset_request_contains_original_value": false,
"post_reset_request_retains_new_context": true,
"after_history_read_contains_original_value": true
},
"history_read": {
"source_event_id": "c2c6396b-4be0-45d0-939b-dc0c3d533231",
"in_active_view": false,
"is_error": false,
"recovered_original_value": true
},
"artifacts": {
"01-notes-reset.png": {
"bytes": 63677,
"sha256": "d3480b3b9cd594bc6b472f00101c2cf2788ca337871536489501716fcf6d63da"
},
"02-history-recovered.png": {
"bytes": 83338,
"sha256": "490e19efb51bdb3db4c6f0f17d7cd99950d65abd1c792cd5cd0f1bc440352cb2"
},
"canvas-scripted-demo.webm": {
"bytes": 242159,
"sha256": "4d07fa855e745678802822dfdfcadef01c3d9c0e8f956754a75404843e7e606d"
}
},
"limitations": [
"Scripted responses determine tool selection and final wording; no inference about real-model behavior.",
"Development UI stack omits automation and VS Code. Ancillary 404 toasts can occur; they do not represent the notes/reset/history tool results.",
"Canvas uses generic NEWCONTEXT and CONVERSATIONHISTORY cards; no dedicated Canvas renderer was added."
],
"process_cleanup": {
"launcher_exit_code": 0,
"provider_exit_code": 143,
"closed_ports": [
18416,
18417,
19416,
31416
],
"original_canvas_checkout_status": "Only the same four pre-existing untracked .DS_Store files."
}
}
44 changes: 44 additions & 0 deletions .pr/validation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
# Validation for model-directed context resets

Source revision: `20532dd3e5fb88e75e880ce07a6afbdff98df1b3`.
Base main: `8ba966d1d`. macOS arm64, Python 3.13.15, SDK/Agent Server 1.53.0 development tree.
The source tree is identical to the tested pre-rebase revision `781da62aff238d3f9609c88337198c853e0ac658`; only commit ancestry changed.

## Local checks

| Check | Result |
| --- | --- |
| `make build` | Passed in both isolated SDK worktrees |
| Changed-file and source-commit pre-commit hooks | Passed, including pyright, lint, import and tool-registration checks |
| `uv run pytest tests/sdk -q -n 4` | 6,982 passed; 9 skipped; 11 xfailed; 1 non-strict xpassed (7,003 collected) |
| New Agent reset lifecycle/capacity tests | 30 passed, sync and async |
| History retrieval tests on history revision `7ba20ec00` | 18 passed |
| `uv run pytest tests/agent_server tests/cross -q -n 4` | 2,992 passed; 9 failed; 1 skipped (3,002 collected) |
| New settings → create → real local Agent Server REST/WebSocket reset test | Passed within the cross-runtime suite |
| Persisted settings compatibility script | 23 golden fixtures and 8 PyPI 1.53.0 baseline payloads passed |
| Public API compatibility script | No breaking changes against PyPI 1.53.0 in SDK, workspace or tools; ACP dependency check skipped without a base-ref environment variable |
| OpenAPI tests / quality checker | 8 passed; existing 62 allowlisted cases unchanged |
| TypeScript coverage suite | 384 passed; final additive type refinement additionally passed focused tests and strict compilation |
| TypeScript build, lint, format, public type budget | Passed (lint retains 7 existing warnings) |
| Both offline examples | Passed; reset example executes ordinary file writes/reads and history retrieval, with zero-cost scripted model responses |
| Real Canvas UI with scripted HTTP provider | 1 passed, 11.828 seconds; five successful tool results, one reset, zero summaries; [recording and assertions](evidence/canvas/README.md) |
| Real-model behavior test `b06_agent_reset_history` | Collector/import and hooks passed; real inference not run without configured model credentials |

## Broad server-suite failures reproduced on untouched main

An independent worktree at `8ba966d1d`, with its own `make build`, reproduced all nine failing node IDs (9 failed, 2 passed in the targeted baseline run):

- Seven tests in `tests/agent_server/canvas_extensions/test_canvas_extension_backend.py`: Linux backend fixtures are unsupported on this macOS host.
- `test_cleanup_stale_tmux_sessions_includes_isolated_sockets`: the tmux executable is absent.
- `test_server_info_reports_configured_conversation_runtime`: the Docker executable is absent.

These results are not a completely green server CI matrix. They establish that the same failures occur without the feature. Unrelated production code and tests were not changed to hide them.

## Behavior and evidence boundaries

The tests assert actual model inputs, tool outcomes, event persistence/reopen/fork, and REST/WebSocket round trips. They cover asynchronous input arriving during generation, full tool batches, Finish precedence, approval rejection, interrupt after a committed tool result, request/condensation persistence gaps, unknown model capacity, protected handoff overflow, and a bounded outer summary rescue.

The scripted examples and Canvas demonstration establish harness behavior. They do not establish how reliably a real model decides when to reset, what to preserve in handoff, or when to search history. The live behavior test reports voluntary resets separately from summary rescues and accepts either sufficient search snippets or explicit reads.

Remote CI, model-backed behavioral evaluation, and maintainer acceptance of the adjusted issue design remain separate requirements. This work does not claim automatic merge readiness.

61 changes: 61 additions & 0 deletions clients/typescript/src/__tests__/agent-reset.test.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
import { Agent } from '../agent/agent';
import { ConversationClient, type CreateConversationPayload } from '../client/conversation-client';
import { SettingsClient } from '../client/settings-client';
import type { AgentResetCondenser, AgentResetCondenserSettings } from '../index';
import type { AgentServerSettingsPatchRequest } from '../models/agent-server-api';

it('forwards agent reset settings through settings updates and conversation creation', async () => {
const condenser: AgentResetCondenserSettings = {
condenser_kind: 'agent_reset',
enabled: true,
};
const fetchMock = vi.fn().mockImplementation(async () =>
Response.json({
agent_settings: { condenser },
conversation_settings: {},
llm_api_key_is_set: false,
})
);
vi.stubGlobal('fetch', fetchMock);
try {
const settings = new SettingsClient({ host: 'https://agent.test' });
const patch: AgentServerSettingsPatchRequest = {
agent_settings_diff: { condenser },
};
const updated = await settings.updateSettings(patch);
expect(updated.agent_settings.condenser).toEqual(condenser);
expect(fetchMock).toHaveBeenNthCalledWith(
1,
'https://agent.test/api/settings',
expect.objectContaining({ method: 'PATCH', body: JSON.stringify(patch) })
);

const client = new ConversationClient({ host: 'https://agent.test' });
const payload: CreateConversationPayload = {
agent_settings: { agent_kind: 'openhands', condenser, tools: [] },
};
await client.createConversation(payload);
expect(fetchMock).toHaveBeenNthCalledWith(
2,
'https://agent.test/api/conversations',
expect.objectContaining({ method: 'POST', body: JSON.stringify(payload) })
);

const runtimeCondenser: AgentResetCondenser = { kind: 'AgentResetCondenser' };
const direct: CreateConversationPayload = {
agent: new Agent({
llm: { model: 'test-model' },
condenser: runtimeCondenser,
include_default_tools: ['NewContextTool', 'ConversationHistoryTool'],
}),
};
await client.createConversation(direct);
expect(fetchMock).toHaveBeenNthCalledWith(
3,
'https://agent.test/api/conversations',
expect.objectContaining({ method: 'POST', body: JSON.stringify(direct) })
);
} finally {
vi.unstubAllGlobals();
}
});
45 changes: 45 additions & 0 deletions clients/typescript/src/__tests__/websocket-callback-client.test.ts
Original file line number Diff line number Diff line change
@@ -1,4 +1,6 @@
import type { WebSocketCallbackClient } from '../events/websocket-client';
import type { ConversationEvent } from '../events/types';
import type { NewContextAction, NewContextObservation } from '../models/agent-reset';

class Socket {
static instances: Socket[] = [];
Expand Down Expand Up @@ -66,4 +68,47 @@ describe('typed conversation event transport', () => {
expect(Socket.instances[1].close).toHaveBeenCalledTimes(1);
expect(vi.getTimerCount()).toBe(0);
});

it('delivers reset handoffs, input boundaries and requests without dropping fields', async () => {
const callback = vi.fn();
const { WebSocketCallbackClient } = await import('../events/websocket-client');
client = new WebSocketCallbackClient({
host: 'https://agent.test',
conversationId: 'conversation',
callback,
});
client.start();
const action = {
kind: 'NewContextAction',
handoff: 'Continue with the deployment check.',
} satisfies NewContextAction;
const observation: NewContextObservation = {
kind: 'NewContextObservation',
content: [{ type: 'text', text: 'Context reset requested.' }],
input_event_id: 'user-1',
};
const events: ConversationEvent[] = [
{
kind: 'ActionEvent',
id: 'action-1',
tool_name: 'new_context',
tool_call_id: 'call-1',
action,
},
{
kind: 'ObservationEvent',
action_id: 'action-1',
tool_name: 'new_context',
tool_call_id: 'call-1',
observation,
},
{ kind: 'CondensationRequest', trigger_action_id: 'action-1' },
{ kind: 'ContextWindowReminderEvent', source: 'environment' },
];

for (const event of events) {
Socket.instances[0].onmessage?.({ data: JSON.stringify(event) } as MessageEvent);
}
expect(callback.mock.calls.map(([event]) => event)).toEqual(events);
});
});
14 changes: 12 additions & 2 deletions clients/typescript/src/events/types.ts
Original file line number Diff line number Diff line change
Expand Up @@ -49,8 +49,14 @@ export type ActionEvent = Omit<
};
export type AgentErrorEvent = AgentServerAgentErrorEvent;
export type CondensationEvent = AgentServerCondensationEvent;
export type CondensationRequestEvent = AgentServerCondensationRequestEvent;
export type CondensationRequestEvent = AgentServerCondensationRequestEvent & {
/** The successful new_context action that requested this reset. */
trigger_action_id?: EventID | null;
};
export type CondensationSummaryEvent = AgentServerCondensationSummaryEvent;
export type ContextWindowReminderEvent = Omit<AgentServerCondensationRequestEvent, 'kind'> & {
kind: 'ContextWindowReminderEvent';
};
export type ConversationErrorEvent = Omit<AgentServerConversationErrorEvent, 'source'> & {
source?: EventSource;
};
Expand Down Expand Up @@ -169,7 +175,11 @@ export interface ThinkEvent extends BaseEvent {
* Union type of all conversation events
*/
export type ConversationEvent =
| AgentServerEvent
| Exclude<AgentServerEvent, { kind: 'CondensationRequest' }>
| CondensationRequestEvent
| ContextWindowReminderEvent
| ActionEvent
| ObservationEvent
| ConfirmationRequestEvent
| ConfirmationResponseEvent
| StuckDetectionEvent
Expand Down
10 changes: 10 additions & 0 deletions clients/typescript/src/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -87,6 +87,7 @@ export type {
PauseEvent,
CondensationRequestEvent,
CondensationSummaryEvent,
ContextWindowReminderEvent,
CondensationEvent,
ConversationStateUpdateEvent,
ConversationErrorEvent,
Expand All @@ -103,6 +104,15 @@ export type {
ConversationEvent,
} from './events/types';

export type {
AgentResetCondenserSettings,
AgentResetCondenser,
NewContextAction,
NewContextObservation,
ConversationHistoryAction,
ConversationHistoryObservation,
} from './models/agent-reset';

// Hooks
export {
HookEventType,
Expand Down
35 changes: 35 additions & 0 deletions clients/typescript/src/models/agent-reset.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
import type { ThinkObservation } from '../generated/agent-server-schema';

/** Requires an Agent Server with model-requested context reset support. */
export interface AgentResetCondenserSettings {
condenser_kind: 'agent_reset';
enabled?: boolean;
}

/** Runtime condenser configuration for a direct Agent specification. */
export interface AgentResetCondenser {
kind: 'AgentResetCondenser';
}

export interface NewContextAction {
kind?: 'NewContextAction';
handoff?: string;
}

export type NewContextObservation = Omit<ThinkObservation, 'kind'> & {
kind: 'NewContextObservation';
input_event_id?: string | null;
};

export interface ConversationHistoryAction {
kind?: 'ConversationHistoryAction';
command: 'search' | 'read';
query?: string | null;
event_id?: string | null;
before_event_id?: string | null;
offset?: number;
}

export type ConversationHistoryObservation = Omit<ThinkObservation, 'kind'> & {
kind: 'ConversationHistoryObservation';
};
Loading
Loading