The roadmap describes the current main branch, which may move ahead of published binaries. The reliability, recovery, and AgentQi Companion work listed as recently completed is available in v0.3.0 and later. Consult each future item for an explicit target release rather than assuming that main is already packaged.
-
Companion token protection and migration (released in v0.3.0): OS-backed token storage now automatically migrates legacy JSON/fallback credentials with read-back verification, preserves recoverable copies on failure, respects Remember token and plaintext opt-in, and uses atomic private file writes. See Companion token storage.
-
Run explanation and recovery view (released in v0.3.0): admin console, Dashboard, and Companion session details combine recorded run/goal state, goal notes, session-scoped pending approvals, checkpoints, and recent tool failure evidence with contextual recovery guidance. Existing timelines remain available for investigation; the view does not authorize or replay actions.
-
Real-run regression import and offline replay (released in v0.3.0): import a redacted, complete text exchange from a gateway trajectory export and run recorded provider/tool fixtures through
RuntimeScenarioRunner, with independent assertions and strict consumption checks. Includes an executable sample; see trajectory replay. -
Browser sessions invalidate after local operator account updates, deletion, or disablement.
-
Goal completion accepts completed tool work; model status updates run alone and blocked transitions require three observations. Resume resets continuation and blocker counters.
-
Goal state is persisted atomically under the configured memory storage directory and restored lazily after restart.
-
Startup recovery pages all runnable sessions by stable ID instead of stopping after its first batch.
-
Dashboard authentication shares the gateway login request contract; typed API failures surface HTTP errors.
-
The production native runtime uses the extracted checkpoint, tool-loop, model, and context services.
-
RuntimeScenarioRunnerexecutes an injected native or MAF runtime and evaluates emitted evidence instead of trusting a supplied trace. -
CLI insights, outbound URL safety validation, and anonymizable trajectory export are implemented. See the capability matrix for optional and experimental lanes.
-
Channel expansion: Discord (Gateway WebSocket + interaction webhook), Slack (Events API + slash commands), Signal (signald/signal-cli bridge) channel adapters with DM policy, allowlists, thread-to-session mapping, and signature validation.
-
Tool expansion (80+ native and optional surfaces): edit_file, apply_patch, message, x_search, memory_get, sessions_history, sessions_send, sessions_spawn, session_status, sessions_yield, agents_list, cron, gateway, profile_write.
-
Tool presets and groups: 4 new built-in presets (full, coding, messaging, minimal) and 7 built-in tool groups (group:runtime, group:fs, group:sessions, group:memory, group:web, group:automation, group:messaging).
-
Chat commands: /think (reasoning effort), /compact (history compaction), /verbose (tool call/token output).
-
Multi-agent routing: per-channel/sender routing with model override, route-scoped prompt instructions, tool presets, and tool allowlist restrictions.
-
Integrations: Tailscale Serve/Funnel, Gmail Pub/Sub event bridge, mDNS/Bonjour service discovery.
-
Plugin installer: built-in
openclaw plugins install/remove/list/searchfor npm/ClawHub packages. -
Security audit closure for plugin IPC hardening, plugin-root containment, browser cancellation recovery, strict session-cap admission, and session-lock disposal.
-
Admin/operator tooling:
- posture diagnostics
- approval policy simulation
- redacted incident export
-
Expanded observability for approval decisions, session evictions/cap rejects, browser cancellation resets, plugin bridge auth/restart behavior, and sandbox lease lifecycle.
-
Optional estimated token admission control.
-
Startup/runtime composition split into explicit service, channel, plugin, and runtime assembly stages.
-
Optional native Notion scratchpad integration with scoped read/write tools (
notion,notion_write), allowlists, and write approvals by default. -
Canvas and A2UI v1 visual workspace with session-scoped websocket command broker, webchat Canvas host, Companion native Canvas tab, A2UI v0.8 JSONL renderer, event feedback, snapshots, and public-bind hardening.
-
Voice memo transcription for inbound audio media with Gemini provider support, degraded fallback behavior, and audio-marker preservation.
-
Checkpoint and resume for long-running native runtime turns, with durable save points after completed tool batches and resume from the latest completed batch after interruption.
-
Background session execution: Channel-activated sessions (including WebChat WebSocket) continue running after the user disconnects, with self-requeue through
MessagePipeline, bounded batches, startup auto-recovery, lifecycle notifications, and identical nativeAgentRuntime/MafAgentRuntimebehavior. -
Review-first learning hardening: richer learning proposal provenance, duplicate suppression, risk levels, validation warnings, and rollback for managed skill drafts and learning-created automation drafts.
These are strong candidates for the next roadmap phases because they extend the current runtime, channel, and operator model without fighting the existing architecture.
- Mixture-of-agents execution
- Fan out a prompt to multiple providers and synthesize a final answer from their outputs.
- Expose this as an optional high-cost/high-confidence runtime mode or explicit tool.
- Keep it profile-driven so it can be limited to selected models and use cases.
-
Daytona execution backend
- Add a remote workspace backend with hibernation and resume support.
- Fit it into the existing
IExecutionBackendand process execution model rather than adding a separate tool path. - Useful for persistent remote development-style sandboxes.
-
Modal execution backend
- Add a serverless execution backend for short-lived compute-heavy tasks.
- Focus on one-shot and bounded process execution first.
- Treat GPU-enabled workloads as an optional extension once the base backend is stable.
The runtime already includes CLI insights, URL safety validation, and trajectory export. The next additions build on those capabilities:
- Durable action reconciliation (released in v0.3.0, opt-in): persisted dispatch journal, stable provider adapter keys, completed-result reuse, and blocking of unknown outcomes before replay. See durable actions. Provider-specific adapters and executor-bypassing jobs need individual integration.
- Expanded regression capture (released in v0.3.0): opt-in bounded automatic capture, structured failed/blocked tool replay, and offline multimodal URL-content verification. See trajectory replay.
- Full-instance backup and restore (released in v0.3.0): offline inventory plans capture durable state and secret-reference manifests, verify checksums, and restore into a new isolated directory with SQLite validation and no dispatch. See instance backup.
- Guided recovery controls (released in v0.3.0): permission-aware goal pause/resume and evidence-backed action reconciliation with revision, approval, and budget checks. Complements the run explanation view; see guided recovery.
These are worthwhile changes, but they can break existing deployments or require new configuration. Recommend implementing behind flags first, then enabling by default in a major release.
-
Require auth on loopback for control/admin surfaces
- Scope:
/ws,/v1/*,/allowlists/*,/tools/approve,/webhooks/* - Goal: reduce “local process / local browser” attack surface.
- Scope:
-
Default allowlist semantics to
strict- Current:
legacymakes empty allowlist behave as allow-all for some channels. - Target:
strictshould be the default for safer out-of-the-box behavior.
- Current:
-
Default Telegram webhook signature validation to
true- Requires
WebhookSecretToken/WebhookSecretTokenRefto be configured. - Improves default webhook authenticity guarantees.
- Requires
Goal: make it straightforward to run Microsoft.SemanticKernel code behind the OpenClaw gateway/runtime while keeping SK integration optional (so the core stays NativeAOT-friendly).
Principles:
- Ship SK support as a separate package (no SK dependency in the core runtime).
- Treat SK execution as "just another tool" so OpenClaw policies (auth, rate limits, tool approval, tracing) still govern it.
- Prefer stable SK surfaces (Kernel + Functions/Plugins) and avoid betting on planners in the first iterations.
- README section describing supported integration patterns (wrap SK as a tool; host SK behind the gateway).
- Add a new optional NuGet package (tentative):
OpenClaw.SemanticKernelAdapter. - Provide
IServiceCollectionextensions to register an SK-backed tool. - Define a small, explicit request/response contract:
- Identify SK function by
(plugin, function)or a single "entrypoint" function name. - Pass args as JSON object; return JSON result + optional text.
- Identify SK function by
- Add a working sample (recommended location):
samples/SemanticKernelInterop/- Demonstrate: OpenClaw tool call -> SK function -> result -> returned via
/v1/responses - Include OpenTelemetry correlation (same trace/span across gateway -> tool -> SK call).
- Demonstrate: OpenClaw tool call -> SK function -> result -> returned via
- Optional startup mapping:
- Load a configured set of SK plugins/functions and expose each as an OpenClaw tool.
- Preserve OpenClaw tool naming rules and add predictable name mapping (e.g.
sk.<plugin>.<function>).
- Enforce governance:
- Per-tool allow/deny lists and per-tool rate limits (in OpenClaw config, not inside SK).
- Explicit secrets boundary: SK connectors should use the same secret ref system (
env:, etc).
- If SK invocation supports streaming in your chosen integration surface:
- Surface streaming responses through OpenClaw without bypassing message/token accounting.
- Bridge OTEL activities:
- Tag tool spans with
sk.plugin,sk.function, duration, and error metadata. - Ensure errors propagate as structured tool failures (not raw exceptions).
- Tag tool spans with
- Document a supported/known-good configuration:
- Which SK features are compatible with trimming/AOT and which are not.
- Add sample trimming config / annotations if required (only in the adapter/sample).
Non-goals (initially):
- Re-implement Semantic Kernel planners inside OpenClaw.
- Promise "drop-in" compatibility for every SK connector/plugin without validation.