Skip to content

Latest commit

 

History

History
139 lines (107 loc) · 11.3 KB

File metadata and controls

139 lines (107 loc) · 11.3 KB

Roadmap

Release Availability

The roadmap describes the current main branch, which may move ahead of published binaries. The reliability, recovery, and AgentQi Companion work listed as recently completed is available in v0.3.0 and later. Consult each future item for an explicit target release rather than assuming that main is already packaged.

Recently Completed

  • Companion token protection and migration (released in v0.3.0): OS-backed token storage now automatically migrates legacy JSON/fallback credentials with read-back verification, preserves recoverable copies on failure, respects Remember token and plaintext opt-in, and uses atomic private file writes. See Companion token storage.

  • Run explanation and recovery view (released in v0.3.0): admin console, Dashboard, and Companion session details combine recorded run/goal state, goal notes, session-scoped pending approvals, checkpoints, and recent tool failure evidence with contextual recovery guidance. Existing timelines remain available for investigation; the view does not authorize or replay actions.

  • Real-run regression import and offline replay (released in v0.3.0): import a redacted, complete text exchange from a gateway trajectory export and run recorded provider/tool fixtures through RuntimeScenarioRunner, with independent assertions and strict consumption checks. Includes an executable sample; see trajectory replay.

  • Browser sessions invalidate after local operator account updates, deletion, or disablement.

  • Goal completion accepts completed tool work; model status updates run alone and blocked transitions require three observations. Resume resets continuation and blocker counters.

  • Goal state is persisted atomically under the configured memory storage directory and restored lazily after restart.

  • Startup recovery pages all runnable sessions by stable ID instead of stopping after its first batch.

  • Dashboard authentication shares the gateway login request contract; typed API failures surface HTTP errors.

  • The production native runtime uses the extracted checkpoint, tool-loop, model, and context services.

  • RuntimeScenarioRunner executes an injected native or MAF runtime and evaluates emitted evidence instead of trusting a supplied trace.

  • CLI insights, outbound URL safety validation, and anonymizable trajectory export are implemented. See the capability matrix for optional and experimental lanes.

  • Channel expansion: Discord (Gateway WebSocket + interaction webhook), Slack (Events API + slash commands), Signal (signald/signal-cli bridge) channel adapters with DM policy, allowlists, thread-to-session mapping, and signature validation.

  • Tool expansion (80+ native and optional surfaces): edit_file, apply_patch, message, x_search, memory_get, sessions_history, sessions_send, sessions_spawn, session_status, sessions_yield, agents_list, cron, gateway, profile_write.

  • Tool presets and groups: 4 new built-in presets (full, coding, messaging, minimal) and 7 built-in tool groups (group:runtime, group:fs, group:sessions, group:memory, group:web, group:automation, group:messaging).

  • Chat commands: /think (reasoning effort), /compact (history compaction), /verbose (tool call/token output).

  • Multi-agent routing: per-channel/sender routing with model override, route-scoped prompt instructions, tool presets, and tool allowlist restrictions.

  • Integrations: Tailscale Serve/Funnel, Gmail Pub/Sub event bridge, mDNS/Bonjour service discovery.

  • Plugin installer: built-in openclaw plugins install/remove/list/search for npm/ClawHub packages.

  • Security audit closure for plugin IPC hardening, plugin-root containment, browser cancellation recovery, strict session-cap admission, and session-lock disposal.

  • Admin/operator tooling:

    • posture diagnostics
    • approval policy simulation
    • redacted incident export
  • Expanded observability for approval decisions, session evictions/cap rejects, browser cancellation resets, plugin bridge auth/restart behavior, and sandbox lease lifecycle.

  • Optional estimated token admission control.

  • Startup/runtime composition split into explicit service, channel, plugin, and runtime assembly stages.

  • Optional native Notion scratchpad integration with scoped read/write tools (notion, notion_write), allowlists, and write approvals by default.

  • Canvas and A2UI v1 visual workspace with session-scoped websocket command broker, webchat Canvas host, Companion native Canvas tab, A2UI v0.8 JSONL renderer, event feedback, snapshots, and public-bind hardening.

  • Voice memo transcription for inbound audio media with Gemini provider support, degraded fallback behavior, and audio-marker preservation.

  • Checkpoint and resume for long-running native runtime turns, with durable save points after completed tool batches and resume from the latest completed batch after interruption.

  • Background session execution: Channel-activated sessions (including WebChat WebSocket) continue running after the user disconnects, with self-requeue through MessagePipeline, bounded batches, startup auto-recovery, lifecycle notifications, and identical native AgentRuntime/MafAgentRuntime behavior.

  • Review-first learning hardening: richer learning proposal provenance, duplicate suppression, risk levels, validation warnings, and rollback for managed skill drafts and learning-created automation drafts.

Runtime and Platform Expansion

These are strong candidates for the next roadmap phases because they extend the current runtime, channel, and operator model without fighting the existing architecture.

Multimodal and Input Expansion

  1. Mixture-of-agents execution
    • Fan out a prompt to multiple providers and synthesize a final answer from their outputs.
    • Expose this as an optional high-cost/high-confidence runtime mode or explicit tool.
    • Keep it profile-driven so it can be limited to selected models and use cases.

Execution and Deployment Options

  1. Daytona execution backend

    • Add a remote workspace backend with hibernation and resume support.
    • Fit it into the existing IExecutionBackend and process execution model rather than adding a separate tool path.
    • Useful for persistent remote development-style sandboxes.
  2. Modal execution backend

    • Add a serverless execution backend for short-lived compute-heavy tasks.
    • Focus on one-shot and bounded process execution first.
    • Treat GPU-enabled workloads as an optional extension once the base backend is stable.

Reliability and Operator Value

The runtime already includes CLI insights, URL safety validation, and trajectory export. The next additions build on those capabilities:

  1. Durable action reconciliation (released in v0.3.0, opt-in): persisted dispatch journal, stable provider adapter keys, completed-result reuse, and blocking of unknown outcomes before replay. See durable actions. Provider-specific adapters and executor-bypassing jobs need individual integration.
  2. Expanded regression capture (released in v0.3.0): opt-in bounded automatic capture, structured failed/blocked tool replay, and offline multimodal URL-content verification. See trajectory replay.
  3. Full-instance backup and restore (released in v0.3.0): offline inventory plans capture durable state and secret-reference manifests, verify checksums, and restore into a new isolated directory with SQLite validation and no dispatch. See instance backup.
  4. Guided recovery controls (released in v0.3.0): permission-aware goal pause/resume and evidence-backed action reconciliation with revision, approval, and budget checks. Complements the run explanation view; see guided recovery.

Security Hardening (Likely Breaking)

These are worthwhile changes, but they can break existing deployments or require new configuration. Recommend implementing behind flags first, then enabling by default in a major release.

  1. Require auth on loopback for control/admin surfaces

    • Scope: /ws, /v1/*, /allowlists/*, /tools/approve, /webhooks/*
    • Goal: reduce “local process / local browser” attack surface.
  2. Default allowlist semantics to strict

    • Current: legacy makes empty allowlist behave as allow-all for some channels.
    • Target: strict should be the default for safer out-of-the-box behavior.
  3. Default Telegram webhook signature validation to true

    • Requires WebhookSecretToken/WebhookSecretTokenRef to be configured.
    • Improves default webhook authenticity guarantees.

Semantic Kernel Interop (Non-Breaking, Optional)

Goal: make it straightforward to run Microsoft.SemanticKernel code behind the OpenClaw gateway/runtime while keeping SK integration optional (so the core stays NativeAOT-friendly).

Principles:

  1. Ship SK support as a separate package (no SK dependency in the core runtime).
  2. Treat SK execution as "just another tool" so OpenClaw policies (auth, rate limits, tool approval, tracing) still govern it.
  3. Prefer stable SK surfaces (Kernel + Functions/Plugins) and avoid betting on planners in the first iterations.

Phase 0 (Done): Documentation

  • README section describing supported integration patterns (wrap SK as a tool; host SK behind the gateway).

Phase 1 (Done): Minimal Adapter Package + Sample (High ROI)

  • Add a new optional NuGet package (tentative): OpenClaw.SemanticKernelAdapter.
  • Provide IServiceCollection extensions to register an SK-backed tool.
  • Define a small, explicit request/response contract:
    • Identify SK function by (plugin, function) or a single "entrypoint" function name.
    • Pass args as JSON object; return JSON result + optional text.
  • Add a working sample (recommended location): samples/SemanticKernelInterop/
    • Demonstrate: OpenClaw tool call -> SK function -> result -> returned via /v1/responses
    • Include OpenTelemetry correlation (same trace/span across gateway -> tool -> SK call).

Phase 2 (Done): "Load SK Plugins as Tools" (Selective Mapping)

  • Optional startup mapping:
    • Load a configured set of SK plugins/functions and expose each as an OpenClaw tool.
    • Preserve OpenClaw tool naming rules and add predictable name mapping (e.g. sk.<plugin>.<function>).
  • Enforce governance:
    • Per-tool allow/deny lists and per-tool rate limits (in OpenClaw config, not inside SK).
    • Explicit secrets boundary: SK connectors should use the same secret ref system (env:, etc).

Phase 3 (Done): Streaming + Observability Polish

  • If SK invocation supports streaming in your chosen integration surface:
    • Surface streaming responses through OpenClaw without bypassing message/token accounting.
  • Bridge OTEL activities:
    • Tag tool spans with sk.plugin, sk.function, duration, and error metadata.
    • Ensure errors propagate as structured tool failures (not raw exceptions).

Phase 4 (Done): NativeAOT/Trimming Guidance (Documentation + Constraints)

  • Document a supported/known-good configuration:
    • Which SK features are compatible with trimming/AOT and which are not.
  • Add sample trimming config / annotations if required (only in the adapter/sample).

Non-goals (initially):

  • Re-implement Semantic Kernel planners inside OpenClaw.
  • Promise "drop-in" compatibility for every SK connector/plugin without validation.