Skip to content

Latest commit

 

History

History
781 lines (591 loc) · 44.5 KB

File metadata and controls

781 lines (591 loc) · 44.5 KB

How Open-Inspect Works

Open-Inspect is a background coding agent system. Unlike interactive coding assistants where you watch the AI work in real-time, Open-Inspect runs sessions in the cloud independently of your connection. You send a prompt, optionally close your laptop, and check the results later.

This guide covers the core architecture, how sessions work, and what happens when you send a prompt. For deployment instructions, see GETTING_STARTED.md.


The Background Model

The key insight behind Open-Inspect is that coding sessions don't need your constant attention.

Traditional coding assistants require you to stay connected:

You type → AI responds → You watch → You respond → Repeat

Open-Inspect decouples your presence from the work:

You send prompt → Session runs in background → You check results when ready

This enables workflows that aren't possible with interactive tools:

  • Fire and forget: Notice a bug before bed, kick off a session, review the PR in the morning
  • Parallel sessions: Run multiple approaches simultaneously without tying up your machine
  • Multiplayer: Share a session URL with a colleague and collaborate in real-time
  • Unlimited concurrency: Your laptop isn't the bottleneck—spin up as many sessions as you need

Sessions

A session is the core unit of work in Open-Inspect. Each session is:

  • Tied to a workspace: The agent works in clones of the repositories you selected — a single repository, an ad-hoc set of up to 10, a saved environment, or no repository at all
  • Persistent: State survives across connections—close the browser, come back later
  • Multiplayer: Multiple users can join, send prompts, and see events in real-time
  • Stateful: Contains messages, events, artifacts, and sandbox state

Session Targets

When creating a session from the web picker you choose what the sandbox works on:

Target What you get
A single repository Today's classic flow: one clone, one branch selector
Multiple repositories An ad-hoc ordered set (up to 10) cloned side by side
An environment A saved, reusable repository set — with optional prebuilt images and secrets
No repository An empty sandbox for scratch work

In multi-repository sessions each repository is cloned into its own directory under /workspace (named after the repository), and the first repository is the primary — it drives defaults like which settings apply. The agent sees all clones side by side and can make coordinated changes across them; pushes and pull requests are per-repository, so one session can produce PRs in several repositories. The session sidebar lists every repository with its branch and any PR created for it.

GitHub-bot sessions open the webhook's repository, unless that repository's metadata names a default environment (defaultEnvironmentId via the repo-metadata API) — then a PR review or @mention opens that environment's full workspace, provided the environment still contains the trigger repository. Slack sessions can target an environment three ways: a routing rule (Settings › Integrations › Slack) launches it from a keyword; a channel association (channelAssociations on the environments API, like repository metadata) routes messages in that channel to it automatically; and the LLM classifier considers environments alongside repositories, using their names and descriptions as signals. The classifier can also select no repository when the task does not require a codebase. Its clarification picker always includes No repository alongside accessible repositories and environments. Linear sessions can target an environment through the team and project mappings ({"environmentId": "env_…"} entries alongside repository entries).

Session Lifecycle

Created → Active → Archived
            ↑
            └── Can be restored from archive

Sessions start when you create one (via web or Slack). They remain active as long as there's work happening or recent activity. You can archive sessions to clean up your list, and restore them later if needed.

What's Stored in a Session

Data Description
Messages Prompts you've sent and their metadata
Prompt attachments Images uploaded with web or Slack prompts
Events Tool calls, token streams, status updates
Artifacts PRs created, screenshots captured
Participants Users who have joined the session
Sandbox state Reference to the current sandbox and its snapshot

Each session gets its own SQLite database in a Cloudflare Durable Object, ensuring isolation and high performance even with hundreds of concurrent sessions.


Environments

An environment is a named, reusable set of repositories — the thing you reach for when the same multi-repository workspace comes up again and again (a frontend + its API, a service + its shared library). Environments are managed under Settings > Environments and appear at the top of the new-session picker when accessible. They can be workspace-owned or team-owned, with names unique within that ownership scope. Ownership is immutable. Team pages also list their own environments.

Team environments require membership in the owning team or a workspace Owner/Administrator role, plus the appropriate environment permission. Managing one requires a lead or workspace Owner/Administrator and environments.manage; secrets, settings, and image mutations also require their respective feature permissions. A team environment can launch only into a session owned by that team. A team's session catalog can include eligible workspace environments too, provided its repository grants cover their members. Unbound bot catalogs do not expose team environments. requireTeamOnCreate, when enabled, also requires a team for new environment definitions. See Authentication and Authorization.

An environment defines:

  • An ordered repository list (up to 10) with a base branch per repository; the first repository is the primary
  • Environment secrets — sessions receive global secrets, then their owning team's secrets (if any), then the environment's secrets; later scopes win. Repository secrets do not flow in (see Secrets Management)
  • Optional prebuilt images — the whole environment (all clones + all setup scripts) is built ahead of time so sessions boot in seconds (see Pre-Built Images)

Sessions snapshot the environment at creation time: editing or deleting an environment never changes what an existing session works on (the session page shows "Environment deleted" if the source is gone). Ad-hoc "Multiple repositories" selections are the unsaved counterpart — same workspace shape, but no environment-scoped secrets or prebuilds; the picker offers to save the set as an environment.


Architecture

Open-Inspect uses a three-tier architecture spanning multiple cloud providers:

┌─────────────────────────────────────────────────────────────────────────┐
│                              Clients                                     │
│                    ┌───────────┬───────────┐                            │
│                    │    Web    │   Slack   │                            │
│                    └─────┬─────┴─────┬─────┘                            │
│                          │           │                                   │
└──────────────────────────┼───────────┼───────────────────────────────────┘
                           │           │
                           ▼           ▼
┌─────────────────────────────────────────────────────────────────────────┐
│                    Control Plane (Cloudflare)                            │
│  ┌────────────────────────────────────────────────────────────────────┐ │
│  │                    Durable Objects (per session)                    │ │
│  │  ┌──────────┐  ┌───────────┐  ┌────────────┐  ┌────────────────┐  │ │
│  │  │  SQLite  │  │ WebSocket │  │   Event    │  │    Sandbox     │  │ │
│  │  │   State  │  │    Hub    │  │   Stream   │  │   Lifecycle    │  │ │
│  │  └──────────┘  └───────────┘  └────────────┘  └────────────────┘  │ │
│  └────────────────────────────────────────────────────────────────────┘ │
│  ┌────────────────────────────────────────────────────────────────────┐ │
│  │                   D1 Database (shared state)                        │ │
│  │           Sessions index, repo metadata, encrypted secrets          │ │
│  └────────────────────────────────────────────────────────────────────┘ │
└───────────────────────────────────┬─────────────────────────────────────┘
                                    │
                                    ▼
┌─────────────────────────────────────────────────────────────────────────┐
│                 Data Plane (Sandbox Backend)                              │
│  ┌────────────────────────────────────────────────────────────────────┐ │
│  │                        Session Sandbox                              │ │
│  │  ┌────────────┐    ┌────────────┐    ┌────────────┐               │ │
│  │  │ Supervisor │───▶│  OpenCode  │───▶│   Bridge   │───────────────┼─┼──▶ Control Plane
│  │  └────────────┘    └────────────┘    └────────────┘               │ │
│  │                           │                                        │ │
│  │                    Full Dev Environment                            │ │
│  │              (Node.js, Python, git, Playwright)                    │ │
│  └────────────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────────────┘

Control Plane (Cloudflare Workers)

The control plane is the coordinator. It doesn't execute code—it manages state and routes messages.

Responsibilities:

  • Session state management (SQLite in Durable Objects)
  • WebSocket connections for real-time streaming
  • Sandbox lifecycle orchestration (spawn, snapshot, restore)
  • GitHub integration (repo listing, PR creation)
  • Authentication and access control

Why Cloudflare? Durable Objects provide per-session isolation with SQLite storage. Each session gets its own lightweight database that can handle hundreds of events per second without impacting other sessions. The WebSocket Hibernation API keeps connections alive during idle periods without incurring compute costs.

Sandbox lifecycle state is authoritative across WebSocket reconnects. Losing the sandbox WebSocket does not stop the sandbox: the bridge reconnects while the control plane schedules a heartbeat check in case the process is actually gone. Explicit lifecycle paths such as inactivity and stale heartbeat persist stopped or stale before closing the connection, which prevents reconnection.

Data Plane (Sandbox Backends)

The data plane is where code actually runs. Each session gets an isolated sandbox with a full development environment.

What's in a sandbox:

  • Debian Linux with common dev tools
  • Node.js 24, Python 3.12, git, curl
  • Package managers: npm, pnpm, pip, uv
  • agent-browser CLI + headless Chrome (for browser automation)
  • OpenCode and the Claude Agent SDK (the coding agent harnesses)

Open-Inspect supports these sandbox backends:

  • Modal: near-instant startup plus filesystem snapshot restore
  • Daytona: persistent stop/start sandboxes via direct REST API calls
  • Vercel Sandboxes: filesystem snapshot restore and prebuilt-image builds via the Vercel Sandbox API
  • OpenComputer: template-based sandboxes with checkpoint-backed prebuilt-image builds via the OpenComputer REST API
  • E2B: template-based sandboxes with persistent pause/resume via direct E2B REST API calls

Prebuilt-image builds are supported on Modal, Vercel, and OpenComputer. Saved filesystem state can be restored on those same providers for session resumes; Daytona and E2B use persistent sandboxes instead. For Daytona and E2B, the control plane stops or pauses the sandbox on inactivity or stale heartbeat, then resumes that same sandbox later.

Clients

Clients are how users interact with sessions. The architecture is client-agnostic—any client that can make HTTP requests and maintain WebSocket connections can participate.

Current clients:

  • Web: Next.js app with real-time streaming, session management, and settings
  • Slack: Bot that responds to @mentions and direct messages, forwards supported image attachments, classifies repos, and posts results
  • GitHub: Bot that reviews PRs and responds to PR @mentions
  • Linear: Agent workflow that starts sessions from Linear issue activity

All clients see the same session state. Send a prompt from Slack or GitHub, watch the results on web. This works because state lives in the control plane, not the client.


The Sandbox Lifecycle

Understanding the sandbox lifecycle explains why Open-Inspect can be fast despite running in the cloud.

Fresh Start (No Snapshot)

When you create a session for a repo without an existing snapshot:

┌─────────┐   ┌──────────┐   ┌──────────┐   ┌──────────────┐   ┌──────────────┐
│ Sandbox │──▶│  Bridge  │──▶│ Git Sync │──▶│ Setup Script │──▶│ Start Script │──┐
│ Created │   │  Starts  │   │ (clone)  │   │  (optional)  │   │  (optional)  │  │
└─────────┘   └──────────┘   └──────────┘   └──────────────┘   └──────────────┘  │
               "starting"      "sync"           "setup"            "start"        │
                                          .openinspect/setup.sh  .openinspect/start.sh
  ┌───────────────────────────────────────────────────────────────────────────────┘
  │   ┌────────────────┐   ┌─────────────┐   ┌───────┐
  └──▶│ Managed Skills │──▶│ Agent Start │──▶│ Ready │
      └────────────────┘   └─────────────┘   └───────┘
            "skills"          "harness"
  1. Sandbox created: The selected backend creates a fresh sandbox from its base runtime
  2. Bridge starts: The runtime starts its bridge process ahead of the repository boot, and the bridge opens its WebSocket to the control plane and reports the starting phase. The supervisor does not wait for that handshake, so the first boot steps can begin a moment before the socket is up; in practice it connects a second or two into the boot. Once connected the sandbox is connecting: it sends a heartbeat every 30 seconds and reports every later phase as it starts and completes, so the control plane can tell a long boot from a dead one. A prompt sent during boot waits here.
  3. Git sync (sync): Clones your repository using brokered SCM credentials from the git credential helper
  4. Setup script (setup): Runs .openinspect/setup.sh for provisioning (if present). A non-zero exit is reported as a warning and the boot continues
  5. Start script (start): Runs .openinspect/start.sh for runtime startup (if present)
  6. Managed skills (skills): Installs the session's managed skills into the workspace
  7. Agent start (harness): The agent harness starts (the OpenCode server, or the Claude Agent staging). The bridge attaches to it and sends ready
  8. Ready: The runtime's ready event, not the connection, marks the sandbox ready. Prompts that were waiting dispatch now

For multi-repository sessions the three steps have different shapes. Git sync is a single phase covering the whole set — the repositories are cloned concurrently into /workspace/<repo-name> — so it names no individual repository. Setup then runs for every repository in position order, and start runs as a second pass in the same order. Each setup and start phase names the repository it is running for.

Watching a boot

The session header names the phase while it runs: "Cloning repository", "Running setup.sh", "Starting services", "Installing skills", "Starting agent". Multi-repository sessions add the repository, as in "Running setup.sh for acme/api". Between phases, its status popover can say what just finished, but it does not display a list of completed phases or their durations. When a script fails, the header's status popover says which phase failed and names its repository when available. Failure reports retain phase, repository, and error or warning metadata, but hook stdout and stderr are discarded rather than collected or shown. A fatal start.sh failure in the session's first repository ends the boot. A setup.sh failure, and a start.sh failure in a later repository, are tolerated instead: the boot continues and the phase completes carrying a warning.

How long a boot may take

There is no fixed limit on setup.sh or start.sh. Two bounds apply instead:

  • Connect watchdog (4 minutes): measured from sandbox creation until the bridge first connects. It covers the provider launching the container, not your scripts.
  • Boot budget (30 minutes by default): measured from sandbox creation until ready. Set SANDBOX_BOOT_TIMEOUT_MS to change it (see Getting Started). When the budget runs out the sandbox is failed and its credentials are revoked, the prompt that was waiting fails with the phase that was running, and the next prompt starts a new sandbox.

A bridge that stops sending heartbeats for 90 seconds during boot is treated as dead the same way. No snapshot is taken in any of these cases, so a half-provisioned workspace never becomes the restore point. The sandbox itself is stopped only on providers that can stop one explicitly; on the others the row is marked stale and its socket detached, and the sandbox ages out on the provider's own timeout.

Restore (From Snapshot)

When restoring from a previous snapshot:

┌─────────────┐   ┌──────────┐   ┌────────────┐   ┌──────────────┐   ┌─────────────┐   ┌───────┐
│  Restore    │──▶│  Bridge  │──▶│ Quick Sync │──▶│ Start Script │──▶│ Agent Start │──▶│ Ready │
│  Snapshot   │   │  Starts  │   │(git fetch) │   │  (optional)  │   │             │   │       │
└─────────────┘   └──────────┘   └────────────┘   └──────────────┘   └─────────────┘   └───────┘
  1. Restore snapshot: The selected snapshot-capable provider restores the filesystem from a saved snapshot or checkpoint
  2. Bridge connects: As in a fresh start, the bridge connects first and reports each phase
  3. Quick sync: Fetches the session's branch from origin. The restored checkout is left as it is, so the snapshot's commit and any uncommitted work survive the restore
  4. Start script: Runs .openinspect/start.sh for runtime startup (if present)
  5. Agent start: Managed skills are installed and the agent harness starts
  6. Ready: Sandbox is ready almost instantly

A restore reports the same phases as a fresh start minus setup, so the header shows the same labels.

A snapshot taken by an older runtime still restores: the control plane accepts any runtime from generation 62 up, and connecting before the boot arrived in generation 68. A session restored from a generation 62–67 snapshot therefore boots the old way — its bridge connects only once the repository boot has finished — so it reports no phases while it boots and the whole boot still has to fit inside the four-minute connect watchdog.

Snapshots include installed dependencies, built artifacts, and workspace state. This is why follow-up prompts in an existing session are much faster than the first prompt.

Prebuilt Image Start

When starting from a pre-built image (built for the session's repository or, for sessions launched from a prebuild-enabled environment, the environment's whole repository set):

  1. Bridge connects: As in a fresh start, the bridge connects first and reports each phase
  2. Incremental git sync: Fast fetch + hard reset to latest branch head (per repository for multi-repository sets)
  3. Setup skipped: .openinspect/setup.sh already ran when the image was built, so no setup phase is reported
  4. Start script runs: .openinspect/start.sh executes for per-session runtime startup
  5. Managed skills (skills): The session's managed skills are installed
  6. Agent start (harness): The agent harness starts and the bridge attaches to it
  7. Ready: The runtime sends ready once the harness is up

If start.sh exists and fails, startup fails fast instead of continuing with a broken runtime.

Preinstalling local MCP dependencies

OpenInspect checks the global npm installation before preparing local npx MCP servers. An exact package version already installed with intact executable links is reused, including after a snapshot restore. Only missing or mismatched packages are installed. Failed or interrupted installs are marked for retry rather than treated as cache hits.

To remove the installation from first-session startup, pin the same package version in the MCP command and the repository's .openinspect/setup.sh (or the setup hook used by its environment):

# .openinspect/setup.sh — replace this example package/version with your MCP dependency
npm install --global @example/mcp-server@1.2.3

Configure the matching MCP command as ["npx", "-y", "@example/mcp-server@1.2.3"] and rebuild the prebuilt image. This uses the existing setup/prebuild lifecycle; MCP settings are not automatically baked into images. Changing a pinned version causes an install until the image is rebuilt with it.

Unversioned packages and tags such as latest are refreshed once per sandbox boot. Successful installs are reused across OpenCode process restarts within that boot, but not across snapshot restores. Remote MCP servers are unaffected, and server commands, arguments, and credentials are passed through unchanged. Unsupported npx option forms are left to npx without eager installation.

The mcp.package_cache event reports hit/miss counts and lookup time; mcp.packages_installed reports preparation time on misses. This optimization removes redundant global installation, not all MCP startup work: npx may still resolve registry metadata or populate its own execution cache, particularly with explicit --package commands.

When Snapshots Are Taken

  • After successful prompt completion: Preserves the workspace state
  • Before sandbox timeout: Saves state before the sandbox shuts down due to inactivity
  • On explicit save: Can be triggered by the control plane
  • Not on a failed boot: None of the automatic boot-failure paths — connect watchdog, boot budget, stale heartbeat, fatal runtime error — takes a snapshot, so a half-provisioned workspace does not become the restore point that way. The runtime also refuses a snapshot command while it is still booting

Sandbox Warming

To minimize perceived latency, sandboxes warm proactively:

  • When you start typing a prompt, the control plane begins warming a sandbox
  • By the time you hit enter, the sandbox may already be ready
  • If restore is fast enough, you won't notice any delay

Tunnel URLs Inside the Sandbox

When a session uses the tunnelPorts sandbox setting, the resolved tunnel URLs are written to /workspace/.tunnels.env so processes started by .openinspect/start.sh (or by the agent later) can read them locally.

# /workspace/.tunnels.env
TUNNEL_SANDBOX_ID=sandbox-acme-app-1783614336426
TUNNEL_3000=https://abc123-3000.modal.host
TUNNEL_5173=https://abc123-5173.modal.host

This dotenv shape works directly with tools that accept an env-file path — node --env-file=..., bun --env-file=..., docker compose --env-file=.... The format is plain KEY=value, so any other dotenv consumer can read it without parsing. The TUNNEL_SANDBOX_ID line names the sandbox the URLs were resolved for; the supervisor uses it to tell a fresh write from a snapshot leftover.

Boot ordering. On every non-build boot, the supervisor:

  1. Clears a file left by a previous sandbox (its TUNNEL_SANDBOX_ID doesn't match), such as one inherited from a snapshot. A file already written for this sandbox is kept — the backend's write can land before the supervisor starts.
  2. Waits up to TUNNEL_WAIT_TIMEOUT_SECONDS (default 30) for fresh URLs.
  3. Runs .openinspect/start.sh.

If the wait times out (for example, because the backend has not resolved tunnel URLs yet), start.sh proceeds without fresh local URLs and the supervisor logs tunnel.env_file_wait_timeout. The control plane still receives and broadcasts the URLs to clients on a separate path. The file is not written when tunnelPorts is empty or in build mode.


How Prompts Flow Through the System

Here's what happens when you send a prompt:

┌──────┐   ┌────────┐   ┌───────────────┐   ┌─────────┐   ┌──────────┐
│ User │──▶│ Client │──▶│ Control Plane │──▶│ Sandbox │──▶│ OpenCode │
└──────┘   └────────┘   └───────────────┘   └─────────┘   └──────────┘
              │                 │                              │
              │                 │         Events stream back   │
              │◀────────────────┼◀─────────────────────────────┘
              │                 │
              ▼                 ▼
         Display to        Broadcast to
           user           all clients

Step by Step

  1. You send a prompt via web or Slack

  2. Control plane queues it: The prompt goes to the session's Durable Object and is added to the message queue. If a sandbox isn't running, one is spawned or restored.

  3. Sandbox receives the prompt: Via WebSocket, the control plane sends the prompt to the sandbox along with author information (for commit attribution).

  4. OpenCode processes it: The agent reads files, makes edits, runs commands—whatever the task requires. Each action generates events.

  5. Events stream back: Tool calls, token streams, and status updates flow back through the WebSocket to the control plane.

  6. Control plane broadcasts: Events are stored in the session database and broadcast to all connected clients in real-time.

  7. Artifacts are created: If the agent creates a PR or captures a screenshot, these are stored as artifacts and announced to clients.

Prompt Queuing

If you send a prompt while the agent is still working on a previous one, it's queued:

Prompt 1 (processing) ──▶ Prompt 2 (queued) ──▶ Prompt 3 (queued)

This lets you send follow-up thoughts while the agent works. Prompts are processed in order.

You can also stop the current execution if the agent is going down the wrong path.

Parent-to-Child Follow-Ups

An agent that created a child with spawn-child can continue that same child session with send-child-prompt. The follow-up enters the child's normal durable queue:

Child prompt 1 (processing) ──▶ Parent follow-up (queued) ──▶ Child continues

The follow-up does not interrupt active work. Completed and failed children can resume, restoring their compatible sandbox snapshot when available. Cancelled children remain terminal, and archived children must be explicitly unarchived before they can accept prompts.

The parent token is never exchanged for the child's sandbox token. The control plane authenticates the parent session, verifies the direct parent-child relationship in D1, verifies it again in the child Durable Object, and attributes the queued prompt to the child owner with source agent.

send-child-prompt returns after the prompt is durably queued. The parent calls get-child-status when it needs the follow-up result. An earlier completed response is labeled as such while newer child work is still running.

The runtime tool is installed when a sandbox starts from a runtime image that includes it. A parent restored from a snapshot created before this capability shipped keeps the older captured runtime and will not see send-child-prompt until it starts in a fresh sandbox built from the newer runtime.


The Agent

The sandbox runtime speaks to its coding agent through one seam, the agent harness. A session runs on exactly one harness, chosen at create:

  • OpenCode (built-in): OpenCode runs as a server inside the sandbox; the supervisor owns the opencode serve process and the bridge talks to it over HTTP/SSE.
  • Claude Agent: the Claude Agent SDK runs inside the bridge and spawns the claude binary as its own child, launched with a clean environment that carries exactly one Anthropic credential. This is the harness that can use a connected Claude subscription. See Using the Claude Agent Harness.

Both harnesses emit the same session events (tokens, tool calls, steps, warnings), so everything above the sandbox is harness-neutral. The bridge owns turn completion: a harness reports the outcome of a turn and the bridge emits the single execution_complete event. Follow-up prompts queue until the running turn ends on both harnesses.

What the Agent Can Do

Capability Description
Read files Explore the codebase, understand context
Edit files Make changes, refactor code
Run commands Execute tests, builds, scripts
Git operations Commit changes, create branches
Web browsing Look up documentation, research errors
Visual verification Use Playwright to check UI changes

How Changes Are Attributed

When the agent makes commits, they're attributed to the user who sent the prompt:

Author: Jane Developer <jane@example.com>
Committer: Open-Inspect <bot@open-inspect.dev>

This ensures your contributions are properly credited in git history.

Creating Pull Requests

When you ask the agent to create a PR:

  1. Agent pushes the branch using brokered SCM credentials from the sandbox credential helper
  2. Control plane receives the branch name
  3. Control plane creates the PR using your GitHub OAuth token (GitHub logins)
  4. PR appears as created by you, not a bot

If you signed in another way (e.g. Google) you have no GitHub OAuth token, so the control plane pushes the branch with the shared GitHub App credentials and returns a manual pull/new URL — the PR is attributed to the App bot rather than to you.

This maintains proper code review workflows—you can't approve your own PRs.


Real-time Events

Sessions stream events to all connected clients via WebSocket.

Event Types

Event Description
sandbox_spawning Sandbox is being created
sandbox_status Sandbox moved between pending, spawning, connecting, warming, ready, snapshotting, stale, stopped, and failed
sandbox_event Tool call, token stream, boot phase (boot_progress), or other agent event
artifact_created PR created, screenshot captured
presence_update User joined or left the session
session_status Session state changed

Multiplayer

Multiple users can connect to the same session:

  • Presence: See who else is watching
  • Shared stream: Everyone sees the same events
  • Attributed prompts: Each prompt is tagged with who sent it
  • Collaborative: One person can start a task, another can refine it

This makes sessions useful for pair programming, live debugging, or teaching.


Snapshots and Performance

Speed is critical for background agents. If sessions are slow, people won't use them.

The Cold Start Problem

Without optimization, starting a session would require:

  1. Spinning up a container (~5-10s)
  2. Cloning the repository (~10-30s for large repos)
  3. Installing dependencies (~30s-5min)
  4. Starting the agent (~5s)

That's potentially minutes before the agent can start working. Because the runtime connects before it clones, you watch those steps happen phase by phase instead of waiting on a silent "Connecting" indicator.

How Snapshots Solve This

Provider snapshots and checkpoints let us capture a sandbox's state after setup:

First session:  Clone ─▶ Install/Build ─▶ Start Runtime ─▶ [Snapshot] ─▶ Work
                              (slow)

Later sessions: [Restore Snapshot] ─▶ Quick sync ─▶ Start Runtime ─▶ Work
                     (fast)

The first session for a repo pays the setup cost. Subsequent sessions restore in seconds when the active provider supports saved filesystem state.

For Vercel, Terraform builds a base-runtime snapshot from the local checkout and wires a deterministic snapshot name into VERCEL_BASE_SNAPSHOT_NAME. Fresh Vercel sandboxes resolve that name to the newest created snapshot instead of cloning and installing the sandbox runtime on every session. OpenComputer uses a managed template plus checkpoints for the same prebuilt-image lifecycle. See Vercel Sandbox Provider and OpenComputer Sandbox Provider for provider-specific details.

Image Prebuilding

For frequently-used repositories — and for environments — images can be prebuilt on a schedule:

  • Clone the repository (or every repository of the environment), install dependencies, run initial build
  • Save as a provider image artifact
  • Sessions start from this artifact, only syncing recent changes

This means even "cold" sessions (no previous snapshot) start from a recent baseline. See Pre-Built Images for details.


Security Model

Open-Inspect is designed for single-tenant deployment where all users are trusted members of the same organization.

Why Single-Tenant?

GitHub git operations use a shared GitHub App installation. This means:

  • The installation defines the workspace's maximum repository reach; workspace permissions and team repository grants further constrain access
  • Open-Inspect does not compare a user's personal GitHub repository permissions with that scope
  • The trust boundary is your organization, not individual users

This follows Ramp's original design, which was built for internal use where all employees have access to company repositories.

Token Architecture

Session sandboxes fetch git credentials through the control plane and cache them on disk. GitHub credentials cover the persisted session repositories, intersected with current owning-team grants only for team-owned sessions. GitLab returns the deployment PAT without per-session narrowing. Modal filesystem snapshots can retain the helper cache; brokerage is not a token-free snapshot guarantee, and grant removal does not immediately revoke issued credentials.

Image builds receive VCS_CLONE_TOKEN because they have no session broker. For GitHub it is scoped to the build repositories, with current owning-team grants applied for team-owned environment builds. For GitLab it is the deployment PAT, not a repository-scoped or single-use credential. See the canonical access and credential boundaries for dependency access, cache behavior, and provider limitations.

Secrets

You can configure environment variables (API keys, credentials) at global, team, repository, or environment scope. Precedence is global, then owning team when present, then the session target's secrets; later layers win collisions:

  • Global secrets apply to all sessions (e.g., ANTHROPIC_API_KEY, DEEPSEEK_API_KEY, ZHIPU_API_KEY, OPENCODE_API_KEY)
  • Team secrets apply to sessions owned by that team, overriding global values
  • Repository secrets apply to repository-targeted sessions and override global and team secrets with the same key; ad-hoc multi-repository sessions receive each selected repository's secrets, with the primary winning collisions
  • Environment secrets apply to sessions launched from that environment — its repositories' repository secrets do not flow in
  • Stored encrypted (AES-256-GCM) in D1 database
  • Injected into sandboxes at startup
  • Never exposed to clients (only key names are visible)

OpenAI and xAI subscription credentials are installation-wide provider accounts. Account rows store display, status, and optional external identity separately from credentials encrypted with PROVIDER_ACCOUNTS_ENCRYPTION_KEY. Each provider has an optional default account and an unattended mode that chooses the default account or API-key mode for Slack, GitHub, Linear, and unpinned automation runs.

The web groups OpenAI and xAI accounts from shared static provider IDs and display metadata; there is no provider-catalog endpoint. The control-plane adapter registry remains authoritative when an account is connected, selected, defaulted, or consumed.

Session creation resolves every subscription provider once and persists an immutable provider account, API-key, or legacy scoped-OAuth auth row in D1, the sole authority for session provider auth. The session Durable Object remains authoritative for lifecycle and sandbox-token authentication but does not replicate provider-account bindings. An interactive session can follow provider policy, select an active account, or choose API-key mode. Automations can pin the same choices or resolve current defaults each run. Child sessions copy their parent's D1 auth rows, and later default changes do not move existing sessions between accounts.

In account mode, the sandbox receives only a managed marker. Provider API keys and legacy OAuth fields for that provider are suppressed. The runtime plugin calls the sandbox-authenticated POST /sessions/:id/provider-auth/:provider/access-token endpoint; the control plane reads the trusted D1 session binding using the sandbox-authenticated session ID, refreshes the encrypted account credential, and returns short-lived access with Cache-Control: no-store. Sandbox startup also reads the complete D1 auth snapshot and fails closed if it is unavailable or incomplete.

Legacy scoped OAuth and provider accounts can coexist. Existing sessions remain pinned to legacy scoped OAuth. New sessions use an explicit choice, then a provider-account default, and otherwise retain legacy scoped OAuth or API-key behavior. Setting a default affects only future sessions; operators may remove legacy keys after legacy-bound sessions are no longer needed. See Using OpenAI Models and Using Grok with a SuperGrok Subscription.

LLM API keys (e.g., ANTHROPIC_API_KEY for Claude models) are added as global secrets. A deployment can instead configure anthropic_api_key in Terraform to inject one fleet-wide key into Modal session sandboxes and OpenComputer sandboxes; a global secret of the same name takes precedence over it, and the other providers read only the secret store.

Opt-in model providers: DeepSeek models require DEEPSEEK_API_KEY, Z.AI Coding Plan models require ZHIPU_API_KEY, and OpenCode Zen and OpenCode Go models require OPENCODE_API_KEY (Go also needs an active Go subscription on that key), as a global secret with any sandbox provider. SuperGrok models require an xAI provider account or XAI_API_KEY mode and must be enabled under Settings > Models.

See Secrets Management for setup instructions.

Deployment Recommendations

  1. Deploy behind SSO/VPN: Control who can access the web interface
  2. Limit GitHub App scope: Only install on repositories you want accessible
  3. Use "Select repositories": Don't give the App access to all org repos

What's Next