Skip to content

Latest commit

 

History

History
315 lines (237 loc) · 15.9 KB

File metadata and controls

315 lines (237 loc) · 15.9 KB

Clerum E2E Testing Guide

Overview

Clerum has two categories of tests:

  • Unit tests — per service, Vitest. Run with npm test inside each service directory.
  • E2E tests — full-stack on minikube with Calico CNI. Run the runtime gate with ./scripts/e2e/e2e-workflow-runtime-gate.sh; run backend compatibility suites with ./scripts/e2e/e2e-workflow-backend-compat.sh.

E2E tests validate the full pipeline (CRD reconcile → NetworkPolicy → MCP discovery → LLM tool-calling + approval flow) against a real cluster. Unit tests validate per-service logic in isolation.

Prerequisites

From root README.md §Testing prerequisites:

  • Docker Desktop running

  • minikube installed (brew install minikube)

  • kubectl configured

  • Node.js 24+ for unit tests

  • .env file at repo root with LLM API keys:

    Variable Required For How to Get
    ZAI_API_KEY LLM tool-calling https://z.ai
    OPENAI_API_KEY Alternative LLM provider https://platform.openai.com/api-keys
    CLAUDE_API_KEY Alternative LLM provider https://console.anthropic.com/
    CLERUM_MODEL_PROVIDER Provider selection zai, openai, claude, or bailian

Copy .env.example to .env and fill in your keys. .env is gitignored.

Cluster bootstrap

1. Create the minikube cluster

minikube start -p clerum-test \
  --driver=docker \
  --cpus=6 \
  --memory=10240 \
  --cni=calico

Important: ~10GB RAM allocated to minikube is recommended. The E2E suite runs 5 composite recipes concurrently deploying MongoDB, PostgreSQL, Redis, and multiple MCP servers. Calico CNI is required so NetworkPolicy Phase 6 actually enforces. Do not pin an older --kubernetes-version: deploy/base ships ValidatingAdmissionPolicy on admissionregistration.k8s.io/v1, which requires Kubernetes 1.30+. The minikube default is what scripts/minikube/start.sh uses.

2. Bootstrap infrastructure

The canonical setup is the Makefile path — it starts the clerum-test profile, installs CRDs, creates secrets/config, acquires images, deploys manifests, waits for rollouts, and seeds local test data (12 idempotent steps):

MINIKUBE_IMAGE_TAG=latest make minikube-setup
make minikube-status

Setup builds nothing by default: it pulls the 23 published service images from ghcr.io/evenfire-ai. MINIKUBE_IMAGE_TAG=latest is required until the next release is tagged and promoted, because the manifests pin that not-yet-cut tag. make minikube-setup-local (equivalently IMAGE_SOURCE=local) builds every image in minikube's Docker daemon instead and needs no override. Which mode a cluster is in is recorded in deploy/minikube/.image-manifest.json, and it decides which image refs the pods run — see ../deploy/minikube.md.

For the E2E fixture set (test user, e2e-* recipes, demo MCP servers), use make minikube-setup-e2e. In ghcr mode that pulls the published images and then builds the two unpublished E2E coordinator fixtures locally.

scripts/bootstrap-cluster.sh is an older, partial bootstrap (no full JWT chain, UIs, or user seed). E2E suites need the full setup above — do not use the script for E2E runs.

3. Sync a running cluster before a gate

Before any cluster-backed E2E gate, sync the running cluster to the current worktree:

make minikube-pre-gate-sync GATE=<gate-name>

Use --force-cluster-sync --skip-port-forwards when deployable code changed and your test runner will hold its own port-forwards. Inner pre-gate-sync may use --skip-port-forwards; never pass that globally into make minikube-t2.

make minikube-pre-gate-sync GATE=<gate-name> ARGS="--force-cluster-sync --skip-port-forwards"

After the deploy sync, verify a clean state before reading E2E results:

make minikube-status
kubectl --context=clerum-test get pods -A --field-selector=status.phase!=Running,status.phase!=Succeeded
kubectl --context=clerum-test -n sandbox-recipes get workflowrecipes
make minikube-verify-network-policy
CONTEXT=clerum-test scripts/minikube/seed-test-data.sh

4. Port-forwards and Vitest

make test-e2e-vitest and make test-e2e-all install tests/e2e dependencies when missing and keep minikube port-forwards alive for the Vitest phase. Direct npx vitest runs still require a held port-forward terminal:

make minikube-pf-all

For Control UI / Desktop on a branch-owned profile, the first-hand entry point (gitignored helper at repo root — do not search for it) is the host-side hold. Implementation: .local-notes/minikube-profiles/branch-profile.sh. HARD DENY: do not ls/cat ~/.cache/clerum/minikube-profiles/. Profile-owned random ports only (never shared :3000/:8090). make minikube-pf-all-bg is a gate refresh only; it must not replace branch-profile-pf. Do not start UI PFs from a sandboxed agent shell. Run the make target on the host. Do not kill this lane's branch-profile-pf. branch-profile-pf-health starts PFs then STOPS them on EXIT — do not use it as the lasting hold.

MINIKUBE_PROFILE=<owned-profile> \
  make -f .local-notes/minikube-profiles/branch.mk branch-profile-pf

MINIKUBE_PROFILE=<owned-profile> \
  make -f .local-notes/minikube-profiles/branch.mk branch-profile-health

With shared-profile (make minikube-pf-all) port-forwards held, verify the localhost ports used by the tests. A branch-owned profile uses the random ports from branch-profile-health, never shared :8090:

curl -sS http://127.0.0.1:8090/health
curl -sS http://127.0.0.1:8091/health
curl -sS http://127.0.0.1:8094/health
curl -sS http://127.0.0.1:8098/health
curl -sS http://127.0.0.1:8080/v1/runtime/health

Running E2E tests

# Runtime gate: agentic HTTP, snippet workflows, custom coordinator, rotation
./scripts/e2e/e2e-workflow-runtime-gate.sh

# Backend compatibility suites
./scripts/e2e/e2e-workflow-backend-compat.sh

# Cleanup runtime gate recipes
./scripts/e2e/e2e-workflow-runtime-gate.sh --cleanup

# Full bash + Vitest E2E
make test-e2e-all

The long-running software-creation suites are opt-in:

E2E_RUN_SOFTWARE_CREATION=1 make test-e2e-vitest

Individual suites

Runtime gate suites live in scripts/e2e/. Backend compatibility suites live in scripts/e2e/workflow-backend-compat/.

# Script Transport Tests Description
1 e2e-637-secret-ownership-bypass.sh mixed gate Recipe Secret ownership bypass probes
2 e2e-agentic-workflow-baseline.sh HTTP gate Agentic workflow baseline
3 e2e-snippet-runtime-smoke.sh snippet gate Fast snippet runtime smoke
4 e2e-snippet-runtime.sh mixed gate Snippet, DB, MCP, HTTP, and negative runtime paths
5 e2e-custom-coordinator-sdk.sh mixed gate Custom coordinator image runtime
6 e2e-workflow-token-rotation.sh mixed gate Runtime token rotation
7 e2e-agentic-stdio-baseline.sh stdio standalone Pure compute stdio baseline
8 workflow-backend-compat/*.sh HTTP/stdio compat Backend and transport compatibility, including stdio PostgreSQL and multi-tool flows

Rows 1–6 are exactly the six suites the runtime gate runs, in order.

scripts/e2e/ also holds many further suites that are not part of the runtime gate or the backend compatibility driver — webhooks, registry, shared filesystem, GFS, workflow approvals, sandbox-UI, ingress, and more — plus e2e-gke-*.sh (production GKE smoke tests) and e2e-prod-*.sh (production recipe regressions). Run those individually.

Backend compatibility suite results

Suite Transport Tests Description
mongodb-mcp-stack HTTP 36 MongoDB StatefulSet + PVC + 24 MongoDB tools
mock-mcp-with-db HTTP 32 PostgreSQL (runAsUser:70) + Mock MCP + cross-namespace NP
mcp-postgres HTTP 37 PostgreSQL + template interpolation ({{workload:field}})
mcp-redis-cache HTTP 33 Redis Deployment + binding NP cross-namespace
mcp-webhook-relay HTTP 32 MCP + CronJob + template resolution in args[]
stdio-mcp-with-postgres stdio 31 stdio + PostgreSQL StatefulSet + VCT + bindings
stdio-mcp-multi-tool stdio 39 2 stdio servers + Redis
Total 240 Phases 0–7; no LLM tool-calling or approval

The driver (scripts/e2e/e2e-workflow-backend-compat.sh) runs exactly these seven suites. The pure-compute stdio suite is the standalone e2e-agentic-stdio-baseline.sh, not a backend compatibility suite.

E2E phases (8 phases per suite)

Each backend compatibility suite validates 8 phases in order (Phase 0 through Phase 7). Any failed phase aborts the suite.

Phase What it tests
0 — Prerequisites Cluster reachable, namespaces, CRDs installed, core deployments healthy
1 — Clean Slate Delete previous recipe resources to guarantee a clean test
2 — Apply Recipe kubectl apply the WorkflowRecipe YAML
3 — Backend StatefulSet/Deployment readiness; data connectivity (pg_isready, mongosh, redis PING)
4 — MCP Delegation McpServer CRD auto-created, managed=false, transport Service, Context allowlist
5 — MCP Server Pod ready, transport protocol started
6 — NetworkPolicy Deny-all enforcement, binding NP cross-namespace, internet egress blocked
7 — mcp-proxy tool contract tools/list / tools/call against mcp-proxy for the delegated server

The backend compatibility suites stop at the mcp-proxy tool contract — they do not drive an LLM or an approval gate. The approval pipeline is covered separately by tests/e2e/e2e-approval-flow.sh (below).

Approval flow in E2E

tests/e2e/e2e-approval-flow.sh exercises the full production approval pipeline without disabling any security — approval is its Phase 5 ("Approve tool call"). No short-circuiting — the E2E runner plays the role of the approving user.

The mcp-host runtime routes are:

POST /v1/runtime/messages
  → response: { status: "waiting_approval", approval: { requestId, taskId } }
     ↓
POST /v1/runtime/approvals/approve
  body: { userId, requestId, channelType, channelId }   (+ x-clerum-edge-* headers)
     ↓
GET /v1/runtime/tasks/:taskId/result → poll until { status: "completed" }

The E2E suite drives these through the tenant gateway instead of hitting mcp-host directly: it approves at POST /api/v1/rpc/hosts/{hostRef}/approvals/approve with { toolCallId } and polls GET /api/v1/rpc/hosts/{hostRef}/tasks/{taskId}/result (see tests/e2e/e2e-approval-flow.sh).

This ensures E2E tests validate the real production flow, not a weakened test-mode path.

Unit test suites

Every service in the root Makefile's TEST_SERVICES list ships its own vitest suite (mcp-host, workflow-recipes, mcp-servers, control-api, host-context-controller, rpc-proxy, desktop-app, and the rest). Exact per-service counts drift as code changes, so this guide does not pin them — count them yourself, or run the suite:

make test-counts          # prints unit-test file and case counts
                          # (it()/test() cases in *.test.ts, excluding
                          #  node_modules, tests/e2e, and dist)

npm test inside each service directory is the source of truth for what actually passes.

Run all unit tests:

make test-unit-all        # from repo root — runs every service test suite
make test-e2e-all         # bash + vitest E2E suites
make test-integration     # cross-service integration tests
make validate-all         # install → unit → setup → integration → e2e

Per-service:

cd mcp-host && npm test
cd workflow-recipes && npm test
cd mcp-servers && npm test
cd control-api && npm test
cd host-context-controller && npm test

Desktop app E2E

Two-phase test strategy — requires port-forwards to a running cluster.

cd desktop-app
cp .env.e2e.example .env.e2e    # set E2E_DEV_LOGIN_EMAIL and E2E_HOST_REF

npm run test:e2e                # Phase 1: IPC harness   (~45s)
npm run test:e2e:playwright     # Phase 2: Playwright/Electron (~2-3 min)
npm run test:e2e:all            # Both phases

Existing mcp-host E2E (requires minikube):

cd tests/e2e && npx vitest run

Troubleshooting

Common test-related issues:

  • 409 on first-time setup → make minikube-setup ARGS="--reset-db --skip-build"
  • Postgres WAL corruption after cold start → make minikube-setup ARGS="--reset-db --skip-build" (auto-detected)
  • Pod ImagePullBackOff → make minikube-verify-images reports the cluster's mode and every missing ref; then re-pull with make minikube-pull-images (ghcr mode, which reuses the tag the cluster recorded) or rebuild with make minikube-build-images (local mode), and make minikube-restart-all. Every pod runs imagePullPolicy: IfNotPresent, so the ref only has to be present in the daemon.
  • 401: "Invalid token" from chatllm → make minikube-gen-keys followed by make minikube-sync-auth-key (auto-syncs + restarts). See the JWT auth chain notes at the top of the root Makefile.
  • Port-forward drops after pod restart → re-run make minikube-pf-desktop
  • Desktop app shows no agents after login → scripts/minikube/seed-test-data.sh
  • NetworkPolicy Phase 6 passes when it shouldn't → ensure minikube was started with --cni=calico; the default kindnet CNI silently ignores NetworkPolicy. Cross-check the rendered policies from the active overlay (kubectl kustomize deploy/overlays/minikube).
  • Approval times out (e2e-approval-flow.sh Phase 4/5) → check mcp-host logs for the awaiting_approval entry. The runner approves with { toolCallId } only; for rpc-proxy callers mcp-host derives the approving user from the x-clerum-edge-user-id header, not from the request body — so check the identity the RPC token was minted for.

Related

  • Root README.md §Development & testing — per-service test commands and suite summary
  • Per-service package.json test scripts (vitest config)
  • ../deploy/minikube.md — prerequisite deployment guide
  • ../architecture/overview.md NetworkPolicy model — required reading before debugging Phase 6