A small, runnable example of a pattern for building AI agents: write the orchestration once, run it in two different runtimes.
- Production runtime: Temporal. Durable execution: automatic retries, timeouts, and replay. Each step becomes a Temporal activity.
- Eval runtime: in-process. A fast loop that runs the same orchestration directly, mocking dependencies but calling the real LLMs.
The same orchestration code runs byte-for-byte in both.
input ──▶ orchestration(steps, input) ──▶ output
│ depends only on a typed
│ `Steps` interface
▼
┌─────────────────────┐
│ Steps interface │ (enrichWithWebData, classify, ...)
└─────────────────────┘
▲ ▲
injected by │ │ injected by
the Temporal │ │ the eval
runtime │ │ runtime
│ │
┌─────────────┴──────┐ ┌──────┴────────────────┐
│ StepsImpl methods │ │ StepsImpl called │
│ run as Temporal │ │ directly in-process; │
│ activities; real │ │ mock dependencies + │
│ dependencies + │ │ real LLM │
│ real LLM │ │ │
└────────────────────┘ └────────────────────────┘
- The orchestration is a plain async function. Its only dependency is a typed
Stepsinterface — it has no idea whether it is running inside Temporal or a test loop. - A runtime is an adapter that supplies a concrete
Stepsand calls the orchestration.
src/
platform/ runtime-agnostic primitives plus the two runtime adapters: temporal/ and eval/
agents/ agents in the platform, with a classify_business example
bin/
worker.ts `npm run worker`
classify.ts `npm run classify`
eval.ts `npm run eval`
Temporal workflows run in a deterministic sandbox with no access to Node.js APIs. Workflow code is bundled, and the bundle fails if it reaches a Node module. So the code is split into two sides that never cross:
- Workflow side: imports only orchestration functions and pure types.
- Worker side: a normal Node process. It instantiates each agent's
StepsImpland registers its methods as activities.
- Node.js 22+ (a
.nvmrcis included, sonvm useselects a compatible version). - Docker (for the Temporal runtime)
- A Google Gemini API key (for real LLM completions) — a free key from Google AI Studio works
cp .env.example .env # then add your GOOGLE_GENERATIVE_AI_API_KEY
npm install
npm run temporal:up # boots Temporal + UI (http://localhost:8080) via Docker
npm run worker # terminal 1 — the long-running worker (production runtime)
# terminal 2 — starts a workflow and opens the Temporal UI on the run
npm run classify -- --name "Google" --website "https://google.com"
npm run eval # the SAME orchestration through the in-process eval runtimeThen clean up:
npm run temporal:down| Command | What it does |
|---|---|
npm run temporal:up |
docker compose up -d, then waits for the Temporal UI on :8080 |
npm run temporal:down |
docker compose down |
npm run worker |
Runs the standalone Temporal worker (mirrors production) |
npm run classify |
Starts the classify_business workflow, opens the UI on the run, prints the result. Accepts --name and --website. Requires npm run worker to be running. |
npm run eval |
Runs the orchestration over dataset.json entirely in-process — no Docker, no worker — and prints the expected vs. actual category for each entry |
npm run eval:braintrust |
Runs the same orchestration through the Braintrust eval runtime and reports results there. Requires BRAINTRUST_API_KEY. |
npm run typecheck |
tsc --noEmit |
npm run lint |
Biome check |
Without
GOOGLE_GENERATIVE_AI_API_KEY, the LLM step fails with an auth/missing-key error.
runEval is the seam. Pointing the eval loop at a hosted eval platform (Braintrust, Laminar, LangSmith, ...) is just wrapping that call — the orchestration does not change. A working Braintrust runner lives in src/bin/braintrust-eval.ts, a sibling of src/bin/eval.ts: Braintrust owns the dataset iteration, scoring, and reporting, while task calls the same runEval that npm run eval uses.
To run it, add a free Braintrust API key to .env:
BRAINTRUST_API_KEY=... # plus GOOGLE_GENERATIVE_AI_API_KEY for the real LLM
npm run eval:braintrust # runs the same orchestration, reports to Braintrust- Runs directly from TypeScript with
tsx; no build step.npm run typecheckrunstsc --noEmit. - The default model is
gemini-2.5-flash(override withGEMINI_MODEL). - The Temporal stack (
docker-compose.yml) usestemporalio/auto-setup, which is intended for local development only.