Skip to content

Repository files navigation

Runtime-agnostic AI workflows

A small, runnable example of a pattern for building AI agents: write the orchestration once, run it in two different runtimes.

  • Production runtime: Temporal. Durable execution: automatic retries, timeouts, and replay. Each step becomes a Temporal activity.
  • Eval runtime: in-process. A fast loop that runs the same orchestration directly, mocking dependencies but calling the real LLMs.

The same orchestration code runs byte-for-byte in both.

The pattern

        input ──▶  orchestration(steps, input)  ──▶  output
                        │  depends only on a typed
                        │  `Steps` interface
                        ▼
              ┌─────────────────────┐
              │   Steps interface   │   (enrichWithWebData, classify, ...)
              └─────────────────────┘
                 ▲                 ▲
   injected by   │                 │   injected by
   the Temporal  │                 │   the eval
   runtime       │                 │   runtime
                 │                 │
   ┌─────────────┴──────┐   ┌──────┴────────────────┐
   │ StepsImpl methods  │   │ StepsImpl called       │
   │ run as Temporal    │   │ directly in-process;   │
   │ activities; real   │   │ mock dependencies +    │
   │ dependencies +     │   │ real LLM               │
   │ real LLM           │   │                        │
   └────────────────────┘   └────────────────────────┘
  • The orchestration is a plain async function. Its only dependency is a typed Steps interface — it has no idea whether it is running inside Temporal or a test loop.
  • A runtime is an adapter that supplies a concrete Steps and calls the orchestration.

Architecture

src/
  platform/       runtime-agnostic primitives plus the two runtime adapters: temporal/ and eval/ 
  agents/         agents in the platform, with a classify_business example
  bin/
    worker.ts     `npm run worker`
    classify.ts   `npm run classify`
    eval.ts       `npm run eval`

The worker/workflow split (and why it matters)

Temporal workflows run in a deterministic sandbox with no access to Node.js APIs. Workflow code is bundled, and the bundle fails if it reaches a Node module. So the code is split into two sides that never cross:

  • Workflow side: imports only orchestration functions and pure types.
  • Worker side: a normal Node process. It instantiates each agent's StepsImpl and registers its methods as activities.

Prerequisites

  • Node.js 22+ (a .nvmrc is included, so nvm use selects a compatible version).
  • Docker (for the Temporal runtime)
  • A Google Gemini API key (for real LLM completions) — a free key from Google AI Studio works

Quickstart

cp .env.example .env   # then add your GOOGLE_GENERATIVE_AI_API_KEY
npm install
npm run temporal:up    # boots Temporal + UI (http://localhost:8080) via Docker

npm run worker         # terminal 1 — the long-running worker (production runtime)

# terminal 2 — starts a workflow and opens the Temporal UI on the run
npm run classify -- --name "Google" --website "https://google.com"

npm run eval           # the SAME orchestration through the in-process eval runtime

Then clean up:

npm run temporal:down

Commands

Command What it does
npm run temporal:up docker compose up -d, then waits for the Temporal UI on :8080
npm run temporal:down docker compose down
npm run worker Runs the standalone Temporal worker (mirrors production)
npm run classify Starts the classify_business workflow, opens the UI on the run, prints the result. Accepts --name and --website. Requires npm run worker to be running.
npm run eval Runs the orchestration over dataset.json entirely in-process — no Docker, no worker — and prints the expected vs. actual category for each entry
npm run eval:braintrust Runs the same orchestration through the Braintrust eval runtime and reports results there. Requires BRAINTRUST_API_KEY.
npm run typecheck tsc --noEmit
npm run lint Biome check

Without GOOGLE_GENERATIVE_AI_API_KEY, the LLM step fails with an auth/missing-key error.

Swapping eval platforms

runEval is the seam. Pointing the eval loop at a hosted eval platform (Braintrust, Laminar, LangSmith, ...) is just wrapping that call — the orchestration does not change. A working Braintrust runner lives in src/bin/braintrust-eval.ts, a sibling of src/bin/eval.ts: Braintrust owns the dataset iteration, scoring, and reporting, while task calls the same runEval that npm run eval uses.

To run it, add a free Braintrust API key to .env:

BRAINTRUST_API_KEY=...      # plus GOOGLE_GENERATIVE_AI_API_KEY for the real LLM
npm run eval:braintrust     # runs the same orchestration, reports to Braintrust

Notes

  • Runs directly from TypeScript with tsx; no build step. npm run typecheck runs tsc --noEmit.
  • The default model is gemini-2.5-flash (override with GEMINI_MODEL).
  • The Temporal stack (docker-compose.yml) uses temporalio/auto-setup, which is intended for local development only.

About

A small, runnable example of a pattern for building AI agents

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages