Skip to content

Commit cfe409a

Browse files
committed
fix(gate): forward real model + tools to /gate pre-flight (T4)
Pre-0.7.7 every SDK /gate call for any workflow with a budget was hard-blocked because the runtime hard-coded the literal string "budget-precheck" as the model. The backend's PolicyEvaluationGraph treated any synthetic cost_limit rule with score > 0.8 as Block, so the pricing lookup never landed on a real model and the rule fired with the wrong score. This commit: * Adds nullrun.set_call_context(model=..., tools=[...]) plus get_call_model / get_call_tools helpers (and the underlying _call_model_var / _call_tools_var contextvars in nullrun.context). * Wires the call context into check_workflow_budget: the /gate payload now carries the real model name (or None when unset) and the user-supplied tool list. tools=[] vs missing-None are distinguished on the wire per gate/internal.rs::check_tool_block. * Transport.check forwards the tools key when set (it was silently dropped pre-fix). * tests/conftest.py reset_runtime clears the new contextvars so a test's set_call_context(...) doesn't leak into the next test's wire payload. * New tests/test_gate_real_path.py pins down the regression: default request allows a clean workflow, real block still honored, no policy-N residue on the wire, set_call_context flows into the body, no-context means no tools key, and the helpers are reachable from nullrun.*. Bumps version to 0.7.7. No breaking changes - new helpers default to None / empty so existing call sites keep working.
1 parent 95225d8 commit cfe409a

8 files changed

Lines changed: 433 additions & 3 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 91 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -7,6 +7,97 @@ Versioning: [Semantic Versioning](https://semver.org/spec/v2.0.0.html)
77

88
---
99

10+
## [0.7.7] - 2026-06-27
11+
12+
Additive patch on top of 0.7.6. Fixes the `/gate` pre-flight so the
13+
backend can compute `projected_cost` and `tool_block` decisions from
14+
real per-call data instead of the previous fake `"budget-precheck"`
15+
sentinel and empty tool list. No breaking changes — new helpers
16+
default to `None` / empty so existing call sites keep working.
17+
18+
### Added
19+
20+
- **`nullrun.set_call_context(model=..., tools=[...])`** — per-call
21+
context the SDK forwards to `/gate` so the backend can enforce
22+
budget tiers and tool-block on real values.
23+
```python
24+
import nullrun
25+
26+
with nullrun.workflow(name="support-bot"):
27+
nullrun.set_call_context(
28+
model="claude-sonnet-4-6",
29+
tools=["shell.run", "code.eval"],
30+
)
31+
32+
@nullrun.protect
33+
def chat(message: str) -> str:
34+
return agent.run(message)
35+
```
36+
- `model` (optional) — LLM model name. Backend uses it to look up
37+
the per-model rate from `tool_pricing` (Postgres) so
38+
`projected_cost` matches what `/track` will compute from real
39+
token counts. Defaults to `None` (backend falls back to
40+
`claude-sonnet-4` default rate).
41+
- `tools` (optional) — list of tool names the call intends to use.
42+
Backend matches each against the workflow's effective
43+
`blocked_tools` aggregate and returns `block` on any match.
44+
`None` leaves whatever was previously set; `[]` clears.
45+
- `nullrun.get_call_model()` and `nullrun.get_call_tools()` are
46+
the read-side helpers (also reachable via
47+
`nullrun.context.get_call_model` / `get_call_tools`).
48+
49+
### Fixed
50+
51+
- **`/gate` pre-flight no longer sends `model="budget-precheck"`.**
52+
Pre-0.7.7 every SDK `/gate` call for any workflow with a budget
53+
was hard-blocked because the runtime hard-coded the literal
54+
string `"budget-precheck"` as the model. The backend's
55+
`PolicyEvaluationGraph.evaluate()` stub treated any synthetic
56+
`cost_limit` rule with score > 0.8 as `Block` (see
57+
`backend/src/policy/graph.rs:448-462`,
58+
`backend/src/proxy/http/gate/internal.rs:619-628`), so the
59+
pricing lookup never landed on a real model and the rule fired
60+
with the wrong score. Now the runtime forwards the model from
61+
`set_call_context(model=...)` (or `None` when unset), and the
62+
backend's `calculate_projected_cost` falls through to the
63+
default rate cleanly.
64+
65+
- **`/gate` pre-flight now forwards the per-call `tools` list.**
66+
`Transport.check` previously dropped the `tools` key from the
67+
wire payload, so even when the user called
68+
`set_call_context(tools=[...])` the backend's
69+
`gate/internal.rs::check_tool_block` had nothing to match
70+
against. The transport now propagates `tools` when the runtime
71+
sets it; `[]` vs missing-`None` are distinguished on the wire
72+
(per `gate/internal.rs::check_tool_block` doc-comment —
73+
"no tools will be called" is different from "I did not tell you
74+
what tools").
75+
76+
### Tests
77+
78+
- **`tests/test_gate_real_path.py`** (new, 226 lines) — regression
79+
test pinning the fix. Three classes:
80+
- `TestGateRealPathRegression` — default request now returns
81+
`allow` (not the old blanket block on the synthetic
82+
`cost_limit` rule), wire payload contains no
83+
`policy-N` residue from the old graph plumbing, and a real
84+
`decision="block"` still raises `WorkflowKilledInterrupt`
85+
(so the fix didn't accidentally remove the real-block path).
86+
- `TestSetCallContext` — `set_call_context(model=...)` flows
87+
into the wire body, `set_call_context(tools=[...])` flows
88+
into the wire body, no-context means no `tools` key at all
89+
(not `[]`), and `set_call_context(tools=[])` clears a
90+
previously-set tool list.
91+
- `TestPackageExports` — the new helpers are reachable from
92+
`nullrun.*`.
93+
94+
- `tests/conftest.py` — `reset_runtime` fixture now also clears
95+
`_call_model_var` and `_call_tools_var` so a test's
96+
`set_call_context(...)` doesn't leak into the next test's wire
97+
payload.
98+
99+
---
100+
10101
## [0.7.6] - 2026-06-27
11102

12103
Additive patch on top of the 0.7.0 thin-client refactor. Brings a

‎pyproject.toml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
44

55
[project]
66
name = "nullrun"
7-
version = "0.7.6"
7+
version = "0.7.7"
88
# Long form used by PyPI page meta-description and search snippets.
99
# Kept under the 200-char preview threshold so the full line is visible
1010
# without an "expand" click. Keywords are matched against likely search

‎src/nullrun/__init__.py‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -331,6 +331,14 @@ def my_agent():
331331
"get_trace_id": ("nullrun.context", "get_trace_id"),
332332
"get_span_id": ("nullrun.context", "get_span_id"),
333333
"get_agent_id": ("nullrun.context", "get_agent_id"),
334+
# T4 (2026-06-27): per-call context for /gate pre-flight. Users
335+
# call `set_call_context(model=..., tools=[...])` inside
336+
# `with workflow(...)` so the backend's budget + tool_block
337+
# enforcement sees real values instead of the previous fake
338+
# `"budget-precheck"` sentinel and empty tool list.
339+
"set_call_context": ("nullrun.context", "set_call_context"),
340+
"get_call_model": ("nullrun.context", "get_call_model"),
341+
"get_call_tools": ("nullrun.context", "get_call_tools"),
334342
# Instrumentation
335343
"NullRunCallback": ("nullrun.instrumentation", "NullRunCallback"),
336344
# NOTE (Sprint 1.2 / B11-B12): `patch_openai` and `unpatch_openai`

‎src/nullrun/context.py‎

Lines changed: 62 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -31,6 +31,17 @@
3131
_agent_id_var: ContextVar[str | None] = ContextVar("agent_id", default=None)
3232
_attempt_index_var: ContextVar[int] = ContextVar("attempt_index", default=0)
3333

34+
# T4 (2026-06-27): per-call context that flows into the /gate pre-flight
35+
# request so the backend can compute projected_cost and tool_block
36+
# decisions from real data instead of the previous fake "budget-precheck"
37+
# sentinel. Both default to None/empty; users opt in by calling
38+
# ``set_call_context(model=..., tools=[...])`` inside a ``with workflow(...)``
39+
# block. When unset, the backend falls back to its default pricing and
40+
# skips tool-block enforcement on /gate (per-key tool_block is enforced
41+
# on /track only — see gate/internal.rs T3).
42+
_call_model_var: ContextVar[str | None] = ContextVar("call_model", default=None)
43+
_call_tools_var: ContextVar[tuple[str, ...]] = ContextVar("call_tools", default=())
44+
3445

3546
# =============================================================================
3647
# Workflow / trace getters
@@ -62,11 +73,62 @@ def get_attempt_index() -> int:
6273
return _attempt_index_var.get()
6374

6475

76+
def get_call_model() -> str | None:
77+
"""Get the LLM model name set via ``set_call_context``.
78+
79+
Used by ``check_workflow_budget`` to send the real model to the
80+
backend's /gate endpoint instead of the previous fake
81+
``"budget-precheck"`` placeholder (which forced the backend's
82+
pricing model to fall through to the default rate and broke any
83+
future per-model budget tiers).
84+
"""
85+
return _call_model_var.get()
86+
87+
88+
def get_call_tools() -> tuple[str, ...]:
89+
"""Get the tool names set via ``set_call_context``.
90+
91+
Used by ``check_workflow_budget`` so the backend's tool_block
92+
enforcement (when added in T3) can match against the workflow's
93+
configured ``blocked_tools`` aggregate.
94+
"""
95+
return _call_tools_var.get()
96+
97+
6598
def set_attempt_index(index: int) -> None:
6699
"""Set current attempt index for retry correlation."""
67100
_attempt_index_var.set(index)
68101

69102

103+
def set_call_context(
104+
model: str | None = None,
105+
tools: list[str] | tuple[str, ...] | None = None,
106+
) -> None:
107+
"""Set per-call context (model name, tool list) for the next /gate
108+
pre-flight check.
109+
110+
T4 (2026-06-27): replaces the previous fake ``model="budget-precheck"``
111+
and ``estimated_tokens=1`` always-default / always-empty pre-flight.
112+
Call inside a ``with workflow(...)`` block before ``@protect`` to
113+
give the backend real data.
114+
115+
Args:
116+
model: LLM model name (e.g. ``"claude-sonnet-4-6"``). Backend
117+
uses this to look up the per-model rate from
118+
``tool_pricing`` (Postgres) so projected_cost matches what
119+
/track will compute from real token counts.
120+
tools: List of tool names the call intends to use. Backend
121+
matches each against the workflow's effective
122+
``blocked_tools`` aggregate (T3 in backend) and returns
123+
block on any match. Pass ``None`` to leave whatever was
124+
previously set, ``[]`` to clear.
125+
"""
126+
if model is not None:
127+
_call_model_var.set(model)
128+
if tools is not None:
129+
_call_tools_var.set(tuple(tools))
130+
131+
70132
def generate_trace_id() -> str:
71133
"""Generate a new trace ID.
72134

‎src/nullrun/runtime.py‎

Lines changed: 26 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1149,7 +1149,7 @@ def check_workflow_budget(self) -> None:
11491149
# running (not silently always-skipped).
11501150
metrics.inc_runtime("check_calls")
11511151

1152-
from nullrun.context import get_workflow_id
1152+
from nullrun.context import get_call_model, get_call_tools, get_workflow_id
11531153

11541154
# Phase 139+: prefer the user-set contextvar (explicit `with
11551155
# workflow(...)` block), fall back to the API key's bound
@@ -1160,15 +1160,39 @@ def check_workflow_budget(self) -> None:
11601160
if not workflow_id:
11611161
return
11621162

1163+
# T4 (2026-06-27): use the real model name from the call
1164+
# context if the user set it via `set_call_context(model=...)`
1165+
# (or via a future `with workflow(..., model=...)` block).
1166+
# Pre-T4 this always sent the literal string "budget-precheck"
1167+
# — a fake sentinel that:
1168+
# 1. forced backend pricing lookup to fall through to the
1169+
# default 3.0 rate, so projected_cost was always computed
1170+
# against the wrong per-model rate;
1171+
# 2. blocked any future per-model budget tier (model-specific
1172+
# caps) from being enforced correctly.
1173+
# Sending `None` is fine — backend `calculate_projected_cost`
1174+
# defaults to claude-sonnet-4 when model is unset, and tool_block
1175+
# enforcement on /gate is best-effort when no tools are sent.
1176+
call_model = get_call_model()
1177+
call_tools = get_call_tools()
1178+
11631179
check_req = {
11641180
"organization_id": self.organization_id or "local",
11651181
"execution_id": workflow_id,
11661182
"operation_id": str(uuid.uuid4()),
11671183
"check_type": "llm",
1168-
"model": "budget-precheck",
1184+
"model": call_model, # may be None if user didn't set it
11691185
"estimated_tokens": 1,
11701186
}
11711187

1188+
# Forward the tool list so backend (T3) can match each tool
1189+
# against the workflow's effective `blocked_tools` aggregate.
1190+
# Only included when the user actually set it — `[]` means
1191+
# "no tools will be called" which is different from "I didn't
1192+
# tell you what tools will be called" (None).
1193+
if call_tools:
1194+
check_req["tools"] = list(call_tools)
1195+
11721196
try:
11731197
response = self._transport.check(check_req)
11741198
except Exception as exc: # noqa: BLE001

‎src/nullrun/transport.py‎

Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1419,6 +1419,18 @@ def check(
14191419
"model": check_request.get("model"),
14201420
"estimated_tokens": check_request.get("estimated_tokens"),
14211421
"operation_id": check_request.get("operation_id") or str(uuid.uuid4()),
1422+
# T4 (2026-06-27): forward the per-call `tools` list so the
1423+
# backend's `gate/internal.rs::check_tool_block` can match
1424+
# each tool against the workflow's effective `blocked_tools`
1425+
# aggregate. Pre-T4 this key was silently dropped here, so
1426+
# `set_call_context(tools=[...])` had no effect on /gate.
1427+
# When unset (None) we omit the key entirely — the backend
1428+
# distinguishes "no tools sent" from "explicit []".
1429+
**(
1430+
{"tools": check_request["tools"]}
1431+
if "tools" in check_request
1432+
else {}
1433+
),
14221434
}
14231435

14241436
headers = {"Content-Type": "application/json"}

‎tests/conftest.py‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -17,6 +17,7 @@ def reset_runtime():
1717
import nullrun.actions as _act
1818
import nullrun.decorators as _dec
1919
import nullrun.runtime as _rt_mod
20+
from nullrun.context import _call_model_var, _call_tools_var
2021
from nullrun.runtime import NullRunRuntime
2122

2223
# Disable polling for all tests via the runtime's internal `polling` flag
@@ -33,6 +34,11 @@ def reset_runtime():
3334
# leaks across the suite (e.g. a test that did `nullrun.init(...)` with
3435
# the prod URL leaves that URL pinned for the next test).
3536
_rt_mod._runtime = None
37+
# T4 (2026-06-27): reset the per-call context (model + tools) so a
38+
# previous test's `set_call_context(...)` doesn't leak into the next
39+
# test's wire payload.
40+
_call_model_var.set(None)
41+
_call_tools_var.set(())
3642

3743
yield
3844

@@ -42,6 +48,8 @@ def reset_runtime():
4248
_dec._runtime = None
4349
_act._action_handler = None
4450
_rt_mod._runtime = None
51+
_call_model_var.set(None)
52+
_call_tools_var.set(())
4553

4654

4755
@pytest.fixture

0 commit comments

Comments
 (0)