Skip to content

Commit 32876a1

Browse files
committed
chore(release): 0.13.10 — vendor extractor edge cases
Bump version 0.13.9 -> 0.13.10. Prepends the v3.24 / 0.13.10 changelog entry to __version__.py and backfills the inline pyproject.toml comment block for 0.13.6..0.13.10 (the previous comment trail stopped at 0.13.5, which made the pyproject side of the release notes drift behind the docstring). The 5 vendor edge cases themselves ship in the preceding commit 3ac7231. No on-wire change. No SDK_MIN_VERSION bump. Backends on 1.0.0 keep working unchanged. Recommended upgrade path: 0.13.9 -> 0.13.10.
1 parent 3ac7231 commit 32876a1

2 files changed

Lines changed: 147 additions & 2 deletions

File tree

‎pyproject.toml‎

Lines changed: 31 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -28,7 +28,37 @@ name = "nullrun"
2828
# the full ``flush_interval`` (5s default). Plus CI hygiene:
2929
# pip cache, ``fail-fast`` matrix, ``pytest-xdist -n auto``. No
3030
# on-wire change; backends on 1.0.0 keep working unchanged.
31-
version = "0.13.9"
31+
# 0.13.6 (2026-07-11): multi-agent span attachment (parent_trace_id)
32+
# on the langgraph callback; new cost_events.parent_trace_id column
33+
# (backend migration 217). Wire-additive — legacy backends ignore
34+
# the field. Pairs with PR #61.
35+
# 0.13.7 (2026-07-12): wire ``parent_trace_id`` end-to-end on
36+
# ``/track`` (v3 single-event + legacy /track/batch). Pre-fix the
37+
# SDK stamped the field in the langgraph callback but dropped it
38+
# at the runtime._enrich_event / _build_v3_track_payload layers.
39+
# No on-wire change for legacy backends; new column required on
40+
# the v3 path. Pairs with PR #64.
41+
# 0.13.8 (2026-07-12): hotfix #2 for parent_trace_id — the
42+
# runtime._enrich_event parent_trace_id fallback used an
43+
# "if not in enriched" guard that missed whenever langgraph.py
44+
# callback's _active_runs lookup missed (run_id drift between
45+
# auto-injected and user-supplied callbacks, or non-langgraph
46+
# stacks). Switched to override semantics: the chain contextvar
47+
# is the single source of truth; both caller-set and
48+
# contextvar-fallback resolve to the same value (idempotent for
49+
# the happy path, closes the drift in the unhappy path).
50+
# Pairs with PR #66.
51+
# 0.13.9 (2026-07-13): crewai 1.15 compatibility — replace
52+
# step_callback kwargs injection (removed upstream) with a
53+
# crewai_event_bus bridge; gate_cache re-capture for fresh
54+
# server-minted execution_id on cache-hit. Pairs with PR #67.
55+
# 0.13.10 (2026-07-13): close 5 vendor extractor edge cases
56+
# missed in the 0.13.9 audit — Cohere v2 nested tool_calls +
57+
# cached_tokens + UPPERCASE finish_reason; Mistral flat
58+
# num_cached_tokens fallback; Gemini 2.5+ thoughtsTokenCount;
59+
# Anthropic 4.5+ extended-thinking tokens; Bedrock Mistral/Llama
60+
# finish_reason paths. No on-wire change; no SDK_MIN_VERSION bump.
61+
version = "0.13.10"
3262
# Kept under the 200-char preview threshold so the full line is visible
3363
# without an "expand" click. Keywords are matched against likely search
3464
# queries ("AI agent cost control", "LLM circuit breaker", etc.).

‎src/nullrun/__version__.py‎

Lines changed: 116 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,120 @@
11
"""NullRun Platform SDK.
22
3+
v3.24 / 0.13.10 (2026-07-13) — close 5 vendor extractor edge cases
4+
missed in the 0.13.9 audit.
5+
6+
1. Cohere v2 tool_calls path: the pre-0.13.10 extractor read
7+
top-level payload["tool_calls"], but Cohere v2 nests the field
8+
under message.tool_calls (OpenAI shape). Every v2 Cohere call
9+
shipped with tool_names=[] and the backend's loop detection
10+
could not see Cohere tool use. Fix walks both v1 (top-level)
11+
and v2 (message.tool_calls) paths. Same patch adds
12+
usage.tokens.cached_tokens (cache-hit read was always 0) and
13+
the UPPERCASE finish_reason vocabulary
14+
(COMPLETE | MAX_TOKENS | TOOL_CALL) — the _FINISH_REASON_MAP
15+
already lower-cased both vocabularies; the missing piece was
16+
the test snapshot.
17+
18+
2. Mistral num_cached_tokens (flat field on usage, not nested
19+
under prompt_tokens_details.cached_tokens like OpenAI's). The
20+
OpenAI extractor only read the nested shape, so Mistral
21+
customers always saw cache_read_tokens=0 even when the
22+
inference cache hit. Fix reads the Mistral flat field as a
23+
fallback inside the same chain. The _openai_extractor host
24+
map (line 567) already covers Mistral so no host-routing
25+
change was needed.
26+
27+
3. Gemini 2.5+ thoughtsTokenCount (reasoning tokens in
28+
usageMetadata) — was hard-coded to 0, so thinking-mode Gemini
29+
calls had no visible reasoning column on the dashboard.
30+
Surfaced as reasoning_tokens while the total stays at
31+
totalTokenCount (reasoning tokens are part of
32+
candidatesTokenCount upstream).
33+
34+
4. Anthropic 4.5+ output_tokens_details.thinking_tokens
35+
(extended-thinking mode) — was hard-coded to 0 for the same
36+
reason. The pre-0.13.10 comment ("reasoning tokens are part
37+
of output_tokens") was correct for the non-thinking baseline,
38+
but the thinking-mode field was still readable and was being
39+
dropped. Now we read the breakdown while keeping the total at
40+
input+output (Anthropic bills thinking tokens at the output
41+
rate upstream).
42+
43+
5. AWS Bedrock finish_reason for the Mistral-on-Bedrock /
44+
OpenAI-compat and Llama-on-Bedrock adapter shapes. The
45+
pre-0.13.10 extractor only read top-level stopReason /
46+
stop_reason (Anthropic + Llama top-level). Mistral's
47+
OpenAI-compat shape puts the field under
48+
choices[0].finish_reason and was always None. The
49+
matched_shape discriminator (already tracked in the
50+
tool-detection block) tells us which body to read from and
51+
the new branch picks choices[0].finish_reason when
52+
matched_shape == 'openai_choices'.
53+
54+
The same audit identified the following as should-fix but
55+
deferred to a follow-up PR (none is a billing gap; all are
56+
visibility / observability gaps):
57+
58+
- Anthropic cache_creation.ephemeral_{1h,5m}_input_tokens
59+
TTL breakdown (different billing rates; Bedrock does not
60+
yet expose the breakdown as of 2026-Q3).
61+
- Anthropic server_tool_use.{web_search_requests,
62+
web_fetch_requests} — server-side tool invocations not
63+
visible to loop detection.
64+
- Gemini multimodal *TokensDetails[] (TEXT vs IMAGE vs
65+
AUDIO) — image-heavy calls mask the real cost driver.
66+
- Cohere billed_units.{search_units, classifications} for
67+
RAG / classify workloads.
68+
- Cohere reasoning models (command-a-reasoning-*).
69+
- Bedrock Converse API (separate envelope from InvokeModel).
70+
71+
Wire format: unchanged. Backends on 1.0.0 keep working
72+
unchanged. Pinning unchanged: SDK_MIN_VERSION_FOR_V3 =
73+
"0.12.0". Recommended upgrade path: 0.13.9 -> 0.13.10.
74+
75+
Tests (8 new in tests/test_extractors.py):
76+
77+
- test_cohere_v2_message_tool_calls_path — v2 nested
78+
message.tool_calls returns the right tool_names.
79+
- test_cohere_v2_cached_tokens — tokens.cached_tokens
80+
surfaces as cache_read_tokens.
81+
- test_cohere_v1_top_level_tool_calls_fallback — v1
82+
callers (legacy top-level tool_calls) keep working.
83+
- test_openai_mistral_num_cached_tokens — Mistral
84+
usage.num_cached_tokens fallback in the OpenAI extractor.
85+
- test_gemini_2_5_thinking_tokens — thoughtsTokenCount
86+
surfaces as reasoning_tokens while the total stays at
87+
totalTokenCount.
88+
- test_anthropic_extended_thinking_tokens —
89+
output_tokens_details.thinking_tokens surfaces
90+
alongside cache_read_input_tokens /
91+
cache_creation_input_tokens already extracted.
92+
- test_bedrock_mistral_finish_reason_via_choices —
93+
Mistral-on-Bedrock OpenAI-compat finish_reason is now
94+
captured.
95+
- test_bedrock_llama_finish_reason_via_top_level —
96+
Llama-on-Bedrock stop_reason snake_case is captured
97+
(already worked, but had no test snapshot before).
98+
99+
Verification locally (origin/master + this commit on top):
100+
101+
* pytest tests/test_extractors.py tests/test_crewai_patch.py
102+
tests/test_runtime.py tests/test_runtime_branches.py
103+
tests/test_track_batch_retry.py
104+
tests/test_track_span_context.py
105+
tests/test_v3_wire_contract.py tests/test_release_polish.py
106+
— 185 passed, 1 skipped (8 new tests net-new from this
107+
commit; no regression on the 177 tests that were green
108+
on master).
109+
* ruff check src/ — "All checks passed!".
110+
* mypy src/ — 11 pre-existing errors (langgraph overload
111+
mismatches at lines 1818, 1821, 1827; same count as
112+
origin/master). No new mypy findings from this release.
113+
114+
No public API change. No SDK_MIN_VERSION bump.
115+
116+
---
117+
3118
v3.23 / 0.13.9 (2026-07-13) — crewai 1.15 compatibility + gate_cache
4119
re-capture.
5120
@@ -579,5 +694,5 @@
579694
580695
"""
581696

582-
__version__ = "0.13.9"
697+
__version__ = "0.13.10"
583698
__platform_version__ = "1.0.0"

0 commit comments

Comments
 (0)