fix(discover): declare tool_call.arguments as a JSON string so constrained providers stop emptying it - #209
Conversation
The `arguments` field in the tool call schema was declared as a JSON object, but this caused providers with strict schema validation to reject valid tool calls because the open-ended object type was interpreted as disallowing any keys. Changing the type to a JSON string preserves compatibility across all provider projections while still accepting object payloads from models that send them. Auto-committed-on: macbook
…jections
The `arguments` field in `tool_call` is now explicitly typed as a JSON string in the schema, ensuring it reaches the model intact under every schema projection. Previously, an open object type was answered as `{}` by schema-constrained providers, and various projections would rewrite or strip `additionalProperties`, causing the field to be lost. The change updates the help text in the search answer to clarify the expected format and adds comprehensive tests verifying that `arguments` remains a string across all supported schema preparations, including Gemini, Anthropic, OpenAI, conservative, and strict variants.
Auto-committed-on: macbook
Reordered the import of `prepare_tool_schema` and `SchemaPreparation` to follow standard Rust convention, and reformatted the long tuple entry for "conservative strict" to improve readability without changing any behavior. Auto-committed-on: macbook
|
Warning Review limit reached
This review includes 2 billable files and costs up to $0.50. Or wait 29 minutes for your next included review. View limit detailsLimit details: You’ve used the included review currently available. Review configuration: ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (2)
Comment |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: dbb7246480
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| "Invoke a tool found with `{TOOL_SEARCH_NAME}`. `name` is the tool's name \ | ||
| and `arguments` is its argument object, matching the schema the search \ | ||
| returned." | ||
| and `arguments` is its argument object encoded as a JSON string, matching \ | ||
| the schema the search returned." |
There was a problem hiding this comment.
Document the string-encoded bridge argument
This changes the model-facing tool_call contract from an argument object to a JSON-encoded string, but docs/modules/harness/tool-discovery.md still describes only tool_call { name, arguments } without explaining the required encoding. Document the string form and include an example so hosts and prompt integrations relying on the module guide do not implement the old wire shape.
AGENTS.md reference: AGENTS.md:L78-L82
Useful? React with 👍 / 👎.
Tiny Sweeper reviewTiny Sweeper reviewed this change across 6 lane(s) and found 0 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below. State: Ready for maintainer review Review snapshot
Completeness: Complete What changedThe review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below. FeaturesNone identified with supported citations. TestsNo supported feature-to-test mapping was produced. Test execution is not inferred. FindingsNo active actionable findings. Before mergeNone. How this fits togetherflowchart LR
n0["answer_tool_search<br/>changed"]:::changed
n1["tool_call_schema<br/>changed"]:::changed
n2["tool_search_schema<br/>changed"]:::changed
n3["catalog"]:::impacted
n4["catalog_with_families"]:::impacted
n5["with_ranker"]:::impacted
n6["bridge_schemas"]:::impacted
n7["bm25_mode_ignores_an_installed_ranker"]:::impacted
n8["...ves_the_ranker_and_reports_bm25_alongside"]:::impacted
n4 -->|calls| n3
n6 -->|calls| n1
n6 -->|calls| n2
n7 -->|calls| n0
n7 -->|tests| n0
n7 -->|calls| n4
n7 -->|tests| n4
n7 -->|calls| n5
n7 -->|tests| n5
n8 -->|calls| n0
n8 -->|tests| n0
n8 -->|calls| n4
n8 -->|tests| n4
n8 -->|calls| n5
n8 -->|tests| n5
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Agent review detailscritique
security
tests
commits
description
e2e
Evidence and run details
|
There was a problem hiding this comment.
tinysweeper found nothing blocking. Approving.
$0.0075 · 115,239 in / 8,635 out · 20,888 cached (18%) · ladder/vectors, gpt-5.6-luna, deepseek/deepseek-v4-flash · 476 embedded
critique: $0.0019 · 39,198 in / 367 out · 2,118 cached (5%) · gpt-5.6-luna
security: $0.0020 · 38,710 in / 925 out · 1,874 cached (5%) · gpt-5.6-luna
tests: $0.0023 · 28,000 in / 4,423 out · 14,336 cached (51%) · deepseek/deepseek-v4-flash
description: $0.0004 · 5,682 in / 478 out · 2,560 cached (45%) · deepseek/deepseek-v4-flash
Summary
tool_call.argumentsis now declared as a JSON-encoded string instead of an open object.A schema-constrained provider answered every
tool_callwith{}, droppingnameas well as the arguments, so tools reached throughtool_searchcould never actually be invoked. That broke OpenHuman's scheduled "email me the AAPL price" job.The failure, as captured on the wire
OpenHuman's capture proxy recorded the managed backend's replies for
openrouter/deepseek/deepseek-v4-flash-0731. The OpenRouter provider Sail Research served every one of them.The
GMAIL_SEND_EMAILstep came back as nativetool_callslike this, three times out of three:{"function": {"name": "tool_call", "arguments": "{}"}}So the loop answered "
tool_callneeds a non-empty stringname" until the model gave up.Why a string, measured
I replayed the same captured request (
tool_callstep, 27 tools, same prompt) against the same route. The schema variant was the only change:tool_call.argumentsschematool_callattemptsname+ args intact{"type":"object"}(before){}{"type":"object","additionalProperties":true}{"type":"string"}(this PR)The open object works on this route, but it doesn't survive our own schema projections:
set_additional_properties_false(strict mode) overridestruetofalse.Conservativeand Gemini cleaners strip the keyword.On any of those routes the schema degrades back to the shape that fails. A string survives every projection; the new test asserts this for gemini, anthropic, openai, conservative, and both strict variants.
API Or Behavior Changes
tool_callschema declaresarguments: {"type": "string"}, and its description and thetool_searchresult text say to pass a JSON object string.unwrap_tool_callis unchanged. It already decoded a JSON-stringargumentsand still accepts an object, so models that send an object keep working.Tests
New:
tool_call_arguments_is_a_string_under_every_schema_projectionNew:
unwrap_tool_call_decodes_the_string_arguments_the_schema_asks_for. It covers the exact nested-quote and escaped-newline shape the provider returned,"{}", and refusing a non-object JSON string.cargo fmt --checkcargo clippy --all-targets -- -D warningscargo clippy --all-targets --all-features -- -D warningscargo build --all-targets(via test/clippy)cargo build --all-targets --all-features(via test/clippy)cargo test --workspace: 0 failurescargo test --all-features: 0 failuresEnd to end in OpenHuman: built on this branch and ran the real cron job on the same Sail Research route, captured through the proxy.
tool_call {name: "GMAIL_SEND_EMAIL", arguments: "{…body, is_html, recipient_email, subject, user_id}"}.Still open: the provider also intermittently empties the first call of a parallel batch (
web_search_tool {}). That call is recovered by the existing argument validation and retry, and it is not addressed here.Documentation
The doc comment on
tool_call_schemarecords the provider behaviour and why a string rather than an open object.unwrap_tool_call's doc notes that both forms are accepted. The discover README describes the function set, which is unchanged.