Skip to content

Latest commit

 

History

History
455 lines (377 loc) · 20.6 KB

File metadata and controls

455 lines (377 loc) · 20.6 KB

Harness Tool Feature

The tool feature owns typed capabilities exposed to agents. It defines tool metadata, JSON-schema-compatible model-visible inputs, hidden runtime injection, validation, execution, retry policy, artifacts, and result formatting.

For composing which tools are visible/callable per run — the ToolSet trait, its Combined/Filtered/Prefixed/Renamed/Prepared/ ApprovalRequired/External adaptors, and how OpenHuman's MCP layer plugs into ExternalToolSet — see toolsets.md.

Source Inspiration

LangChain tool behavior is spread across core tools, v1 tool-node re-exports, agent middleware, and standard tests:

LangChain and LangGraph also distinguish model-visible tool arguments from runtime-injected values such as state, store, context, and stream writers. TinyAgents should make that distinction explicit in Rust types.

Responsibilities

  • Register named tools.
  • Validate tool names and reject duplicates.
  • Expose model-visible JSON schemas.
  • Hide injected runtime parameters from model-visible schemas.
  • Validate model-provided arguments before execution.
  • Validate provider-supplied tool calls against the tools advertised for the current turn before execution.
  • Execute tools with access to state, runtime context, stores, cancellation, and event streams.
  • Record tool lifecycle events.
  • Format tool results as canonical messages.
  • Preserve structured outputs and artifact references.
  • Classify tool errors for retry, user-visible repair, or hard failure.
  • Support serial and bounded-concurrent execution.
  • Support tool selection middleware and dynamic tool exposure.

Core Types

use tinytools::{Tool, ToolPolicy, ToolResult};

#[async_trait]
impl Tool for Weather {
    fn name(&self) -> &str { "weather" }
    fn description(&self) -> &str { "Look up weather for a city." }
    fn parameters_schema(&self) -> serde_json::Value { serde_json::json!({}) }
    async fn execute(&self, args: serde_json::Value) -> anyhow::Result<ToolResult> {
        Ok(ToolResult::success("sunny"))
    }
    fn policy(&self) -> ToolPolicy { ToolPolicy::classified() }
}

tinytools owns the tool trait, result blocks, declarative policy, and tool specification. tinyinference_llm::tool owns model-facing ToolCall, ToolSchema, and ToolFormat. The harness owns only registration, dispatch, and host policy enforcement.

Tool Names

Tool names should default to ASCII snake_case. The registry should reject:

  • empty names
  • duplicate names
  • names with spaces
  • names that exceed provider-safe length limits
  • names requiring provider-specific escaping

The registry may support provider-specific aliases, but canonical events and stores should use the TinyAgents tool name.

Schema Rules

The model-visible input schema must include only arguments the model may choose. Hidden runtime values include:

  • current run context
  • thread id and run id
  • state references
  • store handles
  • event emitters
  • stream writers
  • cancellation handles
  • secrets or provider clients

Hidden values must never appear in the tool schema sent to a model. This avoids teaching the model about internal implementation details and prevents accidental secret exposure.

The local execution boundary validates the schema subset TinyAgents relies on for fail-closed tool dispatch:

  • primitive and compound type checks
  • object properties
  • required object fields
  • additionalProperties: false
  • array items
  • exact-value enum

Richer provider-facing JSON Schema keywords may still be present; the harness passes them through to providers but only enforces the subset above locally. RunPolicy::tool_schemas can opt a run into a provider projection and byte budgets (see tool-discovery.md).

Exposure and Discovery

ToolExposure::Direct tools go on the initial wire request. Deferred tools are indexed into a per-run catalogue and reached through the intrinsic tool_search / tool_call bridge; searched matches are promoted into later typed requests. Hidden tools are host-only. See tool-discovery.md.

Tool Call Formats

ToolSchema carries a format: ToolFormat field so a tool definition can state how it should be shown to a model. The execution boundary is still one typed shape: after parsing a model emission, the harness invokes tools through ToolCall { id, name, arguments }, where arguments is serde_json::Value. That keeps validation, middleware, replay, and provider normalization stable even when the model-facing syntax changes.

TinyAgents supports three model-facing tool formats:

  • ToolFormat::Json — JSON/function-call format. This is the default and is omitted during serialization for backward compatibility. Providers with native function/tool calling, such as OpenAI Chat Completions, can map this directly to their native tool declaration shape.
  • ToolFormat::Xml — XML tag format. A renderer may expose the same tool as <tool_name><field>value</field></tool_name>. The parser must normalize the emitted tag body back into JSON arguments before schema validation.
  • ToolFormat::PType { parameters } — parametric p-type format. This is a compact ordered-parameter syntax for token-sensitive prompts, for example search("rust agents", 5). parameters records the ordered field names that map positional values back into the JSON argument object.

Example:

use serde_json::json;
use tinyinference_llm::tool::{ToolFormat, ToolSchema};

let json_tool = ToolSchema::new(
    "weather",
    "Look up weather for a city.",
    json!({
        "type": "object",
        "required": ["city"],
        "properties": { "city": { "type": "string" } }
    }),
);

let xml_tool = json_tool.clone().with_format(ToolFormat::Xml);

let ptype_tool = ToolSchema::new(
    "search",
    "Search documents.",
    json!({
        "type": "object",
        "required": ["query"],
        "properties": {
            "query": { "type": "string" },
            "limit": { "type": "integer" }
        }
    }),
)
.with_format(ToolFormat::PType {
    parameters: vec!["query".to_string(), "limit".to_string()],
});

Provider adapters should treat ToolFormat as a capability-aware rendering preference:

  • If the provider has native JSON/function calling, send ToolFormat::Json tools as native tool declarations.
  • If the provider does not support the requested format natively, render the declaration into prompt text and parse the model's emitted call back into ToolCall.
  • If a provider only accepts JSON tool declarations, it may fall back to the JSON schema while preserving ToolSchema::format in harness metadata.

Prompt-Guided Message Shape

A model without a native tool channel is driven through its own Jinja chat template by the serving runtime (LM Studio, llama.cpp, Ollama), so the outgoing message list has to satisfy that template, not just the wire schema. Three helpers in tinyinference_llm::prompt_tools normalize it, and both the OpenAI-compatible adapter (for a profile with tool_calling = false, or after a "tools unsupported" 400) and the harness (for a forced Xml, Pformat, Python, or Typescript dialect) apply them:

  • coalesce_tool_results renders assistant tool_calls back into <tool_call> text and folds consecutive tool-role results into one [Tool results] user turn under the <tool_result id="…"> envelope — the tool role and structured tool_calls field are not consumable by these models, and a result body cannot forge a closing tag.
  • ensure_resolvable_user_turn guarantees the list contains a user turn the template can resolve as "the user query", inserting one after any leading system turns when none is present. Qwen 3's template raises No user query found in messages. otherwise, and a prompt-guided tool loop reaches that state legitimately once the real user turn ages out of the window.
  • with_tool_instructions appends the protocol block and catalogue to the system prompt.

The answer is read back through tinytools_agent::parse; see tool-dialect.md.

Execution Lifecycle

  1. Check cancellation, wall-clock deadline, and tool-call limits.
  2. Run before_tool middleware, allowing policy middleware to reject or adjust the pending call.
  3. Validate the final tool name exists.
  4. Validate final arguments against the input schema.
  5. Emit tool.started.
  6. Execute the tool with ToolRuntime.
  7. Run after_tool middleware.
  8. Format result into a ToolMessage.
  9. Persist artifacts if configured.
  10. Emit tool.completed or tool.failed.

Validation failures should produce a model-consumable error message when the agent loop can recover, and a hard error when policy forbids repair.

Provider-supplied tool calls must fail closed:

  • unknown tool names are not executed
  • malformed JSON arguments are not replaced with empty defaults for side-effecting tools
  • tool call ids are preserved in error tool messages
  • allowlist violations emit events and append repairable tool-result messages only when the agent loop policy allows recovery

A hosted run's tool allow-list is fail-closed by default. The resolved AgentDefinition.tools list is collapsed to Option<HashSet<String>> at the host boundary: a declared, non-empty list is enforced by plain membership, and an empty or absent list means "the definition declared nothing" rather than "unrestricted" — under HostCapabilities::fail_closed_tool_allowlist (default true), that denies every registered tool. A host that relied on the old fail-open behavior (empty list = every tool) must opt back in explicitly via HostCapabilities::with_legacy_unrestricted_tool_allowlist. Explicit-model (non-hosted) runs have no allow-list concept and are unaffected.

Unknown-tool recovery

When the model calls a tool that is not registered, the agent loop's behavior is governed by RunPolicy::unknown_tool: UnknownToolPolicy (crates/tinyagents-harness/src/runtime/types.rs):

  • UnknownToolPolicy::Fail — abort the run with TinyAgentsError::ToolNotFound(name). No longer the default (see below); still available for callers that want a hard stop.
  • UnknownToolPolicy::ReturnToolError (default) — inject a tool-error result (naming the requested tool, echoing its arguments, and listing the registered tools) back into the transcript and continue, letting the model retry with a valid tool.
  • UnknownToolPolicy::Rewrite { tool_name } — retarget the unknown call to a fixed compatibility tool and retry the lookup once; if that target is also unregistered, fall back to ReturnToolError behavior.

Each recovery still consumes a tool-call budget slot, so RunLimits::max_tool_calls bounds any unknown-tool loop. Every recovery emits AgentEvent::UnknownToolCall { call_id, requested_name, arguments, recovery } — the original arguments are preserved verbatim so repair/analysis middleware can re-target or replay the intended call, and recovery is a label such as "tool_error" or "rewrite:lookup".

use tinyagents_harness::runtime::{AgentHarness, RunPolicy, UnknownToolPolicy};

let mut harness: AgentHarness<()> = AgentHarness::new();
// ... register a model whose first turn calls the unregistered `missing` ...
harness.with_policy(RunPolicy {
    unknown_tool: UnknownToolPolicy::ReturnToolError,
    ..RunPolicy::default()
});

let run = harness
    .invoke_in_context(&(), ctx, vec![Message::user("go")])
    .await?;
// The injected repair message names the requested tool for the model.
assert!(run.messages.iter().any(|m| m.text().contains("unknown tool `missing`")));
// A single UnknownToolCall event was recorded with recovery == "tool_error".

Invalid tool-argument recovery

Two distinct failures can affect a provider-supplied call's arguments, and they are handled separately:

  • Schema-invalid (well-formed JSON that violates the tool's input schema) is governed by RunPolicy::invalid_args: InvalidArgsPolicy. ReturnToolError (the default) injects a repairable tool-error message (carrying the validation detail and the expected schema) and continues; Fail aborts the turn and is no longer the default. NormalizeThenReturnToolError first repairs common object-schema transport shapes (a JSON object encoded as a string, including markdown fences, or a non-object for an object schema with no required fields), then returns any remaining validation failure as a tool error. Schemas that accept top-level primitives or arrays are left untouched.
  • Unparseable (malformed JSON the provider could not parse into arguments at all) is surfaced by the provider as a ToolCall with invalid: Some(reason) and the raw string preserved in arguments. Small local models (Ollama, LM Studio, llama.cpp, vLLM) emit this occasionally. Before giving up, admission first tries relaxed_json::recover_relaxed_object on the raw string — conservative, meaning-preserving repairs for the shapes those gateways actually produce (unquoted object keys, redundant wrapping braces, leaked chat-template quote tokens; see that module's doc comment). On success the call's invalid flag is cleared, its arguments become the repaired object, AgentEvent::InvalidToolArgs { recovery: "repaired" } is emitted, and the call proceeds through normal (schema) validation as if the provider had sent it clean. Only when the repair also fails does the agent loop fall back to its always-on recovery — independent of InvalidArgsPolicy, since an unparseable payload is a transport-level defect, not a schema violation — injecting the parse reason back to the model as an error tool result so it can retry. That fallback recovery emits AgentEvent::InvalidToolArgs { call_id, tool_name, arguments, error, recovery: "tool_error" } and consumes one tool-call budget slot, so RunLimits::max_tool_calls bounds the retry loop. Because the call always resolves, a malformed argument blob can never become a never-resolving tool call that stalls the loop. See the OpenAI provider README for how the wire parser produces these invalid calls.

Tool policy enforcement

Beyond the model-visible ToolSchema, each tool advertises a structured, serializable ToolPolicy (crates/tinyagents-harness/src/tool/types.rs) via Tool::policy(). The default is unclassified (classified == false), so strict enforcement can fail closed on any tool that has not declared its safety profile.

pub struct ToolPolicy {
    pub classified: bool,
    pub side_effects: ToolSideEffects, // read_only, writes_files, network,
                                       // installs_dependencies, destructive,
                                       // external_service, payment
    pub runtime: ToolRuntime,          // sandbox: SandboxMode, max_result_bytes,
                                       // timeout_ms, max_retries, idempotent, …
    pub access: ToolAccess,            // workspace: WorkspaceAccess, trusted_roots,
                                       // credentials, approval_required, background_safe
}

Build a policy fluently: ToolPolicy::read_only(), ToolPolicy::classified(), then .with_side_effects(…), .with_runtime(…), .with_access(…). A registry snapshot for enforcement comes from ToolRegistry::policies().

Enforcement itself lives in ToolPolicyMiddleware; see Tool policy enforcement in the middleware feature for the exposure/execution hooks and the enforcement builders (require_sandbox, require_approval, enforce_result_bytes, strict, deny_side_effects, require_classification, require_background_safe). The require_sandbox gate reads the run's tinytools::WorkspaceDescriptor to decide whether a SandboxMode::Required tool may run. The host attaches that descriptor to RunContext, commonly from around-agent middleware.

Deferred tool calls: approval and external execution (A2)

A call can leave the loop without a result, in three ways:

Trigger Where Lands in
ToolPolicy.access.approval_required on the tool's declared policy admission, after schema validation DeferredToolRequests::approvals
Err(TinyAgentsError::ApprovalRequired { metadata }) / Err(TinyAgentsError::CallDeferred { metadata }) from Tool::execute or from a before_tool middleware execution / admission approvals / calls, with metadata keyed by call id
ToolRegistry::register_external(ToolSchema) (or AgentHarness::register_external_tool) — a schema-only tool the harness never runs admission DeferredToolRequests::calls

The loop finishes every other call in the batch, appends their results, emits AgentEvent::ToolDeferred { call_id, reason } per deferred call, and exits with AgentRun::deferred = Some(DeferredToolRequests { calls, approvals, metadata }). The assistant's tool-call row stays on the transcript; only the deferred ids lack a tool-result row. A deferred call is not counted against max_tool_calls until it actually runs.

Resolve it with a DeferredToolResults and resume:

let pending = run.deferred.clone().unwrap();
let results = DeferredToolResults::new()
    .approve("call-1")                                 // run with the model's args
    .approve_with_args("call-2", json!({"path": "x"})) // run with edited args
    .deny("call-3", "operator refused")                // tool-error result, no run
    .respond("call-4", ToolResult::success("done"));   // host ran an external tool
assert!(pending.remaining(&results).is_empty());
let run = harness.resume_deferred(&state, ctx, run.messages, results).await?;

ToolApprovalDecision::{Approve, ApproveWithArgs(Value), Deny { message }} and DeferredCallResult::{Result(ToolResult), Retry(String), Failed(String)} are the per-call vocabularies; DeferredToolRequests::remaining(&results) lists what is still unresolved and approve_all() builds a blanket approval. On resume an approved call is re-admitted through the normal pipeline (before_tool, validation, host authorization, the wrap onion) with RunContext::is_call_approved(call_id) set so neither the policy check nor an approval middleware defers it again; a denial and a host-supplied result are answered through the same fold as a recovery (they emit ToolApproved/ToolDenied plus the usual ToolStarted/ToolCompleted pair, run after_tool, and never appear in executed_tools).

Two ways to avoid surfacing the pause at all: register a DeferredToolHandler on the harness (with_deferred_tool_handler) and the loop resolves the batch inline and keeps going; or give HumanApprovalMiddleware::with_approval_outcome a callback returning ApprovalOutcome::{Allow, Deny(msg), Defer} — Defer is exactly the deferral above, Deny answers the model without an interrupt.

A tool that raises ApprovalRequired from inside execute cannot currently see that it was approved: ToolExecutionContext now carries the call_id (B1, below) but no approval flag, so an approved re-execution of such a tool defers again and the loop surfaces it rather than spinning. Prefer the policy flag or the middleware for approval gates.

Execution context and rich returns (B1/B2)

What a tool sees on its ToolExecutionContext (call_id, store, typed state::<S>(), custom() events) and what the loop does with a result's follow_up and metadata are documented in tool-context.md. In one line: follow-up content becomes a user message after the batch's last tool row; metadata reaches ToolCompleted and AgentRun::tool_metadata and never the transcript.

Safety Metadata

Tools should declare safety metadata:

  • read-only versus mutating
  • idempotent versus non-idempotent
  • local-only versus networked
  • filesystem access
  • shell/process access
  • payment or external spend
  • requires user confirmation
  • allowed workspace root
  • redaction policy

Middleware can use this metadata to enforce confirmation, sandboxing, allowlist, or human-in-the-loop policies.

Tool-Effect Ledger And Replay (B5)

Crash-safe bookkeeping of tool-call side effects — a durable started row written before a tool executes, settled once it completes — plus the resume logic that decides whether an interrupted call is safe to re-run. See tool-effects.md.