Skip to content

Commit 305c58a

Browse files
authored
fix: reduce agent loop hallucination and improve tool call reliability (tinyhumansai#156)
* fix: reduce agent loop hallucination and improve tool call reliability - Strengthen tool-use instructions with explicit anti-hallucination rules: "NEVER narrate tool use without emitting tags", "use exact tool names", "only respond without tool call when no tool is needed" - Wire context guard into tool loop: check utilization before each LLM call, abort on context exhaustion (>95% with circuit breaker tripped) - Add 120-second timeout on tool execution to prevent hangs - Add debug/warn/error logging at all loop boundaries: LLM request, response (with token counts), tool call parsing, tool execution, unknown tools, timeouts, and final response Closes tinyhumansai#144 * fix: implement chat(ChatRequest) on ReliableProvider and add bracket tool call parser Root cause: ReliableProvider did not implement the chat(ChatRequest) trait method. The agent loop called provider.chat() which fell through to the default trait implementation — this used chat_with_history() which strips native tool support and sends raw tool-role messages without the required assistant tool_calls, causing the backend Jinja template to reject the request with "Message has tool role, but there was no previous assistant message with a tool call!" Fixes: - Add chat(ChatRequest) to ReliableProvider with full retry/failover logic, matching the existing chat_with_system/chat_with_history implementations. Delegates to inner provider's chat() which properly converts messages to native OpenAI format with tool_calls. - Add [TOOL_CALL]/[/TOOL_CALL] bracket format to the tool call parser (parse.rs) — some models emit this format instead of <tool_call> XML. - Add parse_bracket_tool_call() for the pseudo-syntax format: {tool => "name", args => { --key "value" }} Verified with real staging backend (agentic-v1 model): - Shell tool calls execute successfully - File read tool calls return real content - Knowledge questions return without tool calls - No Jinja template errors Closes tinyhumansai#144
1 parent 64eb513 commit 305c58a

5 files changed

Lines changed: 302 additions & 15 deletions

File tree

‎src/openhuman/agent/loop_/context_guard.rs‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -7,7 +7,7 @@
77
use crate::openhuman::providers::UsageInfo;
88

99
/// Threshold (0.0–1.0) at which auto-compaction is triggered.
10-
const COMPACTION_TRIGGER_THRESHOLD: f64 = 0.90;
10+
pub(crate) const COMPACTION_TRIGGER_THRESHOLD: f64 = 0.90;
1111

1212
/// Threshold above which, if compaction is disabled, the guard returns an error.
1313
const HARD_LIMIT_THRESHOLD: f64 = 0.95;

‎src/openhuman/agent/loop_/instructions.rs‎

Lines changed: 27 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -8,14 +8,35 @@ pub(crate) fn build_tool_instructions(tools_registry: &[Box<dyn Tool>]) -> Strin
88
instructions.push_str("\n## Tool Use Protocol\n\n");
99
instructions.push_str("To use a tool, wrap a JSON object in <tool_call></tool_call> tags:\n\n");
1010
instructions.push_str("```\n<tool_call>\n{\"name\": \"tool_name\", \"arguments\": {\"param\": \"value\"}}\n</tool_call>\n```\n\n");
11+
12+
// Explicit anti-hallucination rules
13+
instructions.push_str("### Rules (MUST follow)\n\n");
1114
instructions.push_str(
12-
"CRITICAL: Output actual <tool_call> tags—never describe steps or give examples.\n\n",
15+
"1. **ALWAYS use <tool_call> tags** when a task requires a tool. \
16+
NEVER narrate what you would do — emit the actual tags.\n",
1317
);
14-
instructions.push_str("Example: User says \"what's the date?\". You MUST respond with:\n<tool_call>\n{\"name\":\"shell\",\"arguments\":{\"command\":\"date\"}}\n</tool_call>\n\n");
15-
instructions.push_str("You may use multiple tool calls in a single response. ");
16-
instructions.push_str("After tool execution, results appear in <tool_result> tags. ");
17-
instructions
18-
.push_str("Continue reasoning with the results until you can give a final answer.\n\n");
18+
instructions.push_str(
19+
"2. **NEVER describe a tool call in prose** (e.g. \"I'll run ls\" or \
20+
\"Let me check the file\") without also emitting the <tool_call> tags. \
21+
If you mention a tool, you must call it.\n",
22+
);
23+
instructions.push_str(
24+
"3. **Use the exact tool names** listed below. \
25+
Do not invent tool names that are not in the list.\n",
26+
);
27+
instructions.push_str("4. You may use **multiple tool calls** in a single response.\n");
28+
instructions.push_str(
29+
"5. After tool execution, results appear in <tool_result> tags. \
30+
Continue reasoning with the results until you can give a final answer.\n",
31+
);
32+
instructions.push_str(
33+
"6. Only respond **without** a tool call when the answer requires \
34+
no tool (e.g. general knowledge, math, or conversation).\n\n",
35+
);
36+
37+
instructions.push_str("### Example\n\n");
38+
instructions.push_str("User: \"what's the date?\"\nCorrect response:\n<tool_call>\n{\"name\":\"shell\",\"arguments\":{\"command\":\"date\"}}\n</tool_call>\n\n");
39+
1940
instructions.push_str("### Available Tools\n\n");
2041

2142
for tool in tools_registry {

‎src/openhuman/agent/loop_/parse.rs‎

Lines changed: 57 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -82,7 +82,51 @@ pub(crate) fn parse_tool_calls_from_json_value(value: &serde_json::Value) -> Vec
8282
calls
8383
}
8484

85-
const TOOL_CALL_OPEN_TAGS: [&str; 4] = ["<tool_call>", "<toolcall>", "<tool-call>", "<invoke>"];
85+
/// Parse bracket-style pseudo-syntax emitted by some models:
86+
/// `{tool => "shell", args => { --command "ls" }}`
87+
///
88+
/// Extracts tool name and converts `--key "value"` args into a JSON object.
89+
fn parse_bracket_tool_call(inner: &str) -> Option<ParsedToolCall> {
90+
static BRACKET_TOOL_RE: LazyLock<Regex> = LazyLock::new(|| {
91+
Regex::new(r#"(?s)\{?\s*tool\s*=>\s*"([^"]+)"\s*,\s*args\s*=>\s*\{(.*?)\}\s*\}?"#).unwrap()
92+
});
93+
static ARG_RE: LazyLock<Regex> = LazyLock::new(|| Regex::new(r#"--(\w+)\s+"([^"]*)"#).unwrap());
94+
95+
let cap = BRACKET_TOOL_RE.captures(inner.trim())?;
96+
let name = cap.get(1)?.as_str().to_string();
97+
let args_body = cap.get(2)?.as_str();
98+
99+
let mut args = serde_json::Map::new();
100+
for arg_cap in ARG_RE.captures_iter(args_body) {
101+
let key = arg_cap.get(1)?.as_str().to_string();
102+
let value = arg_cap.get(2)?.as_str().to_string();
103+
args.insert(key, serde_json::Value::String(value));
104+
}
105+
106+
if name.is_empty() {
107+
return None;
108+
}
109+
110+
tracing::debug!(
111+
tool = name.as_str(),
112+
parse_mode = "bracket_syntax",
113+
"[parse] parsed bracket-style tool call"
114+
);
115+
116+
Some(ParsedToolCall {
117+
name,
118+
arguments: serde_json::Value::Object(args),
119+
})
120+
}
121+
122+
const TOOL_CALL_OPEN_TAGS: [&str; 6] = [
123+
"<tool_call>",
124+
"<toolcall>",
125+
"<tool-call>",
126+
"<invoke>",
127+
"[TOOL_CALL]",
128+
"[tool_call]",
129+
];
86130

87131
pub(crate) fn find_first_tag<'a>(haystack: &str, tags: &'a [&'a str]) -> Option<(usize, &'a str)> {
88132
tags.iter()
@@ -96,6 +140,8 @@ pub(crate) fn matching_tool_call_close_tag(open_tag: &str) -> Option<&'static st
96140
"<toolcall>" => Some("</toolcall>"),
97141
"<tool-call>" => Some("</tool-call>"),
98142
"<invoke>" => Some("</invoke>"),
143+
"[TOOL_CALL]" => Some("[/TOOL_CALL]"),
144+
"[tool_call]" => Some("[/tool_call]"),
99145
_ => None,
100146
}
101147
}
@@ -376,7 +422,16 @@ pub(crate) fn parse_tool_calls(response: &str) -> (String, Vec<ParsedToolCall>)
376422
}
377423

378424
if !parsed_any {
379-
tracing::warn!("Malformed <tool_call> JSON: expected tool-call object in tag body");
425+
// Try parsing bracket pseudo-syntax:
426+
// {tool => "name", args => { --key "value" }}
427+
if let Some(call) = parse_bracket_tool_call(inner) {
428+
calls.push(call);
429+
parsed_any = true;
430+
}
431+
}
432+
433+
if !parsed_any {
434+
tracing::warn!("Malformed tool_call body: could not parse tag content as JSON or bracket syntax");
380435
}
381436

382437
remaining = &after_open[close_idx + close_tag.len()..];

‎src/openhuman/agent/loop_/tool_loop.rs‎

Lines changed: 104 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -6,6 +6,7 @@ use anyhow::Result;
66
use std::fmt::Write as _;
77
use std::io::Write as _;
88

9+
use super::context_guard::{ContextCheckResult, ContextGuard};
910
use super::credentials::scrub_credentials;
1011
use super::parse::{
1112
build_native_assistant_history, find_tool, parse_structured_tool_calls, parse_tool_calls,
@@ -77,7 +78,35 @@ pub(crate) async fn run_tool_call_loop(
7778
tools_registry.iter().map(|tool| tool.spec()).collect();
7879
let use_native_tools = provider.supports_native_tools() && !tool_specs.is_empty();
7980

80-
for _iteration in 0..max_iterations {
81+
let mut context_guard = ContextGuard::new();
82+
83+
for iteration in 0..max_iterations {
84+
// ── Context guard: check utilization before each LLM call ──
85+
match context_guard.check() {
86+
ContextCheckResult::Ok => {}
87+
ContextCheckResult::CompactionNeeded => {
88+
tracing::warn!(
89+
iteration,
90+
"[agent_loop] context guard: compaction needed (>{:.0}% full)",
91+
super::context_guard::COMPACTION_TRIGGER_THRESHOLD * 100.0
92+
);
93+
// Compaction is handled by history management upstream;
94+
// log and continue so the caller can act on it.
95+
}
96+
ContextCheckResult::ContextExhausted {
97+
utilization_pct,
98+
reason,
99+
} => {
100+
tracing::error!(
101+
iteration,
102+
utilization_pct,
103+
"[agent_loop] context exhausted, aborting: {reason}"
104+
);
105+
anyhow::bail!("Context window exhausted ({utilization_pct}% full): {reason}");
106+
}
107+
}
108+
109+
tracing::debug!(iteration, "[agent_loop] sending LLM request");
81110
let image_marker_count = multimodal::count_image_markers(history);
82111
if image_marker_count > 0 && !provider.supports_vision() {
83112
return Err(ProviderCapabilityError {
@@ -115,6 +144,23 @@ pub(crate) async fn run_tool_call_loop(
115144
.await
116145
{
117146
Ok(resp) => {
147+
// Update context guard with token usage from this response.
148+
if let Some(ref usage) = resp.usage {
149+
context_guard.update_usage(usage);
150+
tracing::debug!(
151+
iteration,
152+
input_tokens = usage.input_tokens,
153+
output_tokens = usage.output_tokens,
154+
context_window = usage.context_window,
155+
"[agent_loop] LLM response received"
156+
);
157+
} else {
158+
tracing::debug!(
159+
iteration,
160+
"[agent_loop] LLM response received (no usage info)"
161+
);
162+
}
163+
118164
let response_text = resp.text_or_empty().to_string();
119165
let mut calls = parse_structured_tool_calls(&resp.tool_calls);
120166
let mut parsed_text = String::new();
@@ -127,6 +173,13 @@ pub(crate) async fn run_tool_call_loop(
127173
calls = fallback_calls;
128174
}
129175

176+
tracing::debug!(
177+
iteration,
178+
native_tool_calls = resp.tool_calls.len(),
179+
parsed_tool_calls = calls.len(),
180+
"[agent_loop] tool calls parsed"
181+
);
182+
130183
// Preserve native tool call IDs in assistant history so role=tool
131184
// follow-up messages can reference the exact call id.
132185
let assistant_history_content = if resp.tool_calls.is_empty() {
@@ -156,6 +209,10 @@ pub(crate) async fn run_tool_call_loop(
156209
};
157210

158211
if tool_calls.is_empty() {
212+
tracing::debug!(
213+
iteration,
214+
"[agent_loop] no tool calls — returning final response"
215+
);
159216
// No tool calls — this is the final response.
160217
// If a streaming sender is provided, relay the text in small chunks
161218
// so the channel can progressively update the draft message.
@@ -221,20 +278,62 @@ pub(crate) async fn run_tool_call_loop(
221278
}
222279
}
223280

281+
tracing::debug!(
282+
iteration,
283+
tool = call.name.as_str(),
284+
"[agent_loop] executing tool"
285+
);
286+
224287
let result = if let Some(tool) = find_tool(tools_registry, &call.name) {
225-
match tool.execute(call.arguments.clone()).await {
226-
Ok(r) => {
288+
// Execute with a 120-second timeout to prevent hangs.
289+
match tokio::time::timeout(
290+
std::time::Duration::from_secs(120),
291+
tool.execute(call.arguments.clone()),
292+
)
293+
.await
294+
{
295+
Ok(Ok(r)) => {
227296
if r.success {
297+
tracing::debug!(
298+
iteration,
299+
tool = call.name.as_str(),
300+
output_len = r.output.len(),
301+
"[agent_loop] tool succeeded"
302+
);
228303
scrub_credentials(&r.output)
229304
} else {
230-
format!("Error: {}", r.error.unwrap_or(r.output))
305+
let err_msg = r.error.unwrap_or(r.output);
306+
tracing::warn!(
307+
iteration,
308+
tool = call.name.as_str(),
309+
"[agent_loop] tool returned error: {err_msg}"
310+
);
311+
format!("Error: {err_msg}")
231312
}
232313
}
233-
Err(e) => {
314+
Ok(Err(e)) => {
315+
tracing::error!(
316+
iteration,
317+
tool = call.name.as_str(),
318+
"[agent_loop] tool execution failed: {e}"
319+
);
234320
format!("Error executing {}: {e}", call.name)
235321
}
322+
Err(_) => {
323+
tracing::error!(
324+
iteration,
325+
tool = call.name.as_str(),
326+
"[agent_loop] tool execution timed out after 120s"
327+
);
328+
format!("Error: tool '{}' timed out after 120 seconds", call.name)
329+
}
236330
}
237331
} else {
332+
tracing::warn!(
333+
iteration,
334+
tool = call.name.as_str(),
335+
"[agent_loop] unknown tool requested"
336+
);
238337
format!("Unknown tool: {}", call.name)
239338
};
240339

0 commit comments

Comments
 (0)