Repository navigation
skill_creator's prompt names node_exec/npm_exec but the agent can call neither — belt narrowed out from under the prompt #6507
Description
Activity
- addedpriority: p1Next. Wrong behaviour a user will hit, or a security weakness behind a condition.Next. Wrong behaviour a user will hit, or a security weakness behind a condition.
on Sep 23, 2026 skill_creator's prompt names node_exec/npm_exec but the agent can no longer call either — belt narrowed out from under the prompt
Labelled
priority: p1.Retracting this issue. There is no defect — my reproduction used the wrong feature profile, and the test passes in every profile the product ships.
I ran
cargo test -p openhuman --libwith default features. CI's full-suite lane runs it with the product feature set (test-reusable.yml:234, features fromscripts/ci/product-features.sh). Under that profile:test agent::registry::agents::fleet_prompt_tests::every_prompt_names_at_least_one_tool_it_can_call ... okThe mechanism, and it is not what I said it was.
node_exec/npm_execare registered under#[cfg(feature = "runtime-node")](crates/openhuman-core/src/tools/ops.rs:863-865).runtime-nodeis not indefault:default = ["media", "skills", "flows", "mcp", "channels", "medulla", "http-server", "scheduler-gate", "file-logging", "modules"]but it is in
scripts/ci/product-features.txt. So with default features those two tools do not exist at all — they are absent fromtool_universe(),can_callis therefore false for both, andskill_creatorlooks like an agent whose prompt names nothing callable. Addruntime-nodeand they exist,can_callis true, and the guard is satisfied.So
skill_creator's prompt does not name tools the agent cannot call, in any configuration that ships. The belt and the prompt are consistent.Two specific things I asserted that were wrong:
- The cause. I attributed it to
is_withheld_from— tools packed behinduse_skill(toolpacks/ops.rs:134-140). That was wrong; it is feature registration, not toolpack withholding. I reasoned from a plausible mechanism in the right area instead of confirming which one applied. - The origin. I pointed at
a63c2b885/ perf(agent): orchestrator prompt diet, lead-in before tool calls, plan review off the chat belt #6436hermes-prompt-diet, on the circumstantial basis that it last touched the belt and prompt. That attribution is withdrawn — perf(agent): orchestrator prompt diet, lead-in before tool calls, plan review off the chat belt #6436 did not cause this, because there is nothing to cause. Apologies to anyone who went to read that diff.
What I got right and will keep: the
NAMES_NO_TOOLallowlist is still the wrong place to silence this guard, and the guard itself is sound — it just needs the feature profile its subject requires.The one real remaining issue is narrow and different: the test is implicitly profile-dependent.
every_prompt_names_at_least_one_tool_it_can_callpasses under the product set and fails under default features, with no#[cfg]or skip explaining why. That is worth a line somewhere — a#[cfg(feature = "runtime-node")]gate, or a comment naming the required profile — so the next person running the suite locally with default features does not repeat exactly this. Filing that separately if it is wanted; it is a test-ergonomics nit, not a product bug, and it does not justify keeping this issue open.Closing as invalid. Context: found during #6486 triage; the corrected picture there is three failures under the product profile, not five.
- The cause. I attributed it to
Closing as invalid per the retraction above: no defect, wrong feature profile in my reproduction.
- added a commit that references this issue
on Sep 23, 2026
Summary
skill_creator's prompt namesnode_execandnpm_exec, but the agent can no longer call either. It carries a twelve-tool belt and, by the fleet prompt guard's measure, names none of the tools it can actually call.Found while triaging #6486. Verified at
0f1ecc9d28d85289ee3f87ade13156be02ffbd28.The failing guard, and why it matters
crates/openhuman-core/src/agent/registry/agents/fleet_prompt_tests.rs:344-368:skill_creatorhas joined the silent set. The test's own docstring states the defect it exists to catch:That is what has happened.
The mechanism
skill_creatorcarries a large named belt (registry/agents/skill_creator/agent.toml):shell,file_read,file_write,git_operations,node_exec,npm_exec,python_exec,grep,glob,list,edit,apply_patch, …Its prompt (
prompt.md) backticks exactly two tool names:node_execandnpm_exec— both on that belt.The guard's predicate is
can_call(def, name, &universe) && prompt.contains("{name}")(fleet_prompt_tests.rs:361), andcan_call(:130-143) returns false whenis_withheld_from(&def.id, tool)— i.e. when the tool has been withheld into a toolpack for that agent, reachable only viause_skill(tools/toolpacks/ops.rs:134-140).So the two tools the prompt names as directly callable are no longer directly callable, and no other tool on the belt is named. The prompt and the belt have come apart.
Impact: the model is instructed to call tools it cannot call, and is told nothing about the ones it can. The failure mode is a wasted turn and a confused recovery, not an error — which is exactly why it needs a test rather than a bug report from a user.
Likely origin
skill_creator's belt and prompt last changed ina63c2b885— PR #6436,hermes-prompt-diet(2026-09-22). A prompt-size reduction is precisely the shape the guard's docstring predicts: "the tool bytes fall and nothing else moves."This is a recurrence, not a novelty. The same pattern has been seen before in this repo: a byte win that concealed a capability loss, visible only by diffing the name lists rather than the totals. That argues for a name-list diff as a standing step in any prompt-diet change, not just a fix here.
I have not bisected to confirm #6436 is the exact commit that flipped it — the belt/prompt last-touch is circumstantial, and the withholding could equally have been introduced by a toolpack change on the other side.
The fix that must NOT be applied
fleet_prompt_tests.rs:332carries an allowlist:Adding
skill_creatorto it would make the suite pass and silence the one assertion in it that is doing its job. The guard is not wrong; the prompt/belt pairing is. Flagging explicitly because that is the cheapest-looking fix and it is the wrong one.Suggested resolution
Either:
skill_creatorcan actually call (and describe the packed ones as reachable viause_skill, if that is the intent); ornode_exec/npm_execforskill_creatorif they were packed by accident.Which one depends on whether the withholding was deliberate — a question for whoever owns the toolpack split.
Reproduction
Fails deterministically on a clean checkout of
main, including in isolation. Note that no CI lane onmainruns this test — see #6486 and #5021 — which is why it has gone unnoticed. It sits in a sibling module of a scopeci-lite.ymldoes cover (agent::registry::agents::loader::), which makes the area look covered when it is not.