Skip to content

skill_creator's prompt names node_exec/npm_exec but the agent can call neither — belt narrowed out from under the prompt #6507

Description

@M3gA-Mind

Summary

skill_creator's prompt names node_exec and npm_exec, but the agent can no longer call either. It carries a twelve-tool belt and, by the fleet prompt guard's measure, names none of the tools it can actually call.

Found while triaging #6486. Verified at 0f1ecc9d28d85289ee3f87ade13156be02ffbd28.

The failing guard, and why it matters

crates/openhuman-core/src/agent/registry/agents/fleet_prompt_tests.rs:344-368:

test agent::registry::agents::fleet_prompt_tests::every_prompt_names_at_least_one_tool_it_can_call ... FAILED

assertion `left == right` failed: agents that carry tools but whose prompt names none of them
  left:  ["tools_agent", "tool_maker", "skill_creator", "critic", "archivist", "skill_setup"]
  right: ["tools_agent", "tool_maker", "critic", "archivist", "skill_setup"]

skill_creator has joined the silent set. The test's own docstring states the defect it exists to catch:

An agent with a belt must be told about at least one tool on it.

Catches a belt narrowed out from under its prompt, which otherwise reads as an improvement: the tool bytes fall and nothing else moves.

That is what has happened.

The mechanism

skill_creator carries a large named belt (registry/agents/skill_creator/agent.toml): shell, file_read, file_write, git_operations, node_exec, npm_exec, python_exec, grep, glob, list, edit, apply_patch, …

Its prompt (prompt.md) backticks exactly two tool names: node_exec and npm_exec — both on that belt.

The guard's predicate is can_call(def, name, &universe) && prompt.contains("{name}") (fleet_prompt_tests.rs:361), and can_call (:130-143) returns false when is_withheld_from(&def.id, tool) — i.e. when the tool has been withheld into a toolpack for that agent, reachable only via use_skill (tools/toolpacks/ops.rs:134-140).

So the two tools the prompt names as directly callable are no longer directly callable, and no other tool on the belt is named. The prompt and the belt have come apart.

Impact: the model is instructed to call tools it cannot call, and is told nothing about the ones it can. The failure mode is a wasted turn and a confused recovery, not an error — which is exactly why it needs a test rather than a bug report from a user.

Likely origin

skill_creator's belt and prompt last changed in a63c2b885 — PR #6436, hermes-prompt-diet (2026-09-22). A prompt-size reduction is precisely the shape the guard's docstring predicts: "the tool bytes fall and nothing else moves."

This is a recurrence, not a novelty. The same pattern has been seen before in this repo: a byte win that concealed a capability loss, visible only by diffing the name lists rather than the totals. That argues for a name-list diff as a standing step in any prompt-diet change, not just a fix here.

I have not bisected to confirm #6436 is the exact commit that flipped it — the belt/prompt last-touch is circumstantial, and the withholding could equally have been introduced by a toolpack change on the other side.

The fix that must NOT be applied

fleet_prompt_tests.rs:332 carries an allowlist:

const NAMES_NO_TOOL: &[&str] = &["tools_agent", "tool_maker", "critic", "archivist"];

Adding skill_creator to it would make the suite pass and silence the one assertion in it that is doing its job. The guard is not wrong; the prompt/belt pairing is. Flagging explicitly because that is the cheapest-looking fix and it is the wrong one.

Suggested resolution

Either:

  • update the prompt to name tools skill_creator can actually call (and describe the packed ones as reachable via use_skill, if that is the intent); or
  • un-withhold node_exec/npm_exec for skill_creator if they were packed by accident.

Which one depends on whether the withholding was deliberate — a question for whoever owns the toolpack split.

Reproduction

RUST_MIN_STACK=67108864 cargo test --manifest-path Cargo.toml \
  -p openhuman --lib -- every_prompt_names_at_least_one_tool_it_can_call

Fails deterministically on a clean checkout of main, including in isolation. Note that no CI lane on main runs this test — see #6486 and #5021 — which is why it has gone unnoticed. It sits in a sibling module of a scope ci-lite.yml does cover (agent::registry::agents::loader::), which makes the area look covered when it is not.

Activity

  1. added
    priority: p1Next. Wrong behaviour a user will hit, or a security weakness behind a condition.
    on Sep 23, 2026
  2. tinysweeper commented on Sep 23, 2026

    @tinysweeper

    skill_creator's prompt names node_exec/npm_exec but the agent can no longer call either — belt narrowed out from under the prompt

    Labelled priority: p1.

  3. added theissue type on Sep 23, 2026
  4. M3gA-Mind commented on Sep 23, 2026

    @M3gA-Mind
    CollaboratorAuthor

    Retracting this issue. There is no defect — my reproduction used the wrong feature profile, and the test passes in every profile the product ships.

    I ran cargo test -p openhuman --lib with default features. CI's full-suite lane runs it with the product feature set (test-reusable.yml:234, features from scripts/ci/product-features.sh). Under that profile:

    test agent::registry::agents::fleet_prompt_tests::every_prompt_names_at_least_one_tool_it_can_call ... ok
    

    The mechanism, and it is not what I said it was. node_exec / npm_exec are registered under #[cfg(feature = "runtime-node")] (crates/openhuman-core/src/tools/ops.rs:863-865). runtime-node is not in default:

    default = ["media", "skills", "flows", "mcp", "channels", "medulla", "http-server", "scheduler-gate", "file-logging", "modules"]
    

    but it is in scripts/ci/product-features.txt. So with default features those two tools do not exist at all — they are absent from tool_universe(), can_call is therefore false for both, and skill_creator looks like an agent whose prompt names nothing callable. Add runtime-node and they exist, can_call is true, and the guard is satisfied.

    So skill_creator's prompt does not name tools the agent cannot call, in any configuration that ships. The belt and the prompt are consistent.

    Two specific things I asserted that were wrong:

    What I got right and will keep: the NAMES_NO_TOOL allowlist is still the wrong place to silence this guard, and the guard itself is sound — it just needs the feature profile its subject requires.

    The one real remaining issue is narrow and different: the test is implicitly profile-dependent. every_prompt_names_at_least_one_tool_it_can_call passes under the product set and fails under default features, with no #[cfg] or skip explaining why. That is worth a line somewhere — a #[cfg(feature = "runtime-node")] gate, or a comment naming the required profile — so the next person running the suite locally with default features does not repeat exactly this. Filing that separately if it is wanted; it is a test-ergonomics nit, not a product bug, and it does not justify keeping this issue open.

    Closing as invalid. Context: found during #6486 triage; the corrected picture there is three failures under the product profile, not five.

  5. M3gA-Mind commented on Sep 23, 2026

    @M3gA-Mind
    CollaboratorAuthor

    Closing as invalid per the retraction above: no defect, wrong feature profile in my reproduction.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    priority: p1Next. Wrong behaviour a user will hit, or a security weakness behind a condition.

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions