Skip to content

Five openhuman --lib tests fail on main, hidden behind the Rust Quality step-7 curtain #6486

Description

@M3gA-Mind

Summary

Five cargo test -p openhuman --lib tests fail on main. They are invisible today because Rust Quality dies at an earlier step and the test steps report skipped — so landing #6462 (which unblocks steps 8–15) will surface them as a fresh red, and that red will not be #6462's fault.

Problem

Proven pre-existing at 0f1ecc9d2, not inferred: a warm worktree was detached to that commit, the absence of any unrelated change was confirmed, and the five were run by exact name. Result: 0 passed; 5 failed.

agent::harness::tool_calling::harness_tool_call_parsing_tests::parse_tool_calls_recovers_mismatched_close_tag
agent::tools::spawn_async_subagent::tests::errors_clearly_when_no_parent_thread_for_delivery
agent::subagent_host::subagent_render_tests::render_subagent_system_prompt_honors_identity_safety_and_skills_flags
agent::registry::agents::fleet_prompt_tests::every_prompt_names_at_least_one_tool_it_can_call
agent::registry::agents::orchestrator::session_routing_tests::the_withheld_block_renders_for_a_renamed_session_with_a_filter

All five are prompt/parsing tests. None touches usage, the codec, or any area currently under change.

Why they are invisible. Rust Quality currently fails at step 7 ("Enforce agent runtime ownership boundary"), and every later step in that job reports skipped — the lib test steps among them. Both coverage lanes are additionally gated on rust-quality succeeding (ci-lite.yml:750, :898), so they do not run either. The suite has therefore not been executed on main for days.

Consequence worth stating plainly: #6462 fixes step 7. The moment it lands, steps 8–15 become reachable for the first time in days and these five will appear. That is the gate working, not a regression introduced by #6462 — exactly the same dynamic already recorded on #6423, where the prompt-prefix guard's six budget errors are reachable-but-unreported for the same reason.

Four of the five look like the tool-collapse family. In particular every_prompt_names_at_least_one_tool_it_can_call reports skill_creator as an agent that carries tools while its prompt names none of them — which, if accurate, is a live prompt/belt mismatch rather than a stale test.

Solution (optional)

Triage the five before #6462 merges, so the red is understood in advance rather than investigated under pressure. Establish for each whether the test is stale or the behaviour regressed — the fleet_prompt_tests one in particular asserts a real invariant and a failure there is more likely to be the product than the test.

Acceptance criteria

Related

Found at 0f1ecc9d2 while proving a PR's failures were not its own. #6462 (unblocks the steps), #6451 (the curtain that hid them), #6423 (same reachable-but-unreported dynamic).

Activity

  1. added
    testTest additions, fixes, or harness work.
    rust-coreCore Rust runtime in src/: CLI, core_server, shared infrastructure.
    on Sep 23, 2026
  2. added
    priority: p2Soon. Real but survivable — a rough edge, a gap, a thing that will bite later.
    on Sep 23, 2026
  3. tinysweeper commented on Sep 23, 2026

    @tinysweeper

    Five cargo test -p openhuman --lib tests fail on main but are hidden due to a CI job skipping later steps. These tests will become visible once #6462 unblocks the CI steps. They are pre-existing failures that need triage before the CI fix lands.

    Labelled priority: p2.

  4. added theissue type on Sep 23, 2026
  5. M3gA-Mind commented on Sep 23, 2026

    @M3gA-Mind
    CollaboratorAuthor

    Investigated at 0f1ecc9d28d85289ee3f87ade13156be02ffbd28, on a branch cut fresh from upstream/main.

    The five are real and reproducible. But the mechanism in the title is wrong: they are not hidden behind a step curtain — they are never executed. Correcting that first, because it sends the next person to the wrong fix.

    The five, confirmed

    RUST_MIN_STACK=67108864 cargo test --manifest-path Cargo.toml -p openhuman --lib → 10524 passed; 5 failed; 83 ignored.

    # test first error
    1 agent::harness::tests::harness_tool_call_parsing_tests::parse_tool_calls_recovers_mismatched_close_tag assertion failed: text.is_empty() (:289)
    2 agent::orchestration::tools::spawn_async_subagent::tests::errors_clearly_when_no_parent_thread_for_delivery panics carrying the expected error text (:404)
    3 agent::prompts::tests::subagent_render_tests::render_subagent_system_prompt_honors_identity_safety_and_skills_flags assertion failed: rendered.contains("## Tools") (:169)
    4 agent::registry::agents::fleet_prompt_tests::every_prompt_names_at_least_one_tool_it_can_call skill_creator has joined the silent set (:365)
    5 agent::registry::agents::orchestrator::prompt::tests::session_routing_tests::the_withheld_block_renders_for_a_renamed_session_with_a_filter a packed delegate must render with its route (:61)

    All five are real assertion failures, not environment. Not RUST_MIN_STACK — the run used CI's value and each reports panicked at … assertion with a file:line, whereas a stack overflow aborts the process with fatal runtime error: stack overflow and no per-test result. Not the raw-coverage baseline — that belongs to the aggregated raw_coverage_all target, not --lib. They reproduce across repeated runs and in isolation, so ordering and cross-test global pollution are ruled out as well.

    They are not concealed. No lane on main runs this suite.

    Three facts from the workflow files:

    1. test.yml — the only automatic caller of the full --lib suite — is workflow_dispatch only. Its own header comment reads "PR/push test gate. Delegates to the reusable test-reusable.yml". The entire file is 22 lines and its trigger block is exactly on: workflow_dispatch: {} — no push, no pull_request.
    2. The other caller, ci-full.yml, is release-only (push/pull_request on branches [release], plus dispatch). It never runs for main.
    3. ci-lite.yml does run on main, but its --lib invocations are narrowly filtered (:540-568): core::all:: core::cli:: core::jsonrpc:: core::legacy_aliases:: core::runtime:: agent::registry::agents::loader:: memory::people::contacts_gate_tests:: openhuman::config:: openhuman::platform::socket::event_handlers:: tools::schemas:: tools::ops::tests::, plus mcp::server::resources:: and test_support::introspect::.

    None of the five falls inside those filters, so the recent green CI Lite runs on main and these five failures are both true at once.

    The near-miss is how this survived. One filter is agent::registry::agents::loader::; failure #4 is agent::registry::agents::fleet_prompt_tests:: — a sibling module, not under loader::. Anyone scanning that filter list sees agent::registry::agents:: and reasonably concludes the area is covered.

    So "they will become visible once #6462 unblocks the CI steps" is false. Unblocking a step cannot add a suite that no lane invokes. This is also not the same mechanism as #6451: that is a genuine curtain over work which does run, whereas this is work that does not run. Same symptom — "tests fail on main, CI is green" — two different causes, and a fix for either would miss the other.

    Two gaps here, and only one is documented

    Documented: the scope filters are deliberate. ci-lite.yml:538 says "These filters route AROUND that test by scope, not by skip. Running the full suite is tracked in #5021." #5021 is open and explains the blocker (a task_local stack overflow in agent::harness::session::tests::turn_dispatches_spawn_subagent_through_full_path that also OOMs unscoped).

    Not documented: test.yml having no push/pull_request trigger. #5021's body does not mention test.yml, workflow_dispatch, trigger or pull_request at all — and its subject is the gates-off (--no-default-features --features tokenjuice-treesitter) lane specifically, not the general absence of an automatic full-suite lane. So a workflow that documents itself as the PR/push gate, and is not wired to either, is a separate and currently unexplained gap. Whether that is deliberate is a maintainer's call, so flagging rather than filing.

    Scope caveat on the reproduction

    My run used default features. test-reusable.yml:225-234 runs the suite under the product feature set (scripts/ci/product-features.sh). I have not run the five under that profile, so I cannot say whether all five would also fail there — only that no lane on main runs either profile's full suite today.

    Triage of the five

    #4 is a genuine product defect and must not be fixed by relaxing the assertion. Filed separately — see the linked issue. Briefly: fleet_prompt_tests.rs:340-368 asserts that every agent carrying a belt names at least one tool it can actually call, and its docstring says it exists to catch "a belt narrowed out from under its prompt, which otherwise reads as an improvement: the tool bytes fall and nothing else moves." skill_creator now fails that. Adding it to NAMES_NO_TOOL would silence the one assertion in the suite doing its job.

    A correlation worth recording as checked-and-discarded, because it is convincing enough that the next reader will form it: 66eef8016 (#6447, "drop the integrations sub-agent") touches prompts/sections.rs, prompts/types.rs, the orchestrator prompt, tools_agent/ and deletes skill_delegation.rs — the subject of four of the five. It is not the cause: checking out its first parent 622aea071 and rebuilding, all five still fail. They also are not one cause — the five test files' last-touch commits differ.

    #1, #2, #3 and #5 are not yet classified. Each is a real assertion; whether the test or the product drifted needs reading each subject, which I am doing next and will report rather than assume.

    Suggested re-pointing

    This issue's value is the five-test triage, which stands. Its stated mechanism does not. Suggest re-pointing it at the coverage gap (and at #5021 for the documented half) rather than at #6462/#6451, and treating test.yml's trigger as its own question.

  6. M3gA-Mind commented on Sep 23, 2026

    @M3gA-Mind
    CollaboratorAuthor

    Correction to my comment above. I called test.yml's missing push/pull_request trigger "a separate and currently unexplained gap". It is explained — it was a deliberate consolidation, not a regression, and I should have checked the history before saying otherwise.

    The chain, verified at 0f1ecc9d2:

    1. f64199e4e — "ci: consolidate pull request checks (ci: consolidate pull request checks #3001)", 2026-05-30, by a maintainer, an ancestor of main. It removed push: branches: [main] and pull_request: from test.yml and added pr-ci.yml, trimming triggers across nine other workflows in the same commit. test.yml was deliberately left as a dispatch-only entry point.
    2. 5650c6d33 — "ci: two-lane CI (ci-lite/ci-full) … (ci: two-lane CI (ci-lite/ci-full), manual main→release promotion, universal back-merge #4533)", 2026-07-04. pr-ci.yml does not exist on main; this is the commit that replaced it, creating ci-lite.yml.

    So the lineage is test.yml → pr-ci.yml → ci-lite.yml, across two intentional consolidations — and ci-lite.yml carries only the scoped --lib filters. The full suite was not silently dropped by a broken trigger; it was left out of the lane that replaced the one that replaced it, and that consequence is what #5021 already tracks.

    What is actually stale is the comment at the top of test.yml, which still calls it "PR/push test gate. Delegates to the reusable test-reusable.yml" — inaccurate since #3001, along with its permissions: pull-requests: read and a concurrency group keyed on github.event.pull_request.number. Three leftovers from one removal, not three signals. Worth a tidy-up, not an issue.

    Everything else in my comment stands — the five are real, reproducible and unrun; none falls inside ci-lite.yml's filters; the agent::registry::agents::loader:: vs fleet_prompt_tests:: sibling near-miss is how it survived; and "they will become visible once #6462 unblocks the CI steps" remains false, because unblocking a step cannot add a suite no lane invokes. The right pointer for the coverage half is #5021.

  7. M3gA-Mind commented on Sep 23, 2026

    @M3gA-Mind
    CollaboratorAuthor

    Corrected count: three tests fail under CI's feature profile, not five

    "Five" is a default-features figure — mine. It should not be quoted as the number that matters. I flagged the profile caveat in my first comment and then failed to act on it; this corrects that.

    CI's full-suite lane runs cargo test -p openhuman --lib --features "$(scripts/ci/product-features.sh)" (test-reusable.yml:225-234), i.e. channels,media,inference,voice,web3,documents,modules,flows,skills,mcp,crash-reporting,http-server,scheduler-gate,file-logging,contacts,runtime-node,hosting. Under that profile:

    the_withheld_block_renders_for_a_renamed_session_with_a_filter ... ok
    every_prompt_names_at_least_one_tool_it_can_call ... ok
    render_subagent_system_prompt_honors_identity_safety_and_skills_flags ... FAILED
    parse_tool_calls_recovers_mismatched_close_tag ... FAILED
    errors_clearly_when_no_parent_thread_for_delivery ... FAILED
    
    test result: FAILED. 2 passed; 3 failed
    
    # test under default under product features classification
    1 parse_tool_calls_recovers_mismatched_close_tag fail fail unclassified; parser/vendored-adjacent
    2 errors_clearly_when_no_parent_thread_for_delivery fail fail test drift (superseded string)
    3 render_subagent_system_prompt_honors_identity_safety_and_skills_flags fail fail test drift (vendored dialect wording)
    4 every_prompt_names_at_least_one_tool_it_can_call fail pass profile artefact — not a failure
    5 the_withheld_block_renders_for_a_renamed_session_with_a_filter fail pass profile artefact — not a failure

    So: three real failures, two of which are test drift, and one unclassified.

    Why #4 and #5 are artefacts

    NodeExecTool is registered under #[cfg(feature = "runtime-node")] (crates/openhuman-core/src/tools/ops.rs:863-865), and runtime-node is absent from default:

    default = ["media", "skills", "flows", "mcp", "channels", "medulla", "http-server", "scheduler-gate", "file-logging", "modules"]
    

    but present in scripts/ci/product-features.txt. With default features node_exec/npm_exec do not exist, so they are not in tool_universe(), can_call is false for both, and #4's guard reports skill_creator as naming nothing callable. #5 fails the same way for a documents-gated tool (make_presentation); documents is likewise product-only.

    I filed #6507 as a product defect on the strength of #4 before checking the profile. It is retracted and closed as invalid, including the withdrawn attribution to #6436 — that PR did not cause it, because there was nothing to cause.

    Corrected classifications for the three that remain

    • Feat/landing revamp #3 — test drift, settled. The subagent path does not use ToolsSection; render_helpers/subagent.rs:195-196 calls render_tool_dialect_prompt into the vendored tinyagents dialect, which renders ## Tool Use Protocol + ### Available Tools + Arguments: where the test expects ## Tools + Parameters: + "type". The catalogue is present, and a probe with a rich schema confirmed field names, types, optionality (? suffix) and descriptions all survive — the compact form is denser, not lossier. No defect; the assertions pin superseded wording.
    • #2 — test drift. :404 asserts the guidance contains the literal spawn_subagent; :405 (blocking: true) and :406 (delegate_) pass. The guidance now recommends delegate_*. Fix belongs in the assertion — pending a check that what it now names is callable under the product profile.
    • Feat/gitbooks #1 — unclassified. assertion failed: text.is_empty() after recovering a mismatched close tag. Parser-side, so it may belong upstream in the vendored tinyagents rather than here.

    Everything else in my earlier comments stands

    The five (under default features) are real assertions and not environment — not RUST_MIN_STACK, not the raw-coverage baseline, reproducible and order-independent. No lane on main runs this suite in either profile; the chain is test.yml → pr-ci.yml (#3001) → ci-lite.yml (#4533), two intentional consolidations, with the coverage gap tracked by #5021. And "they will become visible once #6462 unblocks the CI steps" remains false.

    A separate, real finding falls out of this, which I am filing on its own: #4 and #5 are implicitly profile-dependent with no #[cfg] and no comment saying so, so they fail misleadingly under default in a way that reads as a product defect. That is what produced the retracted issue.

  8. M3gA-Mind commented on Sep 23, 2026

    @M3gA-Mind
    CollaboratorAuthor

    Triage complete. Final tally: one upstream defect, two drift, two profile artefacts, zero openhuman product defects.

    That is the opposite of this issue's framing, and the opposite of what I concluded when I started. Both reversals came from probes, not argument.

    # test verdict state
    1 parse_tool_calls_recovers_mismatched_close_tag upstream defect; our test is correct specified below, not started
    2 errors_clearly_when_no_parent_thread_for_delivery drift fixed — PR #6517
    3 render_subagent_system_prompt_honors_identity_safety_and_skills_flags drift specified below, not started
    4 every_prompt_names_at_least_one_tool_it_can_call profile artefact — passes under CI's features no action; see #6512
    5 the_withheld_block_renders_for_a_renamed_session_with_a_filter profile artefact — passes under CI's features no action; see #6512

    Only #2 is changed. The two below are specified rather than half-done.


    Specification — #1: parse_tool_calls_recovers_mismatched_close_tag

    Do not change the assertion. It is correct and it is currently the only thing catching this.

    What happens

    Fixture:

    let response = r#"<tool_call>
    {"name": "shell", "arguments": {"command": "uptime"}}
    </arg_value>"#;

    Probed actual behaviour:

    text        = "</arg_value>"      (len 12)
    calls.len() = 1
    call: name="shell" args={"command":"uptime"}
    

    The recovery works — the malformed call is salvaged with correct name and arguments. The defect is that the unmatched close tag survives into the text half. The test asserts text.is_empty(), which is the right contract.

    Why it matters — the consumer

    OpenHuman calls tinytools_agent::parse_tool_calls in exactly one production place: crates/openhuman-core/src/agent/subagent_host/ops/checkpoint.rs:59, the sub-agent cap-hit checkpoint summary.

    let (prose, _) = tinytools_agent::parse_tool_calls(&raw);
    let text = if prose.trim().is_empty() { deterministic } else { prose };

    "</arg_value>".trim() is not empty, so the residue passes that guard and is handed to the delegating parent as the sub-agent's progress summary — displacing the deterministic fallback that exists for exactly this case. So this is not cosmetic: a malformed fragment substitutes for a useful summary.

    Honest counterweight on likelihood: the checkpoint prompt explicitly says "Do not call tools", so reaching this needs a model to emit a malformed tool call it was told not to emit at all. That is low-probability — and it is precisely the population a recovery path exists to serve.

    Ownership and why this is two pieces of work

    parse_tool_calls is tinytools_agent, i.e. vendor/tinyagents/vendor/tinytools — the nested submodule. Nothing in openhuman parses this; we only consume the result. So the fix is:

    1. a tinytools change — strip the unmatched close tag from the text half during recovery;
    2. a gitlink bump here.

    The openhuman half cannot land until the upstream half does. Unstarted deliberately: two repos at once is worse than a clear specification.

    Fails under both feature profiles, so this one is not a profile artefact.


    Specification — #3: render_subagent_system_prompt_honors_identity_safety_and_skills_flags

    The cheap fix is the wrong fix. Do not repoint the assertions at the dialect's current headings.

    What happens

    The test asserts rendered.contains("## Tools"), "Parameters:" and "\"type\"". The subagent path does not use ToolsSection — render_helpers/subagent.rs:195-196 calls render_tool_dialect_prompt into the vendored tinyagents dialect, which renders:

    ## Tool Use Protocol
    …
    ### Available Tools
    
    Arguments are shown as `{name: type, optional?: type}`.
    
    **rich_probe_tool**: probe desc
    Arguments: `{deep?: boolean, limit?: integer, path: string}`
      - limit: max rows
      - path: the file path
    

    So the catalogue is present and the test's stated intent is satisfied; only its string expectations are stale.

    Why not just update the strings

    Changing ## Tools → ### Available Tools and Parameters: → Arguments: pins the vendored dialect's current wording — the very thing that just drifted. The next tinyagents bump breaks it again and the next person makes the same edit.

    What to assert instead: the information, not the presentation

    The test's own comment states the contract — a Json-dispatch subagent must get the parameter schema inline so the model knows what to emit. So assert that the schema's content reaches the prompt: the tool name, its field names, and its types.

    A probe with a rich schema (path: string required + description, limit: integer + description, deep: boolean) established exactly what is stable across the compact form:

    present absent
    tool name, path, limit, deep the heading ## Tools
    string, integer, boolean the label Parameters:
    optionality, via the ? suffix the literal word required
    descriptions, for fields that have one the raw JSON schema

    An assertion over field names and types cannot be broken by a heading rename, and still fails if the catalogue empties — which is the failure this test exists to catch. Note the fixture matters: the current TestTool schema is {"type": "object"} with no properties (mod_tests.rs:40-42), so it cannot distinguish "fields dropped" from "no fields"; a schema with real properties is needed for the assertion to mean anything.

    Also worth recording from that probe: the compact form is denser, not lossier — no capability loss, so there is no product bug hiding behind #3.


    Everything established earlier still stands

    Five are real assertions under default features and not environment — not RUST_MIN_STACK, not the raw-coverage baseline, reproducible and order-independent. No lane on main runs this suite in either profile (test.yml → pr-ci.yml #3001 → ci-lite.yml #4533, two intentional consolidations; coverage gap tracked by #5021). And "they will become visible once #6462 unblocks the CI steps" remains false — unblocking a step cannot add a suite no lane invokes.

  9. senamakel commented on Oct 9, 2026

    @senamakel
    Member

    Triage: needs opinion — “Five openhuman --lib tests fail on main, hidden behind the Rust Quality step-7 curtain” needs a maintainer decision on product scope/priority or fresh reproduction evidence before its status can be settled. Please advise whether to pursue, narrow, or close it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    priority: p2Soon. Real but survivable — a rough edge, a gap, a thing that will bite later.rust-coreCore Rust runtime in src/: CLI, core_server, shared infrastructure.source: teamIssue opened by a repository membertestTest additions, fixes, or harness work.triage: needs-opinionValid issue awaiting a maintainer product or priority decision

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions