Repository navigation
Five openhuman --lib tests fail on main, hidden behind the Rust Quality step-7 curtain #6486
Description
Activity
- addedtestTest additions, fixes, or harness work.Test additions, fixes, or harness work.rust-coreCore Rust runtime in src/: CLI, core_server, shared infrastructure.Core Rust runtime in src/: CLI, core_server, shared infrastructure.
on Sep 23, 2026 - addedpriority: p2Soon. Real but survivable — a rough edge, a gap, a thing that will bite later.Soon. Real but survivable — a rough edge, a gap, a thing that will bite later.
on Sep 23, 2026 Five cargo test -p openhuman --lib tests fail on main but are hidden due to a CI job skipping later steps. These tests will become visible once #6462 unblocks the CI steps. They are pre-existing failures that need triage before the CI fix lands.
Labelled
priority: p2.Investigated at
0f1ecc9d28d85289ee3f87ade13156be02ffbd28, on a branch cut fresh fromupstream/main.The five are real and reproducible. But the mechanism in the title is wrong: they are not hidden behind a step curtain — they are never executed. Correcting that first, because it sends the next person to the wrong fix.
The five, confirmed
RUST_MIN_STACK=67108864 cargo test --manifest-path Cargo.toml -p openhuman --lib→10524 passed; 5 failed; 83 ignored.# test first error 1 agent::harness::tests::harness_tool_call_parsing_tests::parse_tool_calls_recovers_mismatched_close_tagassertion failed: text.is_empty()(:289)2 agent::orchestration::tools::spawn_async_subagent::tests::errors_clearly_when_no_parent_thread_for_deliverypanics carrying the expected error text ( :404)3 agent::prompts::tests::subagent_render_tests::render_subagent_system_prompt_honors_identity_safety_and_skills_flagsassertion failed: rendered.contains("## Tools")(:169)4 agent::registry::agents::fleet_prompt_tests::every_prompt_names_at_least_one_tool_it_can_callskill_creatorhas joined the silent set (:365)5 agent::registry::agents::orchestrator::prompt::tests::session_routing_tests::the_withheld_block_renders_for_a_renamed_session_with_a_filtera packed delegate must render with its route(:61)All five are real assertion failures, not environment. Not
RUST_MIN_STACK— the run used CI's value and each reportspanicked at … assertionwith a file:line, whereas a stack overflow aborts the process withfatal runtime error: stack overflowand no per-test result. Not the raw-coverage baseline — that belongs to the aggregatedraw_coverage_alltarget, not--lib. They reproduce across repeated runs and in isolation, so ordering and cross-test global pollution are ruled out as well.They are not concealed. No lane on
mainruns this suite.Three facts from the workflow files:
test.yml— the only automatic caller of the full--libsuite — isworkflow_dispatchonly. Its own header comment reads "PR/push test gate. Delegates to the reusabletest-reusable.yml". The entire file is 22 lines and its trigger block is exactlyon: workflow_dispatch: {}— nopush, nopull_request.- The other caller,
ci-full.yml, isrelease-only (push/pull_requeston branches[release], plus dispatch). It never runs formain. ci-lite.ymldoes run onmain, but its--libinvocations are narrowly filtered (:540-568):core::all:: core::cli:: core::jsonrpc:: core::legacy_aliases:: core::runtime:: agent::registry::agents::loader:: memory::people::contacts_gate_tests:: openhuman::config:: openhuman::platform::socket::event_handlers:: tools::schemas:: tools::ops::tests::, plusmcp::server::resources::andtest_support::introspect::.
None of the five falls inside those filters, so the recent green
CI Literuns onmainand these five failures are both true at once.The near-miss is how this survived. One filter is
agent::registry::agents::loader::; failure #4 isagent::registry::agents::fleet_prompt_tests::— a sibling module, not underloader::. Anyone scanning that filter list seesagent::registry::agents::and reasonably concludes the area is covered.So "they will become visible once #6462 unblocks the CI steps" is false. Unblocking a step cannot add a suite that no lane invokes. This is also not the same mechanism as #6451: that is a genuine curtain over work which does run, whereas this is work that does not run. Same symptom — "tests fail on main, CI is green" — two different causes, and a fix for either would miss the other.
Two gaps here, and only one is documented
Documented: the scope filters are deliberate.
ci-lite.yml:538says "These filters route AROUND that test by scope, not by skip. Running the full suite is tracked in #5021." #5021 is open and explains the blocker (atask_localstack overflow inagent::harness::session::tests::turn_dispatches_spawn_subagent_through_full_paththat also OOMs unscoped).Not documented:
test.ymlhaving nopush/pull_requesttrigger. #5021's body does not mentiontest.yml,workflow_dispatch,triggerorpull_requestat all — and its subject is the gates-off (--no-default-features --features tokenjuice-treesitter) lane specifically, not the general absence of an automatic full-suite lane. So a workflow that documents itself as the PR/push gate, and is not wired to either, is a separate and currently unexplained gap. Whether that is deliberate is a maintainer's call, so flagging rather than filing.Scope caveat on the reproduction
My run used default features.
test-reusable.yml:225-234runs the suite under the product feature set (scripts/ci/product-features.sh). I have not run the five under that profile, so I cannot say whether all five would also fail there — only that no lane onmainruns either profile's full suite today.Triage of the five
#4 is a genuine product defect and must not be fixed by relaxing the assertion. Filed separately — see the linked issue. Briefly:
fleet_prompt_tests.rs:340-368asserts that every agent carrying a belt names at least one tool it can actually call, and its docstring says it exists to catch "a belt narrowed out from under its prompt, which otherwise reads as an improvement: the tool bytes fall and nothing else moves."skill_creatornow fails that. Adding it toNAMES_NO_TOOLwould silence the one assertion in the suite doing its job.A correlation worth recording as checked-and-discarded, because it is convincing enough that the next reader will form it:
66eef8016(#6447, "drop the integrations sub-agent") touchesprompts/sections.rs,prompts/types.rs, the orchestrator prompt,tools_agent/and deletesskill_delegation.rs— the subject of four of the five. It is not the cause: checking out its first parent622aea071and rebuilding, all five still fail. They also are not one cause — the five test files' last-touch commits differ.#1, #2, #3 and #5 are not yet classified. Each is a real assertion; whether the test or the product drifted needs reading each subject, which I am doing next and will report rather than assume.
Suggested re-pointing
This issue's value is the five-test triage, which stands. Its stated mechanism does not. Suggest re-pointing it at the coverage gap (and at #5021 for the documented half) rather than at #6462/#6451, and treating
test.yml's trigger as its own question.Correction to my comment above. I called
test.yml's missingpush/pull_requesttrigger "a separate and currently unexplained gap". It is explained — it was a deliberate consolidation, not a regression, and I should have checked the history before saying otherwise.The chain, verified at
0f1ecc9d2:f64199e4e— "ci: consolidate pull request checks (ci: consolidate pull request checks #3001)", 2026-05-30, by a maintainer, an ancestor ofmain. It removedpush: branches: [main]andpull_request:fromtest.ymland addedpr-ci.yml, trimming triggers across nine other workflows in the same commit.test.ymlwas deliberately left as a dispatch-only entry point.5650c6d33— "ci: two-lane CI (ci-lite/ci-full) … (ci: two-lane CI (ci-lite/ci-full), manual main→release promotion, universal back-merge #4533)", 2026-07-04.pr-ci.ymldoes not exist onmain; this is the commit that replaced it, creatingci-lite.yml.
So the lineage is
test.yml→pr-ci.yml→ci-lite.yml, across two intentional consolidations — andci-lite.ymlcarries only the scoped--libfilters. The full suite was not silently dropped by a broken trigger; it was left out of the lane that replaced the one that replaced it, and that consequence is what #5021 already tracks.What is actually stale is the comment at the top of
test.yml, which still calls it "PR/push test gate. Delegates to the reusabletest-reusable.yml" — inaccurate since #3001, along with itspermissions: pull-requests: readand aconcurrencygroup keyed ongithub.event.pull_request.number. Three leftovers from one removal, not three signals. Worth a tidy-up, not an issue.Everything else in my comment stands — the five are real, reproducible and unrun; none falls inside
ci-lite.yml's filters; theagent::registry::agents::loader::vsfleet_prompt_tests::sibling near-miss is how it survived; and "they will become visible once #6462 unblocks the CI steps" remains false, because unblocking a step cannot add a suite no lane invokes. The right pointer for the coverage half is #5021.Corrected count: three tests fail under CI's feature profile, not five
"Five" is a default-features figure — mine. It should not be quoted as the number that matters. I flagged the profile caveat in my first comment and then failed to act on it; this corrects that.
CI's full-suite lane runs
cargo test -p openhuman --lib --features "$(scripts/ci/product-features.sh)"(test-reusable.yml:225-234), i.e.channels,media,inference,voice,web3,documents,modules,flows,skills,mcp,crash-reporting,http-server,scheduler-gate,file-logging,contacts,runtime-node,hosting. Under that profile:the_withheld_block_renders_for_a_renamed_session_with_a_filter ... ok every_prompt_names_at_least_one_tool_it_can_call ... ok render_subagent_system_prompt_honors_identity_safety_and_skills_flags ... FAILED parse_tool_calls_recovers_mismatched_close_tag ... FAILED errors_clearly_when_no_parent_thread_for_delivery ... FAILED test result: FAILED. 2 passed; 3 failed# test under defaultunder product features classification 1 parse_tool_calls_recovers_mismatched_close_tagfail fail unclassified; parser/vendored-adjacent 2 errors_clearly_when_no_parent_thread_for_deliveryfail fail test drift (superseded string) 3 render_subagent_system_prompt_honors_identity_safety_and_skills_flagsfail fail test drift (vendored dialect wording) 4 every_prompt_names_at_least_one_tool_it_can_callfail pass profile artefact — not a failure 5 the_withheld_block_renders_for_a_renamed_session_with_a_filterfail pass profile artefact — not a failure So: three real failures, two of which are test drift, and one unclassified.
Why #4 and #5 are artefacts
NodeExecToolis registered under#[cfg(feature = "runtime-node")](crates/openhuman-core/src/tools/ops.rs:863-865), andruntime-nodeis absent fromdefault:default = ["media", "skills", "flows", "mcp", "channels", "medulla", "http-server", "scheduler-gate", "file-logging", "modules"]but present in
scripts/ci/product-features.txt. With default featuresnode_exec/npm_execdo not exist, so they are not intool_universe(),can_callis false for both, and #4's guard reportsskill_creatoras naming nothing callable. #5 fails the same way for adocuments-gated tool (make_presentation);documentsis likewise product-only.I filed #6507 as a product defect on the strength of #4 before checking the profile. It is retracted and closed as invalid, including the withdrawn attribution to #6436 — that PR did not cause it, because there was nothing to cause.
Corrected classifications for the three that remain
- Feat/landing revamp #3 — test drift, settled. The subagent path does not use
ToolsSection;render_helpers/subagent.rs:195-196callsrender_tool_dialect_promptinto the vendored tinyagents dialect, which renders## Tool Use Protocol+### Available Tools+Arguments:where the test expects## Tools+Parameters:+"type". The catalogue is present, and a probe with a rich schema confirmed field names, types, optionality (?suffix) and descriptions all survive — the compact form is denser, not lossier. No defect; the assertions pin superseded wording. - #2 — test drift.
:404asserts the guidance contains the literalspawn_subagent;:405(blocking: true) and:406(delegate_) pass. The guidance now recommendsdelegate_*. Fix belongs in the assertion — pending a check that what it now names is callable under the product profile. - Feat/gitbooks #1 — unclassified.
assertion failed: text.is_empty()after recovering a mismatched close tag. Parser-side, so it may belong upstream in the vendored tinyagents rather than here.
Everything else in my earlier comments stands
The five (under default features) are real assertions and not environment — not
RUST_MIN_STACK, not the raw-coverage baseline, reproducible and order-independent. No lane onmainruns this suite in either profile; the chain istest.yml→pr-ci.yml(#3001) →ci-lite.yml(#4533), two intentional consolidations, with the coverage gap tracked by #5021. And "they will become visible once #6462 unblocks the CI steps" remains false.A separate, real finding falls out of this, which I am filing on its own: #4 and #5 are implicitly profile-dependent with no
#[cfg]and no comment saying so, so they fail misleadingly underdefaultin a way that reads as a product defect. That is what produced the retracted issue.- Feat/landing revamp #3 — test drift, settled. The subagent path does not use
Triage complete. Final tally: one upstream defect, two drift, two profile artefacts, zero openhuman product defects.
That is the opposite of this issue's framing, and the opposite of what I concluded when I started. Both reversals came from probes, not argument.
# test verdict state 1 parse_tool_calls_recovers_mismatched_close_tagupstream defect; our test is correct specified below, not started 2 errors_clearly_when_no_parent_thread_for_deliverydrift fixed — PR #6517 3 render_subagent_system_prompt_honors_identity_safety_and_skills_flagsdrift specified below, not started 4 every_prompt_names_at_least_one_tool_it_can_callprofile artefact — passes under CI's features no action; see #6512 5 the_withheld_block_renders_for_a_renamed_session_with_a_filterprofile artefact — passes under CI's features no action; see #6512 Only #2 is changed. The two below are specified rather than half-done.
Specification — #1:
parse_tool_calls_recovers_mismatched_close_tagDo not change the assertion. It is correct and it is currently the only thing catching this.
What happens
Fixture:
let response = r#"<tool_call> {"name": "shell", "arguments": {"command": "uptime"}} </arg_value>"#;
Probed actual behaviour:
text = "</arg_value>" (len 12) calls.len() = 1 call: name="shell" args={"command":"uptime"}The recovery works — the malformed call is salvaged with correct name and arguments. The defect is that the unmatched close tag survives into the text half. The test asserts
text.is_empty(), which is the right contract.Why it matters — the consumer
OpenHuman calls
tinytools_agent::parse_tool_callsin exactly one production place:crates/openhuman-core/src/agent/subagent_host/ops/checkpoint.rs:59, the sub-agent cap-hit checkpoint summary.let (prose, _) = tinytools_agent::parse_tool_calls(&raw); let text = if prose.trim().is_empty() { deterministic } else { prose };
"</arg_value>".trim()is not empty, so the residue passes that guard and is handed to the delegating parent as the sub-agent's progress summary — displacing thedeterministicfallback that exists for exactly this case. So this is not cosmetic: a malformed fragment substitutes for a useful summary.Honest counterweight on likelihood: the checkpoint prompt explicitly says "Do not call tools", so reaching this needs a model to emit a malformed tool call it was told not to emit at all. That is low-probability — and it is precisely the population a recovery path exists to serve.
Ownership and why this is two pieces of work
parse_tool_callsistinytools_agent, i.e.vendor/tinyagents/vendor/tinytools— the nested submodule. Nothing in openhuman parses this; we only consume the result. So the fix is:- a
tinytoolschange — strip the unmatched close tag from the text half during recovery; - a gitlink bump here.
The openhuman half cannot land until the upstream half does. Unstarted deliberately: two repos at once is worse than a clear specification.
Fails under both feature profiles, so this one is not a profile artefact.
Specification — #3:
render_subagent_system_prompt_honors_identity_safety_and_skills_flagsThe cheap fix is the wrong fix. Do not repoint the assertions at the dialect's current headings.
What happens
The test asserts
rendered.contains("## Tools"),"Parameters:"and"\"type\"". The subagent path does not useToolsSection—render_helpers/subagent.rs:195-196callsrender_tool_dialect_promptinto the vendored tinyagents dialect, which renders:## Tool Use Protocol … ### Available Tools Arguments are shown as `{name: type, optional?: type}`. **rich_probe_tool**: probe desc Arguments: `{deep?: boolean, limit?: integer, path: string}` - limit: max rows - path: the file pathSo the catalogue is present and the test's stated intent is satisfied; only its string expectations are stale.
Why not just update the strings
Changing
## Tools→### Available ToolsandParameters:→Arguments:pins the vendored dialect's current wording — the very thing that just drifted. The next tinyagents bump breaks it again and the next person makes the same edit.What to assert instead: the information, not the presentation
The test's own comment states the contract — a Json-dispatch subagent must get the parameter schema inline so the model knows what to emit. So assert that the schema's content reaches the prompt: the tool name, its field names, and its types.
A probe with a rich schema (
path: stringrequired + description,limit: integer+ description,deep: boolean) established exactly what is stable across the compact form:present absent tool name, path,limit,deepthe heading ## Toolsstring,integer,booleanthe label Parameters:optionality, via the ?suffixthe literal word requireddescriptions, for fields that have one the raw JSON schema An assertion over field names and types cannot be broken by a heading rename, and still fails if the catalogue empties — which is the failure this test exists to catch. Note the fixture matters: the current
TestToolschema is{"type": "object"}with no properties (mod_tests.rs:40-42), so it cannot distinguish "fields dropped" from "no fields"; a schema with real properties is needed for the assertion to mean anything.Also worth recording from that probe: the compact form is denser, not lossier — no capability loss, so there is no product bug hiding behind #3.
Everything established earlier still stands
Five are real assertions under
defaultfeatures and not environment — notRUST_MIN_STACK, not the raw-coverage baseline, reproducible and order-independent. No lane onmainruns this suite in either profile (test.yml→pr-ci.yml#3001 →ci-lite.yml#4533, two intentional consolidations; coverage gap tracked by #5021). And "they will become visible once #6462 unblocks the CI steps" remains false — unblocking a step cannot add a suite no lane invokes.- a
- added a commit that references this issue
on Sep 23, 2026 - added a commit that references this issue
on Sep 23, 2026 - addedtriage: needs-opinionValid issue awaiting a maintainer product or priority decisionValid issue awaiting a maintainer product or priority decision
on Oct 9, 2026 Triage: needs opinion — “Five openhuman --lib tests fail on main, hidden behind the Rust Quality step-7 curtain” needs a maintainer decision on product scope/priority or fresh reproduction evidence before its status can be settled. Please advise whether to pursue, narrow, or close it.
- addedsource: teamIssue opened by a repository memberIssue opened by a repository member
on Oct 9, 2026
Summary
Five
cargo test -p openhuman --libtests fail onmain. They are invisible today becauseRust Qualitydies at an earlier step and the test steps reportskipped— so landing #6462 (which unblocks steps 8–15) will surface them as a fresh red, and that red will not be #6462's fault.Problem
Proven pre-existing at
0f1ecc9d2, not inferred: a warm worktree was detached to that commit, the absence of any unrelated change was confirmed, and the five were run by exact name. Result: 0 passed; 5 failed.All five are prompt/parsing tests. None touches usage, the codec, or any area currently under change.
Why they are invisible.
Rust Qualitycurrently fails at step 7 ("Enforce agent runtime ownership boundary"), and every later step in that job reportsskipped— the lib test steps among them. Both coverage lanes are additionally gated onrust-qualitysucceeding (ci-lite.yml:750,:898), so they do not run either. The suite has therefore not been executed onmainfor days.Consequence worth stating plainly: #6462 fixes step 7. The moment it lands, steps 8–15 become reachable for the first time in days and these five will appear. That is the gate working, not a regression introduced by #6462 — exactly the same dynamic already recorded on #6423, where the prompt-prefix guard's six budget errors are reachable-but-unreported for the same reason.
Four of the five look like the tool-collapse family. In particular
every_prompt_names_at_least_one_tool_it_can_callreportsskill_creatoras an agent that carries tools while its prompt names none of them — which, if accurate, is a live prompt/belt mismatch rather than a stale test.Solution (optional)
Triage the five before #6462 merges, so the red is understood in advance rather than investigated under pressure. Establish for each whether the test is stale or the behaviour regressed — the
fleet_prompt_testsone in particular asserts a real invariant and a failure there is more likely to be the product than the test.Acceptance criteria
mainwith the test steps actually executing.Related
Found at
0f1ecc9d2while proving a PR's failures were not its own. #6462 (unblocks the steps), #6451 (the curtain that hid them), #6423 (same reachable-but-unreported dynamic).