Repository navigation
Conversation
Update the vendored tinyagents submodule to the latest commit. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
Co-authored-by: Medulla <medulla@tinyhumans.ai>
# Conflicts: # crates/openhuman-core/src/security/keyring/encrypted_file_backend.rs
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
|
Closing this Jev-assisted tool-recovery implementation at the maintainer's request, together with #7341 and its implementation PRs: #7345, tinyhumansai/tinyagents#375, and tinyhumansai/tinytools#61. The investigation did not establish an incremental recovery or task-completion benefit from this adviser for the observed workload. The maintainer's recent local sessions used native structured tool calls. Observed failures primarily involved command execution, missing runtime state, unavailable services, and a declared JSON Schema anyOf requirement that local validation did not enforce. Recognized schema failures also take the existing deterministic recovery path before this adviser. These findings call for focused validation and state-handling changes rather than proceeding with this decision-model integration. This is a decision not to pursue this implementation; it does not claim that decision-assisted recovery can never be useful. The existing deterministic safeguards and independently merged compatibility fixes remain in place. |
Summary
Implements the opt-in host adapter for #7341. TinyTools owns the typed recovery contract and conservative validation in tinyhumansai/tinytools#61; TinyAgents consumes it in tinyhumansai/tinyagents#375. OpenHuman owns configuration, route/credential snapshots, decision bounds, admission and ephemeral advice through its existing repeated-failure driver.
Keywords remain the default. Compare records metadata without changing results or nudges; Jev advises on unresolved declared read-only failures. Trusted facts, ordinary policy/approval gates, failure accounting and uncertain-write exclusions remain authoritative. Alternate advice uses only the current admitted callable read-only surface; the main model's eventual call still passes the normal tool gates. Classification and alternate requests share a three-second/run-remaining deadline and consume the bounded per-run decision allowance. No automatic re-execution.
Dependencies
Draft only because TinyAgents #375 depends on TinyTools #61 landing first. All referenced gitlinks are published upstream. Mark ready after the dependency changes land; do not merge against unpublished dependency commits.
Verification
Fresh compiled host tests pass: 85 core recovery tests (serial), 158 middleware tests plus one existing ignored test, one embed recovery test, all 15 embed seam tests and six transport tests. The broad parallel recovery filter exposed existing global proxy-state interference; all 85 tests passed from the same compiled executable when serialized. The admission fixture was updated to match the merged harness contract: a denied call terminates the run, so success, denied alternate and hidden/denied call cases execute independently. No production gate was weakened.
Static Rust layout, crate chain, feature forwarding, formatting, generated docs and whitespace checks pass. A fresh TinyHumans host-library build passed. Fresh checks with default features disabled passed both with Jev enabled and with Jev disabled. Instrumented host coverage is still running; no host coverage percentage is claimed yet. Generated architecture references were refreshed against the merged base using the existing capability reporter.
Eleven actual OpenRouter evaluations used synthetic minimized observations, one attempt each, all within the three-second ceiling (462–1755 ms). Transient read, wrong arguments and clear wrong tool returned accepted advice; ambiguous wrong tool abstained at low confidence and unknown evidence abstained. Trusted permission/uncertain-write fixtures made zero endpoint requests, and an injected timeout separately verified evaluator-failure abstention. The clear wrong-tool case also selected the permitted alternate on three additional live repetitions. Aggregate live usage: 8681 input / 1461 output tokens. Credentials, user payloads, raw arguments and provider error text are absent from logs.
A live two-stage classification/alternate flow completed in 1844 ms total under one shared three-second deadline. One additional trial linked the actual compiled OpenHuman TinyHumans adapter and accepted the permitted alternate in 1271 ms. Those eleven adviser-only trials verify TinyTools and the host transport mapping. The host model-loop fixture uses scripted inference and proves alternate admission and subsequent denial without external services. Recovery thresholds remain provisional; no benchmark gain or default-mode rollout is claimed.
A subsequent full live public Embed Runtime/Agent turn used OpenRouter
openai/gpt-4.1-miniand two synthetic in-memory read-only tools. It completed in 6.795 seconds: one arithmetic failure, wrong-tool classification guidance on the next main-model request, one permitted status-reader call, then a ready reply. Three main-model requests used 841 input / 48 output tokens; two Jev requests used 1493 / 251 tokens and took 673 / 604 ms. Boolean request inspection verified the original error was preserved and the advisory nudge was ephemeral. Jev selected none for alternate advice, so this proves classification guidance reaches the live loop, without claiming that an explicit alternate recommendation caused success. This full-turn fixture linked the actual compiled embed runtime/middleware and exact copied host transport mapper; the earlier adapter-only trial separately linked the actual compiled TinyHumans mapper. An initial exploratory full turn is also retained, but may have batched tool calls and is excluded from the sequential demonstration. Across all live trials: 15 Jev requests and five main-model requests, 12,968 input / 2,105 output tokens.Refs #7341