Skip to content

fix: stabilize onboarding and thread selection e2e - #7268

Merged
senamakel merged 299 commits into
tinyhumansai:mainfrom
senamakel:ci-e2e-green-followup
Oct 10, 2026
Merged

senamakel merged 299 commits into
tinyhumansai:mainfrom
senamakel:ci-e2e-green-followup

Conversation

@senamakel

@senamakel senamakel commented Oct 10, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Expose the runtime tenant helper through the embed host facade so the merged upstream server builds.
  • Make harness initialization E2E checks resilient to optional RPC support and use stable test IDs.
  • Prevent stale async thread loads from overriding later user selections, including worker-thread selection, with a regression test.
  • Stabilize onboarding E2E setup and interactions.

Validation

  • typecheck, lint, format check, coverage tests, build, Rust tests, and clippy passed.
  • Full 64-shard browser E2E matrix completed; the onboarding case that initially caught a missing import passed when rerun after correction.
  • Focused Conversations unit suite: 81 tests passed.

Summary by CodeRabbit

  • Bug Fixes
    • Preserved your selected conversation when thread loading finishes, preventing delayed results from unexpectedly switching the active conversation.
    • Prevented in-progress thread creation and loading from overriding a newer conversation selection. If you select a thread while the list is loading, it remains active even when the response does not include it.
    • Improved app readiness checks when initialization status is unavailable, while continuing to wait for initialization to complete when a status is provided.

@coderabbitai

coderabbitai Bot commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Important

Review skipped

We couldn't safely recover the incremental review. No full review was started, and the last reviewed checkpoint was preserved. Retry later, or explicitly request a full review by commenting @coderabbitai full review.

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: f11ddf29-f7f2-46be-9bf1-7d84c412d674


📥 Commits

Reviewing files that changed from the base of the PR and between ac098ea and 1e7aa0e.



📒 Files selected for processing (4)
  • app/src/store/__tests__/threadSlice.test.ts
  • app/src/store/threadSlice.ts
  • app/test/playwright/helpers/core-rpc.ts
  • app/test/playwright/specs/chat-thread-selection-race.spec.ts


Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 4 remain after this review.




📝 Walkthrough
📝 Walkthrough

Walkthrough

The changes track thread-selection intent in Redux and prevent asynchronous thread results from overriding newer selections. They also update initialization and browser tests, Rust tests and code organization, and CI records.

Changes

Thread selection intent

Layer / File(s) Summary
Redux selection intent
app/src/store/threadSlice.ts, app/src/store/__tests__/threadSlice.test.ts
Thread state tracks selection intent changes. Thread-list fulfillment checks the captured version before clearing or retaining a selected thread missing from the response. Tests cover intent increments and pending requests.
Conversation loading and selection flow
app/src/features/conversations/Conversations.tsx, app/src/pages/__tests__/Conversations.render.test.tsx, app/test/playwright/specs/chat-thread-selection-race.spec.ts
Thread creation and initial loading compare the Redux intent version. Tests cover a selection made while thread loading is pending.

Initialization and browser test updates

Layer / File(s) Summary
Initialization readiness and onboarding
app/src/components/InitProgressScreen/*, app/test/playwright/helpers/core-rpc.ts, app/test/playwright/specs/onboarding-modes.spec.ts
Initialization controls receive test identifiers. The readiness helper handles an unavailable status method. Tests check initialization states and onboarding navigation.
Chat composer and thread interactions
app/test/playwright/helpers/chat-drive.ts, app/test/playwright/specs/chat-composer-primary-slot.spec.ts, app/test/playwright/specs/chat-scroll-stick.spec.ts
Chat tests update new-thread actions, composer input, and thread-selection checks.
Other browser state checks
app/test/playwright/specs/aui-context-usage.spec.ts, app/test/playwright/specs/skills-registry.spec.ts, app/test/playwright/specs/voice-mode.spec.ts
The context-usage assertion changes. Skills setup dismisses the walkthrough earlier, and voice-mode setup waits for the control to be enabled.

Rust code, tests, and CI records

Layer / File(s) Summary
Rust test and tool checks
crates/openhuman-core/src/tools/impl/system/shell_tests_runtime_and_sandbox_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware_failure_policy_class_tests.rs, crates/openhuman-core/src/web3/x402/seams.rs, crates/openhuman-embed/tests/isolation_autonomy.rs, tests/x402_twit_sh_live.rs
Tests update sandbox and denial checks, module feature gates, proxy-condition expression, and x402 request-tool setup.
Turn harness and security-gate changes
crates/openhuman-core/src/agent/tinyagents/harness_assembly*, crates/openhuman-core/src/agent/tinyagents/host/security_gate*
AssembledTurnHarness moves to a submodule. decision_for_outcome moves to the security-gate outcome module.
CI baselines and exemptions
scripts/ci/agent-runtime-boundary-baseline.json, scripts/ci/saas-ambient-baseline.json, scripts/ci/module-pin-exemptions.json, scripts/ci/check-openhuman-rust-layout.mjs
CI records update tracked locations and occurrences. Rust layout pins and module exemptions also change.

Priority: ⬇️ Low

Estimated code review effort: 3 (Moderate) | ~25 minutes

Change: Bug fix

Sequence Diagram(s)

sequenceDiagram
  participant Conversations
  participant ReduxThreadSlice
  participant ThreadListRequest
  Conversations->>ReduxThreadSlice: capture selectionIntentVersion
  Conversations->>ThreadListRequest: loadThreads
  Conversations->>ReduxThreadSlice: setSelectedThread
  ThreadListRequest->>ReduxThreadSlice: fulfill with captured version
  ReduxThreadSlice->>ReduxThreadSlice: preserve newer selection
Loading

Possibly related PRs

  • tinyhumansai/openhuman#6995: Changes initial thread selection before the chat route, connecting to the asynchronous selection flow tested here.

Suggested labels: test, rust-core, infra-ci-release

Suggested reviewers: graycyrus



Merge Risk: ⚪ Minimal · up to 1e7aa

Worker-thread selections remain protected during pending loads. No actionable merge-blocking issue remains after normal checks.

Security Architecture Review

Security architecture risk: 🔵 Low · up to 1e7aa

The change strengthens protection against stale responses replacing newer thread choices. No introduced security boundary bypass was established in the inspected changes. Validation of identity transitions and overlapping requests remains limited.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The demonstrated effect is on browser-side selection, thread rows, and subsequent thread-targeted loading. The inspected production diff does not change RPC credentials or transport, and retaining a row is not evidence of a server authorization grant. Server-side tenant enforcement was not verified by this pass.

Trust Boundaries and Controls

  • observed — The version used to decide whether selection was superseded is captured from local Redux state and appended after spreading the RPC response. A response cannot substitute its own request-version value through that payload spread. This is a selection-ordering control, not an identity or authorization boundary.

Resilience and Maintainability Implications

  • inferred — Selection intent does not serialize loads sharing the same version or establish response ownership across identity/cache resets. The separate proactive creation path also selects its result without this intent guard. The base comparison identifies these as pre-existing limitations, not introduced security findings; identity-transition exposure and overlapping-load behavior lack dedicated test evidence.

Hardening Proposals

  • proposed — As a separate ownership-hardening measure, bind asynchronous thread results to an identity/cache generation distinct from selection intent, and reject results from a previous owner after reset. This would address a pre-existing response-provenance gap rather than remediate a demonstrated regression in this PR.

Pre-merge checks | Passed 4 | Failed 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage Warning Docstring coverage is 66.67% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 18 functions across 25 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check Passed The title clearly and concisely summarizes the main changes: stabilizing onboarding and thread-selection end-to-end tests.
Linked Issues check Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check Passed Check skipped because no linked issues were found for this pull request.



✨ Finishing Touches 💡 2
⚔️ Resolve merge conflicts 💡
  • Resolve merge conflict in branch ci-e2e-green-followup


🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR



  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

A rabbit checks the threads in flight,
And keeps the chosen one in sight.
The buttons gain a test-id tag,
While Rust and CI update their flags.
The browser waits, then clicks with care,
A tidy patch hops through the lair.

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-10T18:43:24.976172Z 1da9b39 New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@tinysweeper

tinysweeper Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

⚠️ Review failed for 1da9b39b2753. pull request exceeds review limits: 611 changed files (limit 500), 60089 changed lines (limit 50000)

Last completed report

Tiny Sweeper review

This pull request stabilizes onboarding and thread-selection e2e behavior plus supporting Rust refactors. It replaces a component-local selection-intent ref with a Redux `selectionIntentVersion` so stale async thread loads/creations are dropped, adds stable testids to the harness-init dialog with a state-aware Playwright helper, and restructures Rust modules. Reviewers report most earlier findings resolved; two open concerns remain: legacy line-count pins in the layout checker were raised again and two new vendored-pin exemptions were added, and the aui-context-usage 'Output' assertion was removed without a corresponding adapter change in this diff.

State: Changes requested
Priority: high
Reviewed head: 1e7aa0e1f9cc
Updated: 1791645385 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 9 Active findings 7
Tests 16 Noted findings 0
Documentation 0 Resolved findings 55
Configuration 3 Pending checks/questions 4

Completeness: Complete
Test assessment: Test coverage is assessed from changed tests and lane evidence; execution is not claimed without trusted check data.

What changed

No supported behavioral explanation was produced.

Features

  • Modified — Redux-backed thread selection intent guards against stale async continuations: Explicit sidebar selections and thread clears are no longer overwritten when an older `loadThreads` response or thread-create continuation resolves; superseded loads are ignored or merged so the newer selection survives. (app/src/store/threadSlice.ts#const threadSlice = createSlice({, app/src/store/threadSlice.ts#function appendMessageToCache(, app/src/features/conversations/Conversations.tsx#const Conversations = ({)
  • Modified — Harness init dialog exposes stable test IDs and state-appropriate dismissal: E2E suites can reliably locate the dialog and the correct continuation button (run-in-background vs continue-anyway) and tolerate cores lacking the optional init status method. (app/src/components/InitProgressScreen/InitProgressScreen.tsx#export default function InitProgressScreen({, app/test/playwright/helpers/core-rpc.ts#export async function waitForAppReady(page: Page): Promise<void> {)
  • Modified — Context-usage breakdown drops the 'Output' prompt partition: The e2e assertion now expects only System prompt, Tool schemas, and Your input, on the rationale that provider output may be trimmed from the final request; a reviewer notes no corresponding adapter change appears in this diff. (app/test/playwright/specs/aui-context-usage.spec.ts#test.describe('assistant-ui context usage on the chat path', () => {)
  • Internal refactor — TinyAgents assembly and security-gate modules restructured: `AssembledTurnHarness` moves to `harness_assembly/assembled.rs` and `decision_for_outcome` to `security_gate/outcome.rs` with visibility narrowed to `pub(in crate::agent::tinyagents)`; behavior is unchanged. (crates/openhuman-core/src/agent/tinyagents/harness_assembly/assembled.rs, crates/openhuman-core/src/agent/tinyagents/host/security_gate/outcome.rs, crates/openhuman-core/src/agent/tinyagents/harness_assembly.rs#use crate::agent::tinyagents::verify_before_finish;)

Tests

  • component-render — The Conversations render suite adds a test that an explicit sidebar selection made while initial thread loading is pending survives the pending load resolving, with a new `selectionIntentVersion: 0` field in the empty thread state.: Pins the behavior at the component/store integration level; would fail if the selection were overwritten. (app/src/pages/__tests__/Conversations.render.test.tsx#describe('Conversations — smoke render (Replace the onboarding bot with a managed guided walkthrough using react-joyride #1123 welcome-lock removal)', () => {, app/src/pages/__tests__/Conversations.render.test.tsx#async function openSidebar() {)
  • slice-unit — New threadSlice thunk tests verify that clearing all threads invalidates a superseded load (threads stay empty and selection null) and that a newer selection is preserved when an older thread-list response omits it; synchronous reducer tests assert `selectionIntentVersion` increases on `clearAllThreads` and `clearStaleThread`.: Directly exercises the version-invalidation semantics; would fail if the reducer applied stale responses. (app/src/store/__tests__/threadSlice.test.ts#describe('threadSlice loadThreads thunk', () => {, app/src/store/__tests__/threadSlice.test.ts#describe('threadSlice synchronous reducers', () => {)

Findings

  • medium · critique · Ignore superseded loads after clearing a stale thread — A `loadThreads` request can start while `t-1` is present, then `clearStaleThread('t-1')` removes it and increments `selectionIntentVersion`, leaving another thread such as `t-2` in (app/src/store/threadSlice\.ts:621)
  • medium · critique · Disambiguate the continuation-button locator — This locator is used by `isVisible()` inside the polling loop and by the eventual click. If the overlay renders two continuation controls with this test id, Playwright treats the l (app/test/playwright/helpers/core\-rpc\.ts:283)
  • medium · critique · Handle non-terminal initialization states without timing out — If `harness_init_status` returns a valid non-terminal state such as `pending` or `queued`, this code proceeds into the ten-second polling loop even when the initialization dialog i (app/test/playwright/helpers/core\-rpc\.ts:271)
  • medium · tests · Do not raise the legacy line-count allowances again — The per-file line-count exemptions in the layout checker were bumped again in this pull request (runtime_session.rs 1406→1412, lifecycle 1306→1323, ops.rs 1206→1209, progress_bridg (scripts/ci/check\-openhuman\-rust\-layout\.mjs:43)
  • high · e2e · End-to-end job `Storage e2e on MongoDB` will not run on this change — `Storage e2e on MongoDB` in `.github/workflows/storage-mongodb.yml` will not run for this pull request: its workflow's `paths:` filter matches nothing this pull request changed. Th (\.github/workflows/storage\-mongodb\.yml:27)
  • medium · e2e · Keep legacy pins from increasing — Every previously pinned file allowance was raised again this revision (runtime_session 1406→1412, lifecycle 1306→1323, runner 1145→1152, tools/ops 1206→1209, progress_bridge 1304→1 (scripts/ci/check\-openhuman\-rust\-layout\.mjs:43)
  • medium · e2e · Restore or justify removal of the Output breakdown assertion — The end-to-end assertion `await expect(popover).toContainText('Output')` was deleted and the comment rewritten to claim the breakdown now emits only three rows, because "output is (app/test/playwright/specs/aui\-context\-usage\.spec\.ts:121)

Resolved this pass

  • Invalidate selection when clearing all threads
  • Ignore superseded loads after clearing all threads
  • Normalize missing persisted selection versions
  • Recognize standard unknown-method errors
  • Harness helpers match two init continuation buttons
  • Two buttons share the harness-init-continue testid
  • Two buttons still share the harness-init-continue testid
  • Select the visible init continuation button
  • Track the continuation button after status changes
  • Invalidate selection when clearing all threads
  • Ignore superseded loads after clearing all threads
  • Normalize missing persisted selection versions
  • Recognize standard unknown-method errors
  • Harness helpers match two init continuation buttons
  • Two buttons still share the harness-init-continue testid
  • Two buttons share the harness-init-continue testid
  • Select the visible init continuation button
  • Track the continuation button after status changes
  • Two buttons share the harness-init-continue testid
  • Select the visible init continuation button
  • Harness helpers match two init continuation buttons
  • Dismiss the walkthrough before clicking Connections
  • Preserve forced clicks for transiently obstructed thread buttons
  • Use a stable locator for the new conversation action
  • Invalidate selection when clearing all threads
  • Ignore superseded loads after clearing all threads
  • Normalize missing persisted selection versions
  • Recognize standard unknown-method errors
  • Run the production timeout test in a supported configuration
  • End-to-end job `Storage e2e on MongoDB` will not run on this change
  • Do not add new legacy exceptions
  • Keep legacy pins from increasing
  • Invalidate selection when clearing all threads
  • Ignore superseded loads after clearing all threads
  • Normalize missing persisted selection versions
  • Two buttons share the harness-init-continue testid
  • Select the visible init continuation button
  • Track the continuation button after status changes
  • Recognize standard unknown-method errors
  • Harness helpers match two init continuation buttons
  • Preserve forced clicks for transiently obstructed thread buttons
  • Dismiss the walkthrough before clicking Connections
  • Use a stable locator for the new conversation action
  • Preserve forced clicks for transiently obstructed thread buttons
  • Dismiss the walkthrough before clicking Connections
  • Two buttons share the harness-init-continue testid
  • Select the visible init continuation button
  • Invalidate selection when clearing all threads
  • Track the continuation button after status changes
  • Recognize standard unknown-method errors
  • Run the production timeout test in a supported configuration
  • Use a stable locator for the new conversation action
  • Normalize missing persisted selection versions
  • Ignore superseded loads after clearing all threads
  • Harness helpers match two init continuation buttons

Pending checks: Rust E2E (mock backend), Build Playwright E2E Artifact, E2E (Playwright / web lane), Desktop E2E (full suite, 3 OS)

Before merge

  • Address End-to-end job `Storage e2e on MongoDB` will not run on this change (\.github/workflows/storage\-mongodb\.yml).
  • Wait for Rust E2E (mock backend), Build Playwright E2E Artifact, E2E (Playwright / web lane), Desktop E2E (full suite, 3 OS).

How this fits together

flowchart LR
  n0["ThreadState<br/>changed<br/>1 finding"]:::flagged
  n1["appendMessageToCache<br/>changed<br/>1 finding"]:::flagged
  n2["expect"]:::impacted
  n3["toBe"]:::impacted
  n4["bootAuthenticatedPage"]:::impacted
  n5["dismissWalkthroughIfPresent"]:::impacted
  n1 -->|uses| n0
  n4 -->|uses| n2
  n5 -->|uses| n2
  n5 -->|calls| n3
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading
Agent review details

critique

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: Reviewed 3 files; 3 findings. (11 earlier finding(s) still open) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: app/src/store/threadSlice\.ts — Ignore superseded loads after clearing a stale thread
  • Evidence: app/test/playwright/helpers/core\-rpc\.ts — Disambiguate the continuation-button locator
  • Evidence: app/test/playwright/helpers/core\-rpc\.ts — Handle non-terminal initialization states without timing out

security

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: Reviewed 3 files; 0 findings. (11 earlier finding(s) still open) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Positive: The thread-selection race is covered at three levels — slice unit, render, and a Playwright spec that deliberately holds the threads_list response in flight — each of which would fail under the old clobbering behavior.
  • Positive: The InitProgressScreen tests assert both the presence and the absence of each action button per dialog state, pinning the state-button mapping rather than only that a button exists.
  • Positive: The `seams.rs` negation rewrite in `direct_connection_allowed` is a De Morgan equivalence, preserving the original semantics.
  • Positive: The e2e helper tracks which continuation button is visible via testids rather than ambiguous role/text locators, addressing earlier ambiguity findings.
  • Positive: Most earlier review findings were resolved in this revision, including distinct init-dialog locators, unknown-method tolerance, selection-intent invalidation on clear, and feature-gating of module-dependent tests.
  • Lane summary: This increment resolves most of the earlier findings: the harness-init dialog now has distinct, stable testids matched by both the component tests and the E2E helper, unknown-method errors from the init-status RPC are tolerated, thread-clearing and superseded loads invalidate the selection intent in Redux with reducer tests, and the production timeout test is feature-gated. The change itself looks sound: the selectionIntentVersion migration is behaviour-backed by new unit and Playwright race tests, and the x402 boolean rewrite is a correct De Morgan equivalent. One prior concern remains: the diff again raises the legacy line-count allowances in the layout checker rather than holding them flat. (2 earlier finding(s) still open) (1 observation(s) grouped into shared inline comments) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: scripts/ci/check\-openhuman\-rust\-layout\.mjs — Do not raise the legacy line-count allowances again

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: This revision resolves the earlier thread-selection and harness-init findings: clearing/clearing-stale threads now bumps selectionIntentVersion with tests, superseded loads are ignored via a version captured at request time, the init helper uses distinct stable testids and tracks the visible button, and unknown-method errors are recognized. What remains are the two pin/exception concerns from earlier revisions, which the latest diff does not change; I could not verify the previously reported CI job concern from the diff. (7 earlier finding(s) still open) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

e2e

  • Conclusion: Failure
  • Scope reviewed: all assigned evidence
  • Lane summary: This revision fixes nearly all previously raised findings: the init dialog now has distinct, stably-locatable testids and the Playwright helper drives whichever button is visible; walkthrough dismissal precedes the Connections click; thread clearing and superseded loads bump a selectionIntentVersion that both the slice and Conversations honour, and the new Playwright race spec drives that behaviour end to end; the unknown-method fallback is in place; the timeout-class tests are gated behind the modules feature. Two concerns remain: the layout-line pins grew again and two new vendored-pin exemptions were added, and the aui-context-usage e2e assertion that the breakdown contains an 'Output' row was deleted without any corresponding adapter change in this pull request to justify it. The Storage (MongoDB) workflow not triggering is correct here — no storage path is touched. (1 finding discarded for not matching a changed line) Waiting on end-to-end jobs: `Rust E2E (mock backend)`, `Build Playwright E2E Artifact`, `E2E (Playwright / web lane)`, `Desktop E2E (full suite, 3 OS)`. (3 earlier finding(s) still open) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Unresolved questions/checks: Rust E2E (mock backend), Build Playwright E2E Artifact, E2E (Playwright / web lane), Desktop E2E (full suite, 3 OS)
  • Evidence: \.github/workflows/storage\-mongodb\.yml — End-to-end job `Storage e2e on MongoDB` will not run on this change
  • Evidence: scripts/ci/check\-openhuman\-rust\-layout\.mjs — Keep legacy pins from increasing
  • Evidence: app/test/playwright/specs/aui\-context\-usage\.spec\.ts — Restore or justify removal of the Output breakdown assertion
Evidence and run details
  • Models: gpt-5.6-luna, glm-5.3-flash
  • Spend: $0.021991
  • Tokens: 459752 input · 28950 output · 44438 cached · 0 embedding
Head State Pass summary
7c65b1ac6f28 changes requested 8 active finding(s), 16 resolved finding(s) (at 1791635754)
16d83ae06071 changes requested 12 active finding(s), 179 resolved finding(s) (at 1791641783)
ac098eac3216 changes requested 5 active finding(s), 86 resolved finding(s) (at 1791643692)
012d8f10fed5 changes requested 3 active finding(s), 59 resolved finding(s) (at 1791644609)
1e7aa0e1f9cc changes requested 7 active finding(s), 55 resolved finding(s) (at 1791645385)

tinysweeper 0.1.0

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

             $0.0466 · 646,519 in / 33,425 out · 86,556 cached (13%) · gpt-5.6-luna, glm-5.3-flash
critique:    $0.0274 · 337,338 in / 18,318 out · 45,973 cached (14%) · gpt-5.6-luna, glm-5.3-flash
security:    $0.0186 · 245,703 in / 9,395 out  · 35,783 cached (15%) · gpt-5.6-luna
tests:       $0.0001 · 14,995 in  / 550 out    · 1,536 cached (10%)  · glm-5.3-flash
description: $0.0001 · 14,296 in  / 650 out    · 1,408 cached (10%)  · glm-5.3-flash
e2e:         $0.0002 · 19,121 in  / 1,994 out  · 1,728 cached (9%)   · glm-5.3-flash

Comment thread app/test/playwright/helpers/chat-drive.ts Outdated
Comment thread app/test/playwright/specs/skills-registry.spec.ts
Comment thread app/src/components/InitProgressScreen/InitProgressScreen.tsx Outdated
@tinysweeper tinysweeper Bot added the priority: p2 Soon. Real but survivable — a rough edge, a gap, a thing that will bite later. label Oct 10, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 08cd9d4b45

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread app/test/playwright/specs/onboarding-modes.spec.ts Outdated
Comment thread app/src/features/conversations/Conversations.tsx Outdated
senamakel and others added 24 commits October 10, 2026 06:46
Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…s/openhuman-core/src/profiles/m

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…sts.rs

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…sts.rs

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…nstall.rs

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Introduce an approval gate that intercepts tool calls and requires
explicit user approval before they proceed. Gate state is tracked
separately so pending approvals can be queried and resolved.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Moved the host profile logic out of the profiles module into its own
host.rs file so the host-specific behaviour can be maintained and tested
in isolation. No functional changes were made.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Introduce a host module in the MCP layer to manage server lifecycle and
connections. This provides the foundation for coordinating MCP servers
within the core.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…s/openhuman-core/src/profiles/m

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Cleaned up unused imports and dead code across the session store, cron origin delivery, and agent storage modules to keep the codebase free of compiler warnings.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Introduce lifecycle operations for profiles so they can be created, activated, and torn down through a dedicated module. The ops layer now delegates to these routines, keeping profile state transitions consistent across callers.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Adds tests exercising the profile lifecycle paths, including activation and
deactivation transitions, to lock in the expected behaviour and guard against
regressions.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Document that current_agent_id is genuinely per-agent and never a tenant
key, since a SaaS profile's default agent has none. Tables and stores that
must keep tenants apart should key on current_tenant instead, as the
remaining callers only log the agent id.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Adds property-based tests and host-level tests covering gateway profile
resolution, exercising edge cases that the existing unit tests did not
reach.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add coverage for the approval gate's tenant scoping so that approvals
issued for one tenant cannot be reused by another. The new tests
exercise the gate's tenant checks directly to guard against regressions
in cross-tenant approval handling.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Adds test coverage for the MCP host agent and the skills write root
behaviour, exercising the paths that were previously untested.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Adds coverage for the write_root skill path, exercising how skill files are
written to the root directory so regressions in that flow are caught.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Reformat long log and function signatures, collapse a short method
chain, and wrap an assert so the touched files match rustfmt output.
No behaviour changes.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Moved profile lifecycle logic into its own module to keep the profiles
namespace organized and make the lifecycle behaviour easier to locate.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
… roots

Add tests asserting that desktop session keys and transcript roots stay
unchanged without a profile, that an explicit session agent overrides the
default key and root, and that two profiles' default agents resolve to
distinct keys and transcript stores.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Update the ambient-context baseline to reflect shifted line numbers and drop two entries that no longer match any occurrence.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Document the profile system's purpose and structure so contributors can understand how profiles are defined and used within the core crate.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
coderabbitai[bot]
coderabbitai Bot previously approved these changes Oct 10, 2026
senamakel and others added 26 commits October 10, 2026 18:29
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Update the app lockfile to pick up newer versions of async-imap, imap-proto, base64, curve25519-dalek, reqwest, and the tinychannels, tinymemory, and tinywallet crates, along with a consolidated windows-sys 0.60.2 and the removal of the now-unused nom 7 and minimal-lexical entries.

Auto-committed-on: macbook
Co-authored-by: Medulla <medulla@tinyhumans.ai>
feat(ui): centralize toast messages and history
Combine the separate default and named imports from threadSlice into a single statement to tidy up the test file's import block.

Auto-committed-on: macbook
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Move the threadGoal and threadTodos reducer imports into alphabetical
order alongside the other store imports in the conversation test files,
and expand the combineReducers call in the approval test onto multiple
lines. No test behaviour changes.

Auto-committed-on: macbook
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…smatch

Auto-committed-on: macbook
Co-authored-by: Medulla <medulla@tinyhumans.ai>
chore(i18n): remove unused translation keys
# Conflicts:
#	crates/openhuman-core/src/web3/x402/seams.rs
Revamp conversation UI with assistant-ui and stabilize rendering
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…fault

feat(storage): sqlite by default, with a one-shot import of the small stores
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ts-6905

fix(inference): trust custom provider CA certificates
feat(embed): add dynamic runtime APIs and verified documentation
fix(tauri): allow WebSocket connections to remote cores
fix(security): block literal credential paths in command tools
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@senamakel
senamakel merged commit fb558fe into tinyhumansai:main Oct 10, 2026
12 of 15 checks passed

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

"expect": "v0.7.4-6-gaccb920a",
"reason": "The vendored tree includes six unreleased commits; the shipped v0.7.4 artifact remains the last published runtime module. Remove this exception when the next artifact is published and both pins advance."

P1 Badge Refresh exemptions after advancing module submodules

In the self-hosted static plan (scripts/ci/self-hosted/lanes-plan.mjs:236-239), module-pins is always run, but this exemption still describes the pre-merge wallet commit accb920a/v0.7.4 even though the final gitlink is 6ae706c and the registry is v0.8.0; the tinychannels entry likewise describes 5e3b2044/v0.1.12 while its final gitlink/registry are 642687bb/v0.1.13. git describe -h documents --abbrev=<n> as “use digits to display object names,” and classifyPin rejects both stale and changed exemptions, so the always-on lane cannot pass with initialized submodules. Remove reconciled entries or update both exact descriptions and reasons for the final pins.

AGENTS.md reference: AGENTS.md:L612-L614

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

infra-ci-release CI, release automation, packaging, build containers, and test harnesses. priority: p1 Next. Wrong behaviour a user will hit, or a security weakness behind a condition. rust-core Core Rust runtime in src/: CLI, core_server, shared infrastructure. test Test additions, fixes, or harness work.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant