fix(aios_connectors/resolver): treat cross-account session as DETACHED, not SESSION_MISSING, after reparent - #2378
Conversation
…D, not SESSION_MISSING, after reparent `reparent_connection` (PR #696) rewrites child rows' `account_id` to the destination but leaves their `session_id` / `target_id` pointing at source-account sessions (sessions are account-scoped and are not reparented). Post-reparent, the resolver returned the stale cross-account id with `drop=None`; `handle_inbound` then hit `append_event`'s account-scoped seq allocation (0 rows) -> `NotFoundError` -> `InboundDrop.SESSION_MISSING` -> HTTP 404. The connector-http runner treats 404 as non-fatal and drops/acks, so every previously-routed chat on the reparented connection lost messages under a generic `session_missing` reason that did not point at reparent as the cause. The reparent docstring's risk model only considered children whose `account_id` is *left on source* (the case it fixes); it never considered the opposite — children whose `account_id` is *moved to destination* while their `session_id` is *left on source* — which is exactly what the carry-over CTE produces. `_session_is_archived` (PR #541) conflated a 0-row (cross-account) session lookup with "live", letting the cross-account id sail past the `DETACHED` guard the docstring credits with preventing a bad outcome and into the 404. Fix: `_session_is_archived` now treats a 0-row session the same as an archived one (return `True`), so the post-reparent inbound surfaces as `ResolveDrop.DETACHED` -> `InboundDrop.DETACHED` (HTTP 422) — the terminal signal the reparent docstring already says is the expected outcome when a child points at the source account, and the read-time tenancy re-check the carry-over tests credit the resolver with performing. This is a diagnostic improvement (it does not restore delivery to pre-reparent chats; that requires a session-reparent primitive, out of scope). Scope honesty: under the aios-connector-http runner the fix changes the dropped reason from `session_missing` to `detached` but not whether the message is dropped — both 404 and 422 are non-fatal there. The operator-facing gain is that the per-message `connector.inbound.refused` log now carries `drop_reason=detached`, which points at reparent rather than a generic "session not found". Tests: new `tests/integration/test_resolver_reparent_cross_account.py` pins tier-1 ledger + tier-3 single_session + `handle_inbound` returning DETACHED for cross-account session ids post-reparent, plus a live same-account control. New `tests/e2e/test_reparent_cross_account_inbound.py` drives the full multipart POST to `/v1/connectors/runtime/inbound` over a live uvicorn socket, pinning the 422 `drop_reason=detached` (not 404) response. Both are red-then-green verified. The false "the resolver re-checks tenancy at read time" comment in `test_reparent_unique_index.py` is corrected to cross-reference the test that now actually verifies it. Co-authored-by: Detail <noreply@detail.dev>
|
Held — and this one is NOT held for the reason the label suggests. Sweeping the ten PRs carrying
The other nine each carry a posted verdict (seven pass, two fail). This PR carries the label that means "a machine reviewed it and now a human decides" — but the first half of that never happened. The label was doing the work of a verdict. I found it only because I checked each PR's review comment individually instead of trusting the shared label. A batch labelled uniformly is not uniform, and the gate label is the thing that made them look alike. DispositionHeld, with the reason corrected: this needs an uncorrelated review before it needs a merge decision. It should not inherit the "just needs a human to say go" framing from its neighbours — there is no green to approve. Also true of all ten, including this one: the branch is No fix-round dispatch: the fixround driver is deliberately disabled under the aios#2396 spend freeze. When the freeze lifts, this PR's first step is review, not fix and not merge. |
|
Relabelled: Re-verified today: this PR has no adversarial review of any kind.
CI is green at head Green CI on a never-reviewed PR is not a merge-ready PR. CI proves the tests that exist pass; it says nothing about whether the change is correct, whether its tests actually discriminate, or whether it introduces the defect it set out to fix. That judgment is what the uncorrelated review exists to supply, and here it has simply never happened. Asking a human to approve a merge on a diff no reviewer has examined would make the approval latency without detection — the exact thing the chairman retired his own merge gate to eliminate. The label promised a decision that could not responsibly be made. Correct next step is a review, then (if it passes) merge under the standing rule. Review dispatch is paused under the aios#2396 spend freeze, so this waits — deliberately, not stranded. |
Uncorrelated adversarial review: PASS as scoped — with two caveats that must not be lost at merge.This PR had no adversarial review of any kind (0 review comments, 0 formal reviews) for 8 days. Dispatched one rather than merging it unreviewed or holding it indefinitely. The reviewer had no hand in authoring it and returned its verdict privately; findings are reproduced here in full. The defect is real — reproduced by execution, not readingNo Docker in the reviewer's sandbox, so it ran a local Postgres 15 and pointed the suite at it, bypassing the testcontainer fixture. Reverting only the resolver hunk at head Observed over a live uvicorn socket, not asserted against an internal enum. The stale-pointer premise is confirmed structurally too: the reparent CTE rewrites Mutation evidence — three mutants, all killed, kills tier-resolved
Red→green baseline: reverting the hunk fails 3 of 4 new integration tests + 1 of 2 new e2e; restoring returns 4/4 and 2/2. Regression sweep: 94 passed, 13 skipped (all Docker-gated), 0 failed. The positive control genuinely discriminates: Caveat 1 — the reported harm is NOT closed
Caveat 2 — a same-class hole remains, and it poisons the ledgerTier-2 On my own framingI briefed the reviewer with a It also self-corrected an intermediate belief of its own: it suspected CI's green contexts never executed the new e2e file because of the DispositionMergeable as a diagnostic improvement. Relabelling |
…r-treat-cross-account-s-c2d3b0
Code reviewVerdict: PASS — scoped correctly, and the central invariant holds on inspection. Head verified: Scope of re-verificationI confirmed the substantive diff is unchanged from the head the prior eumemic-bot review examined, so its red→green baseline and mutation evidence carry over rather than being re-run:
CI at this head: 9/9 green, including The load-bearing property, checked directlyThe change is only defensible if the resolver's refusal predicate is the exact complement of
Refuse ⟺ ¬(row exists for this account ∧ not archived) ⟺ ¬(append would succeed). Exact complement, same The 0-row case genuinely arises: the reparent CTE ( Claims in the PR body I spot-checked
Non-blocking observations (no fix required for this PR)
What I could not evaluateThe sandbox has no Docker and no installed project environment (no Blocking issues: none. |
Detail bug report: View on Detail
Summary
After
reparent_connectionmoves a connection across accounts, the carry-over CTE rewrites child rows'account_idto the destination but leaves theirsession_id/target_idpointing at source-account sessions (sessions are account-scoped and are not reparented). The resolver's_session_is_archivedconflated a 0-row (cross-account) session lookup with "live", so it returned the stale cross-account id withdrop=None;handle_inboundthen hitappend_event's account-scoped seq allocation (0 rows) →NotFoundError→InboundDrop.SESSION_MISSING→ HTTP 404. Since the connector-http runner treats 404 as non-fatal, every previously-routed chat on the reparented connection silently dropped under a genericsession_missingreason that did not point at reparent as the cause.The reparent docstring's risk model only considered children whose
account_idis left on the source (the case it fixes); it had a blind spot for the opposite — children whoseaccount_idis moved to the destination while theirsession_idstays on source — which is exactly what the CTE produces.Fix:
_session_is_archivednow treats a 0-row (cross-account) session the same as an archived one (returnsTrue), so the post-reparent inbound surfaces asResolveDrop.DETACHED→InboundDrop.DETACHED(HTTP 422) — the terminal signal the reparent docstring already says is the expected outcome when a child points at the source account, and the read-time tenancy re-check the carry-over tests credit the resolver with performing.This is a diagnostic improvement: under the aios-connector-http runner it changes the dropped reason from
session_missingtodetachedbut not whether the message is dropped (both 404 and 422 are non-fatal there). The operator-facing gain is that the per-messageconnector.inbound.refusedlog now carriesdrop_reason=detached, which points at reparent rather than a generic "session not found". It does not restore delivery to pre-reparent chats — that requires a session-reparent primitive, explicitly out of scope.Substrate state changes
None — code-only change.
Test plan
New coverage (red-then-green verified):
tests/integration/test_resolver_reparent_cross_account.py— pins tier-1 ledger, tier-3 single_session, andhandle_inboundreturningDETACHED(notSESSION_MISSING) for cross-account session ids post-reparent, plus a live same-account session control that still routes (guards against over-eager DETACH).tests/e2e/test_reparent_cross_account_inbound.py— drives the full multipart POST to/v1/connectors/runtime/inboundover a live uvicorn socket, asserting the HTTP response is 422 withdrop_reason=detachedafter reparent (was 404 pre-fix), and that a live same-account inbound still returns 201.test_reparent_unique_index.pyto cross-reference the test that now actually verifies that re-check.No regressions: existing resolver archived-session tests (the archived-session
DETACHEDpath is unchanged by this fix), the reparent carry-over suite, the per_chat spawn path, the full unit suite (5898 passed), and the full integration suite (1138 passed) all green. mypy, ruff, and the pooled-await lint are clean.Live-traffic verification run: Confirmed the end-to-end POST over a real uvicorn socket is red before the fix (HTTP 404,
error_type: not_found— exactly the reported surface) and green after (422,drop_reason=detached).Could not fully verify the "per_chat new chat delivers post-reparent" scenario: I attempted it and found a related, distinct cross-account defect —
reparent_connection's carry-over does not rewritebindings.session_template_id, andsession_templatesare account-scoped and not reparented, so a reparented per_chat connection whose template stayed on the source account raisesNotFoundError: session template ... not found(a 500) on the next new-chat inbound viaget_session_template. This is the same defect class as thesession_idbug but forsession_template_id, falls outside the recommended fix (which targets only_session_is_archived), and is worth filing separately. The intended guarantee — that this fix does not regress the per_chat spawn path — was verified via the existing per_chat spawn integration test (the fix doesn't touch_spawn_per_chat_session).Risk / rollback
Low — the change only widens
_session_is_archivedto returnTruefor one additional case (0-row lookup) that previously returnedFalse; the live-session path is unchanged (theorshort-circuits onrow["archived_at"] is not None). Revert the single resolver change to roll back.Automatic Fixes PRs can be configured here.