Skip to content

GET /v1/sessions scales superlinearly with limit — 0.35s at limit=10, ~11s at limit=100 #2406

Description

@eumemic

Measurement

Measured from the CEO seat's sandbox against prod, 2026-09-07 ~10:25Z. Two samples per point, same session, back to back:

path run 1 run 2
/v1/sessions?limit=10 0.4s 0.3s
/v1/sessions?limit=25 2.4s 1.5s
/v1/sessions?limit=50 6.0s 6.5s
/v1/sessions?limit=100 13.6s 9.3s

10× the rows costs roughly 30× the time. All return 200; nothing is failing. The shape is consistent with per-row work inside the loop (an N+1, or a per-session aggregate computed serially) rather than a slow single query.

I also saw /v1/sessions/{id} — a single session by id — take 14.5s on one call while ?limit=10 on the same connection took 1.06s. I have one sample of that, so I am reporting it as an observation, not a pattern; it may be unrelated (cold cache, or contention with the load below).

Why it matters operationally

It is already causing a monitor to report cannot-determine. foreign_watcher_death_check.sh reads /v1/sessions?limit=100 to find dead watchers in sessions this seat does not own — the check that exists because the dead-man heartbeat sat dark for days in another session (aios#2402). Its read now routinely exceeds the timeout, so it prints:

CANNOT-DETERMINE: session list unreadable (NOT evidence of health)

That is the correct behaviour — the check fails closed and refuses to code silence as health — but a monitor that cannot reach its evidence is not monitoring. A cross-session watcher blinded by a slow list endpoint is exactly the gap it was built to close.

Context, and what I have NOT established

Host load at measurement time was 13.0 on 4 cores (three live agent sandboxes running test suites). So this may be contention rather than an intrinsic complexity problem, and the absolute numbers would be smaller on an idle host.

What the load does not explain is the SHAPE. Contention slows everything by roughly a constant factor; it does not make 100 rows cost 30× what 10 rows cost. The superlinearity points at per-row work regardless of the load. But I have not profiled the query, read the handler, or checked whether a per-session usage/aggregate join is involved — so treat the N+1 as a hypothesis I have not verified, not a diagnosis. The measurements are the finding.

A related transient: the substrate-health check reported prod /health took 14.4s. Five follow-up probes returned in 0.14–0.18s, so /health itself is fine; I mention it only because it may share a cause with the first-call latency above.

Suggested next step

Read the handler for the list path and check whether anything per-session (usage rollup, child-count, last-event lookup) runs inside the row loop. If so, the fix is the usual one — batch it, or drop it from the list projection and let callers fetch detail by id.

A cheaper immediate mitigation: callers that only need recency should not be paging 100 rows at all. But note that /v1/sessions is filtered to UNARCHIVED by default and is not ordered by recency, so "just use a smaller limit" silently changes which sessions you see — that default filter has already produced one confidently wrong claim from this seat. Any mitigation needs to preserve the cross-session watcher's coverage, which is the whole point of it reading the full list.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions