Skip to content

Retry attempts are invisible: the transcript shows neither that a retry happened nor how many were spent #6796

Description

@7jrxt42BxFZo4iAnN4CX

Problem

Codewhale retries the provider in several places, and by default none of them is visible in the
transcript. An operator watching the TUI cannot tell these three states apart:

  • (a) the turn is retrying and will recover;
  • (b) the turn is retrying and will exhaust its budget;
  • (c) the turn already failed and needs a manual resend ("продолжи").

That ambiguity costs time twice — first while waiting, then while diagnosing after the fact. The
counters already exist in memory (turn.stop_diagnostics.stream_resumes,
transparent_stream_retries, empty_stop_retries, reasoning_only_reprompts), but they reach
neither the transcript nor the turn inspector, and the retry events themselves only go to
crate::logging, which is stderr-only and gated on CODEWHALE_LOG_LEVEL
(crates/tui/src/logging.rs:73).

Current state, verified against upstream/main at 568abae0e:

  • transport retry — callback in crates/tui/src/client.rs:3876 →
    logging::warn("Isolated HTTP retry reason=… attempt=… delay=…s");
  • transparent stream retry — crates/tui/src/core/engine/turn_loop.rs:5879 →
    logging::info("Transparent stream retry {n}/{limit} …");
  • stream-open retry (Turn fails with no retry when the SSE request never receives response headers #6699) — turn_loop.rs:1957, and mid-stream resume — turn_loop.rs:2188:
    Event::status("Reconnecting…") only when attempt == 2, with no attempt number, no limit
    and no reason.

So a single retry is completely invisible; the second and later ones print a bare Reconnecting…;
and the number of attempts never lands anywhere the operator can read afterwards.

Proposed solution

  1. Emit one transcript line per retry attempt using the existing status channel
    (Event::status), naming the kind (transport / transparent / resume / empty-stop), the
    attempt/limit, and a one-line reason. Example below.
  2. Emit a closing line when the turn recovers (recovered after N retries) and when the budget is
    exhausted (gave up after N retries · last: …), so the two cases are distinguishable without
    opening JSON.
  3. Surface the aggregate counters in the turn inspector — attempts per kind, plus whether the budget
    was exhausted or never entered. That last distinction is the one that currently cannot be made at
    all, and it is exactly what makes a post-mortem guesswork.
  4. Respect the existing quiet controls (notifications.quiet, docs/CONFIGURATION.md:2542): a quiet
    run keeps the counts and only suppresses the per-attempt lines.
  5. Headless (codewhale exec, stream-json) needs the same events on the stream channel — a script
    cannot read stderr.

Use case

Example — operator-visible

Today, a turn that retried ten times and recovered is indistinguishable from one that failed
outright:

⚠ Provider returned an empty response
Reconnecting…
[turn ends]

(the only signal is Reconnecting…, from the second attempt onward; it never says which attempt, of
how many, or why; for transport and transparent retries there is no signal at all)

Wanted:

↻ retry 1/10 · provider error frame: "Provider returned an empty response" — nothing streamed, re-issuing
↻ retry 2/10 · transport: connect timeout, delay 4.0s
✓ recovered after 2 retries (4.1 s)

and when the budget runs out:

✗ gave up after 10 retries · last error: Provider returned an empty response

When this matters

Long autonomous runs (Operate / goal runs, unattended overnight work) are exactly where retries
happen most and where the operator is least able to watch. Without the line, a silently retrying
turn and a dead turn look identical, and the only way to reconstruct what happened is to read
turn_outcomes out of the session JSON after the fact.

Alternatives considered

  • Verbose logging only (CODEWHALE_LOG_LEVEL=info): available today, but it is stderr, it is
    off by default, and it does not survive in the transcript the operator scrolls through. It also
    cannot be seen in a headless stream.
  • A dedicated retry pane/rail: heavier than the problem; a status line plus the counters in the
    existing turn inspector covers it.
  • Notifications only (the existing sound/toast path): does not answer "how many attempts, which
    kind, did it recover", and it is the wrong channel for a turn that is still running.

Impact

Every interactive session that hits a flaky provider or a slow proxy. Two concrete costs today:
waiting on a turn without knowing whether it is alive, and post-mortem archaeology — during this
report's own investigation the retry counters had to be reconstructed from session files and code
because neither the transcript nor the log carried them.

Additional context

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    needs-triageNew external report awaiting maintainer triage; repro, logs and version output help

    Projects

    • Status
      Backlog

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions