Skip to content

Owned readers stall after a stream is recreated below the saved cursor #110

Description

@derekste

Owned stream readers do not resume if a Redis key is deleted and recreated with IDs below the saved cursor. The exact cursor is correctly preserved across ordinary reads, but there is no stream-reset detection, so the new stream is filtered indefinitely until its IDs exceed the old stream's maximum.

Affected revision: 904e0ef (redis-adapter #109 on #108); the reader logic is also present in 0bb9c35. This affects the IOC release path tracked in fermi-ad/redis-pvxs-ioc#4, #24 and #99, but does not establish the root cause of the historical cache/allocator report.

Reproduction with a private standalone Redis:

  1. Write {test}:value with ID 1000-0.
  2. Use subscribeStream("value", callback, "0-0") and wait for 1000-0.
  3. Delete the stream and recreate it with ID 1-0.
  4. No callback arrives for 1-0 during the four-second regression deadline. The test fails with “recreated stream must resume below the previous cursor”. Writing further IDs below 1000-0 remains stalled.

Expected: observe stream deletion/reset, fence queued data from the old stream, resume the recreated stream, and expose the discontinuity. A mere connection outage must preserve the exact cursor and must not replay old samples. Trim/deletion and connection recovery should be observable without changing the legacy source timestamp encoding.

A focused upstream fix and recovery_test.cpp regression are in progress on dev/stream-recovery-status; the fix PR and passing regression/CI evidence will be linked here. Keep this issue open until the fix merges and validates.

Activity

  1. derekste commented on Sep 29, 2026

    @derekste
    MemberAuthor

    The confirmed stream deletion/recreation stall is addressed in redis-adapter #111, stacked after #108 and #109. The regression fails on 904e0ef and passes on 68e01c7; native lifecycle/recovery/write tests and a private loopback cluster routing check passed. Full Linux run: https://github.com/fermi-ad/redis-adapter/actions/runs/36624795890 (pending at this update).

    The fix detects observed stream reset/deletion, fences old queued callbacks, resumes lower IDs, and exposes separate connection/inspection/continuity counters. It preserves exact cursors over ordinary outages and leaves source timestamp encoding unchanged. Read-only XINFO permissions and the limits of detecting between-probe stream replacement are documented.

    This is a separate confirmed continuity defect. It does not establish the cause of the historical cache/tcache error, and no cache issue is being closed on that basis.

  2. bigsamich commented on Sep 30, 2026

    @bigsamich
    Contributor

    Verified at #111 (0c3b0e6) with default options:

    • the recreated 1-0 is delivered after about 1 s, with streamResets=1 epoch=2;
    • queued callbacks from the old epoch are fenced;
    • an ordinary outage preserves the exact cursor, with exactly-once delivery.

    Caveats found in review:

    • legacy addValuesReader/addListsReader readers still stall (documented);
    • readerProbeMs=0 disables recovery;
    • with cxn.timeout=0, detection never runs, so the stall remains;
    • on a Cluster connection, owned readers read nothing at all.

    As with #105, Fixes #110 only takes effect if #111 targets main when it merges.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions