RedisAdapter version or revision
The mechanism is on main at ff2d5b8. Reproduced below at the #108 → #109 → #111 stack head 0c3b0e6.
Redis version and deployment mode
Redis 7.4.11 standalone (TCP).
Observed behavior
RedisConnection::connect() publishes what it built even when the attempt failed:
{
std::lock_guard<std::mutex> lk(_mtx);
_cluster = cluster; // both null after a failed attempt
_singler = singler;
}
if (cluster || singler) return true;
RedisAdapter::reconnect() runs this from its background thread whenever any call fails: a write, a get, connected(), or the watchdog. During an outage the attempt usually runs while the server is still down. It therefore replaces the existing redis-plus-plus client with null pointers, although that client would re-dial its broken pooled sockets by itself on next use.
From then on every RedisConnection call fails locally without touching the network or logging. That includes the reader loop, which #108 changed to retry every 50 ms on the premise that "redis++ reconnects its socket on the next read". Nothing recovers until some later call triggers a successful reconnect.
In practice:
- A process that makes even one adapter call during an outage stops receiving subscription data when the server comes back.
- The first write after recovery fails, even though the server has been back for seconds.
- Consumers that poll
connected() recover within one poll period. Listen-only consumers do not recover at all.
Expected: a failed reconnect attempt leaves the existing client in place, so redis-plus-plus's own re-dial can recover.
Reproduction
One reader on {rst}:r with cxn.timeout = 300. The server is shut down for about 1.5 s and started again. One write is attempted while it is down. Then 3 entries are written to the subscribed stream from another client.
write while the server is down -> RA_NOT_CONNECTED
listen-only for 3 s after the server came back: reader delivered 0 of 3 new entries
FIRST write after the restart -> RA_NOT_CONNECTED
reader delivered 3 of 3 after that
Control, identical except that no adapter call is made while the server is down:
listen-only for 3 s after the server came back: reader delivered 3 of 3 new entries
writes that failed after the server had been back for >3 s: 0
In other words, the adapter's own reconnect is what breaks recovery.
The same mechanism turns a MISCONF episode into a read blackout, because Redis also refuses PING under MISCONF. One connected() probe fails, the reconnect's own PING fails, and the clients are nulled. Every read then returns RA_NOT_CONNECTED, although the server is still serving reads.
Build and runtime environment
- Ubuntu 24.04, gcc 13.3
- pinned redis-plus-plus 1.3.15 and hiredis 1.3.0
- a private Redis 7.4.11 in a container
Suggested fix
Publish the new clients only when the attempt produced a working one (a 4-line change):
+ if (cluster || singler)
{
std::lock_guard<std::mutex> lk(_mtx);
_cluster = cluster;
_singler = singler;
+ return true;
}
-
- if (cluster || singler) return true;
Validated at the stack head:
- the stall case delivers 3 of 3 with 0 failed writes;
- the restart case delivers 3 of 3;
- the MISCONF blackout is gone;
- the gtest suite, lifecycle, WriteErrors and WriteTransportFailure are unchanged.
Going further, and closer to the reviewers' conclusion on #91 to "use the built-in facility", reconnect() could replace the client only when none exists. That also passes the recovery scenarios, but it changes the exact probe count asserted in #109's write_failure_test.py:87.
Related: #105 ("follow the agreed redis-plus-plus reconnect policy for genuine connection loss"), #109 (defines when reconnect fires), #91 discussion.
RedisAdapter version or revision
The mechanism is on
mainatff2d5b8. Reproduced below at the #108 → #109 → #111 stack head0c3b0e6.Redis version and deployment mode
Redis 7.4.11 standalone (TCP).
Observed behavior
RedisConnection::connect()publishes what it built even when the attempt failed:{ std::lock_guard<std::mutex> lk(_mtx); _cluster = cluster; // both null after a failed attempt _singler = singler; } if (cluster || singler) return true;RedisAdapter::reconnect()runs this from its background thread whenever any call fails: a write, a get,connected(), or the watchdog. During an outage the attempt usually runs while the server is still down. It therefore replaces the existing redis-plus-plus client with null pointers, although that client would re-dial its broken pooled sockets by itself on next use.From then on every
RedisConnectioncall fails locally without touching the network or logging. That includes the reader loop, which #108 changed to retry every 50 ms on the premise that "redis++ reconnects its socket on the next read". Nothing recovers until some later call triggers a successful reconnect.In practice:
connected()recover within one poll period. Listen-only consumers do not recover at all.Expected: a failed reconnect attempt leaves the existing client in place, so redis-plus-plus's own re-dial can recover.
Reproduction
One reader on
{rst}:rwithcxn.timeout = 300. The server is shut down for about 1.5 s and started again. One write is attempted while it is down. Then 3 entries are written to the subscribed stream from another client.Control, identical except that no adapter call is made while the server is down:
In other words, the adapter's own reconnect is what breaks recovery.
The same mechanism turns a MISCONF episode into a read blackout, because Redis also refuses PING under MISCONF. One
connected()probe fails, the reconnect's own PING fails, and the clients are nulled. Every read then returnsRA_NOT_CONNECTED, although the server is still serving reads.Build and runtime environment
Suggested fix
Publish the new clients only when the attempt produced a working one (a 4-line change):
Validated at the stack head:
Going further, and closer to the reviewers' conclusion on #91 to "use the built-in facility",
reconnect()could replace the client only when none exists. That also passes the recovery scenarios, but it changes the exact probe count asserted in #109'swrite_failure_test.py:87.Related: #105 ("follow the agreed redis-plus-plus reconnect policy for genuine connection loss"), #109 (defines when reconnect fires), #91 discussion.