7 min read
How We Found and Fixed 20+ Silent Failures in Production¶
Date: 2026-07-08
Tags: rust, reliability, silent-failures, redis, websocket, scheduler, observability
The Problem¶
Our card game backend had been running in production for months. Players would occasionally report that game state didn't update, or that a bot stopped responding mid-game. The logs showed nothing alarming. No errors. No panics. Everything looked fine.
But it wasn't. We had silent failures — errors that occurred but were never reported, logged, or surfaced to anyone. They manifested as subtle degradations: a missing WebSocket event here, a stalled bot there, a cache that never invalidated.
We conducted a systematic audit of the entire backend and found over 20 distinct silent failure points. Here's what we found and how we fixed them.
What We Found¶
1. Fire-and-Forget Redis Publishing (15+ sites)¶
Every game event — card plays, round completions, game finished — was broadcast to clients via Redis Pub/Sub. The publish calls looked like this:
If Redis was briefly unavailable (network blip, reconnection, memory pressure), the event was permanently lost. No retry. No buffer. No alert. Clients simply never received the update.
2. WebSocket Registration Before Confirmation¶
When a client connected via WebSocket, we registered them in the connection manager before sending the welcome message. If the welcome failed (serialization error, channel full), the connection was orphaned — registered but never confirmed, consuming resources forever.
3. Double-Remove Race Condition¶
Two separate tokio tasks managed each WebSocket connection: a forwarding task (sending events to the client) and a stream handler (receiving messages from the client). When a client disconnected, both tasks detected it and both called remove_connection. Metrics were double-counted, and PlayerDisconnected events were sometimes skipped entirely.
4. Scheduler Worker Killed by Any Task Exit¶
The scheduler worker ran 6 background tasks (game cancellation, stall detection, freeze checking, leaderboard refresh, DB metrics, run checking). The main loop used JoinSet::join_next() — if any task exited, the entire process called std::process::exit(1), killing all other tasks.
5. Bot Chain Broke on Transient DB Errors¶
The synchronous bot chain executed bot moves sequentially. A single transient database error (connection timeout, deadlock) caused a break, stalling the entire game with no recovery.
6. Redis Subscriber Shards Died Silently¶
Redis event forwarding used N sharded subscriber tasks. If a shard failed to subscribe (Redis temporarily down), it simply returned — dying forever with no retry. Events for games hashed to that shard were permanently lost.
7-12. Cache Writes, Token Validation, Snapshots, Emails...¶
Cache write errors were silently ignored (let _ = redis.set_ex(...)). Game token validation fell back to "allow" on Redis errors (security regression). Auth token blacklist checks used unwrap_or(false), which silently allowed blacklisted tokens when Redis was down. Game state snapshots failed silently, leaving players with blank screens. Run completion failures left runs stuck in "active" forever. Email send failures for unfreeze and kick notifications were silently swallowed.
What We Built¶
Phase 1: Retry + Buffer for Redis Events¶
We added a publish_with_retry method with exponential backoff (3 retries, 100ms base, 1s max) and a PublishResult enum so callers can decide fallback behavior:
pub enum PublishResult {
Published,
RetryExhausted(String),
}
pub async fn publish_with_retry(&mut self, channel: &str, msg: &str) -> PublishResult {
for attempt in 0..self.max_retries {
match self.publish(channel, msg).await {
Ok(()) => return PublishResult::Published,
Err(e) if attempt + 1 < self.max_retries => {
REDIS_PUBLISH_RETRIES_TOTAL.with_label_values(&[channel_prefix]).inc();
tokio::time::sleep(backoff(attempt)).await;
}
Err(e) => {
REDIS_PUBLISH_FAILURES_TOTAL.with_label_values(&[channel_prefix]).inc();
self.enqueue_event(channel, msg);
return PublishResult::RetryExhausted(e.to_string());
}
}
}
// unreachable, but the compiler doesn't know
}
On top of retries, we added an in-memory event buffer (bounded VecDeque, max 1000 events). If retries are exhausted, events are queued in memory. A background flush task retries buffered events every 5 seconds. If the buffer overflows, we log a critical error and increment a metric.
Phase 2: WebSocket Lifecycle Fixes¶
Orphan fix: Connection registration (add_connection) now happens after the welcome message succeeds. If welcome fails, we return the upgrade response without registering — the connection drops naturally.
Race fix: Added a disconnected: bool flag to TrackedConnection and a force_disconnect method. The forwarding task calls force_disconnect, which always publishes PlayerDisconnected if the player identity was set. The stream handler's remove_connection checks the flag first to avoid double actions.
Heartbeat monitoring: Every connection tracks last_pong. The cleanup task checks both idle time and heartbeat timeout (default 90s). Timed-out connections log a warning, publish PlayerDisconnected, and increment ws_heartbeat_timeouts_total.
Phase 3: Self-Healing Scheduler¶
Each task in the scheduler now runs inside a supervision loop:
tasks.spawn(async move {
let mut restart_count = 0u32;
let mut window_start = Instant::now();
loop {
run_task().await;
warn!("Task exited, restarting");
restart_count += 1;
if window_start.elapsed() > Duration::from_secs(300) {
restart_count = 0;
window_start = Instant::now();
}
if restart_count > task_max_restarts {
error!("Too many restarts, sleeping 60s");
sleep(Duration::from_secs(60)).await;
restart_count = 0;
window_start = Instant::now();
}
sleep(Duration::from_secs(1)).await;
}
});
The std::process::exit(1) in scheduler_worker.rs was removed. Tasks restart themselves; only Ctrl+C triggers shutdown.
Phase 4: Bot Chain Retry¶
Transient DB errors in the bot chain now retry up to 3 times with 500ms backoff:
let mut retries = 0u32;
loop {
match execute_one_bot_move(...).await {
Ok(o) => break o,
Err(e) if retries < 3 => {
retries += 1;
BOT_CHAIN_RETRIES_TOTAL.inc();
sleep(Duration::from_millis(500 * retries as u64)).await;
}
Err(e) => {
BOT_CHAIN_BREAKS_TOTAL.inc();
error!("Bot failed after {} retries: {}", retries, e);
return;
}
}
}
Phase 5: Redis Subscriber Reconnect¶
Each shard now runs inside an infinite retry loop with exponential backoff:
let mut attempt: u64 = 0;
loop {
match redis.psubscribe(&["game:*", "room:*"]).await {
Ok(ps) => {
REDIS_SUBSCRIBER_SHARDS_ACTIVE.inc();
attempt = 0;
process_messages(ps).await;
REDIS_SUBSCRIBER_SHARDS_ACTIVE.dec();
}
Err(e) => {
attempt += 1;
let delay = (1u64 << attempt.min(5)).min(30);
error!("Shard subscribe failed (attempt {}), retrying in {}s", attempt, delay);
sleep(Duration::from_secs(delay)).await;
}
}
}
Phase 6: Fail-Closed Security + Observability¶
- Game token validation: On Redis error, deny the connection (was: allow). No config flag — unconditional fail-closed.
- Auth blacklist: Changed
unwrap_or(false)tounwrap_or(true)on Redis error. Assume token is blacklisted when we can't check. - Cache writes: All
let _ = redis.set_ex(...)calls now log atwarnlevel and incrementGAME_STATE_CACHE_WRITE_ERRORS_TOTAL. - Email failures: User/profile lookup failures log at
warnwithEMAIL_SEND_ERRORS_TOTALmetric (labeled by type:unfreeze,stall_warning,stall_kicked). - Snapshot failures: Distinguish
NotFoundfromDbErr, send error messages to WebSocket clients. - Run completion: Added 3-attempt retry for run lookup with 1s delay.
The Metrics Dashboard¶
We added 12 new Prometheus metrics to track system health:
| Metric | What It Tells You |
|---|---|
redis_publish_retries_total |
Redis is having transient issues |
redis_publish_failures_total |
Redis is down, events are buffering |
redis_buffer_overflow_total |
Event buffer is full, events are being dropped |
redis_subscriber_shards_active |
How many subscriber shards are connected |
ws_heartbeat_timeouts_total |
Clients are disconnecting without clean close |
bot_chain_retries_total |
Transient DB errors in bot execution |
bot_chain_breaks_total |
Bot chain is failing after all retries |
run_completion_errors_total |
Run finalization is failing |
email_send_errors_total |
Email delivery issues by type |
game_state_cache_write_errors_total |
Redis cache writes are failing |
ws_token_validation_redis_errors_total |
Redis errors during token validation |
ws_auth_blacklist_redis_errors_total |
Redis errors during blacklist check |
Set up alerts on any non-zero values for redis_buffer_overflow_total, bot_chain_breaks_total, redis_subscriber_shards_active < expected_shards, and sustained growth on any error counter.
Configuration: Every Threshold Is Tuneable¶
All retry counts, delays, buffer sizes, and timeouts are environment variables with sensible defaults:
REDIS_PUBLISH_MAX_RETRIES=3
REDIS_PUBLISH_RETRY_BASE_DELAY_MS=100
REDIS_PUBLISH_RETRY_MAX_DELAY_MS=1000
REDIS_EVENT_BUFFER_MAX_SIZE=1000
REDIS_EVENT_BUFFER_FLUSH_INTERVAL_SECS=5
WS_HEARTBEAT_INTERVAL_SECS=30
WS_HEARTBEAT_TIMEOUT_SECS=90
BOT_CHAIN_MAX_RETRIES=3
BOT_CHAIN_RETRY_BASE_DELAY_MS=500
SCHEDULER_TASK_MAX_RESTARTS=3
RUN_COMPLETION_MAX_RETRIES=3
REDIS_SUBSCRIBER_RETRY_MAX_DELAY_SECS=30
Key Takeaways¶
-
let _ =is the most dangerous pattern in Rust. Everylet _ =discarding aResultis a potential silent failure. Replace them with proper error handling, logging, and metrics. -
Fire-and-forget needs retry infrastructure. If an operation matters (and event publishing matters), it needs retry logic with backoff, a fallback path, and observability.
-
Failures should fail closed, not open. A Redis error during token validation should deny the connection, not allow it. Security degrades under failure; design against this.
-
Every background task needs a supervisor. If a scheduler task exits, restart it. If it keeps exiting, pause and alert. Never kill the entire process because one task finished.
-
Metrics turn silent failures into loud ones. Adding counters for every error path means you can set up alerts and catch problems before users do.