a month ago
Service: Redis (managed template), US East, hobby plan, 1 replica, volume "redis-volume". Currently running redis:8.2.1, which works.
There are two issues here. The second one is the serious one.
== 1. The CVE-2025-49844 patch to redis:8.2.9 fails identically every time ==
Ten consecutive failed deployments, all with the same trace:
Deployment failed during network process (04:45)
Initialization (00:00) OK
Deploy (00:07) OK
Network > Healthcheck (04:33) Healthcheck failure
Post-deploy Not startedMost recent attempt: deployment 012c981b-98d2-4605-88a2-b7cdf1367a21 at 2026-09-09 09:29 PDT, which I triggered manually with "Patch now" so I could watch it. Nine earlier automatic attempts ran across 2026-09-05 to 2026-09-07.
redis:8.2.1 has run on this service for about 5 months without trouble. Nothing in the service configuration changed between working and failing.
== 2. A failed patch leaves the database with NO running deployment, while still reporting Online ==
When the 8.2.9 deployment fails and rolls back, the previous redis:8.2.1 deployment is left in COMPLETED state rather than ACTIVE. No container is running. The service card and the dashboard still show "Online". Clients get connect ETIMEDOUT. It stays that way until someone manually hits Redeploy on the old deployment.
On 2026-09-05 this went unnoticed for four days:
2026-09-05 13:27:20 UTC Redis receives SIGTERM, clean shutdown, final RDB saved
("User requested shutdown", "Redis is now ready to exit")through 2026-09-06 10:46 PDT roughly 50 "Stopping Container" entries as 8.2.9
repeatedly fails its healthcheckafter that Redis logs are completely silent for three days,
while the dashboard continues to show "Online"Our worker service could not reach Redis that entire time. It exits non-zero when Redis is unreachable, so it burned through restartPolicyMaxRetries and Railway then stopped restarting it permanently. Reports and scheduled jobs were down for four days.
I reproduced this deliberately today with "Patch now" and it behaved exactly the same way: the failure left 8.2.1 as COMPLETED and Redis offline until I redeployed it by hand.
== Questions ==
-
Why does redis:8.2.9 fail the network healthcheck on this service when 8.2.1 passes? Is there anything on your side in the healthcheck logs for those deployments?
-
Can a failed patch be made to leave a running deployment behind? Right now a failed security patch takes the database completely offline with no automatic recovery.
-
Related: can the service status stop reporting "Online" when no container is actually running? That incorrect status is the main reason this went unnoticed for four days.
-
Will this patch retry automatically in the next maintenance window (Sat 10:00 - Sun 18:00 UTC)? After clicking "Patch now" the banner disappeared, so I no longer have the Skip this week or Reschedule options. I would like to avoid a repeat while this is being investigated. Is there a way to pause it?
2 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 26 days ago
a month ago
Dealing with the outage first, since that's the urgent half.
To get Redis running again right now: go to the service's Deployments tab,
find the last successful 8.2.1 deployment (the one that ran for five
months), and redeploy that specific deployment rather than triggering a new
one. Redeploying a known-good historical deployment restores that exact
image and config, instead of re-attempting the patch and re-entering the
failure loop. Your workers having exhausted their restart policies means
they'll need a nudge too once Redis is actually answering.
On why 8.2.9 fails where 8.2.1 passed — I can't see Railway's healthcheck
logs, so this is a hypothesis rather than an answer, but it fits your
timings unusually well and it's cheap for Railway to confirm or kill.
Your failure signature is Deploy OK at 00:07, healthcheck failure at 04:33.
That's roughly four and a half minutes of the healthcheck not getting a
satisfactory response from a Redis that, by the deploy step's account,
started fine. Redis accepts connections while it's still loading a dataset
from disk, but it answers commands with -LOADING until the load completes.
If Railway's healthcheck treats a -LOADING reply as a failure rather than
as "not ready yet, retry", then any instance whose dataset takes longer to
load than the healthcheck window will fail — deterministically, on every
attempt, exactly as yours does.
That would also explain the part that otherwise looks strange: why this
started failing after five months of working. Nothing about 8.2.9 needs to
be broken. The dataset on your volume just has to have grown past the point
where it loads inside the healthcheck budget, and the patch is simply the
first restart in months that had to load it cold.
If that's what's happening, the fix is on Railway's side — either the
healthcheck tolerating -LOADING, or a longer healthcheck timeout for
volume-backed Redis — and it would affect every managed Redis instance with
a large enough dataset, not just yours. Worth asking them directly whether
the healthcheck distinguishes -LOADING from a hard failure, and what the
timeout is. It's a yes/no question they can check quickly.
On your second issue, I'd push hard on it separately from the patch
question, because it's the more serious bug of the two. A failed patch
leaving the previous deployment in COMPLETED rather than rolling back to it
means a failed patch takes the service fully down. The dashboard continuing
to report "Online" with no running container turned what should have been a
visible failure into three days of silence — you had no signal until
downstream services had already burned through their restart budgets. Those
are arguably two separate platform bugs (no rollback-on-patch-failure, and
a status indicator that reflects intent rather than actual container state),
and they're worth filing as such rather than as footnotes to a version
question.
Practical hardening in the meantime, given the dashboard can't be trusted as
a liveness signal: monitor Redis with an external check that actually issues
a PING and alerts on failure, rather than relying on service status. That
would have turned a four-day outage into a four-minute one regardless of how
the patching question resolves.
a month ago
Two parts: what to do right now to get Redis back up, and a script that turns
my -LOADING theory from a guess into evidence.
RIGHT NOW: Deployments tab → find the 8.2.1 deployment that ran for five
months → its "..." menu → Redeploy. That restores the known-good image
directly instead of re-entering the patch/fail loop. Note this has to be
done from the dashboard — Railway's CLI redeploy only re-runs the current
deployment, it can't target a specific historical one by ID, so there's no
faster CLI path here.
WHY 8.2.9 IS PROBABLY FAILING: your timing — Deploy OK at 00:07, healthcheck
failure at 04:33 — is consistent with Redis still loading its dataset from
disk when the healthcheck gives up. Redis accepts TCP connections while
loading but answers -LOADING to commands until it's done; if Railway's
healthcheck treats that as a failure instead of "not ready yet," any restart
whose dataset takes longer than the healthcheck window to load will fail
every time, deterministically — which explains both the 8.2.9 failures and
why 5 months of 8.2.1 uptime never hit this until now (nothing had forced a
cold restart on this dataset size before).
Script to confirm it — polls INFO persistence and timestamps every state
change, so you get a real record covering the failure window instead of a
guess:
#!/usr/bin/env bash
# watch-loading.sh — run via: railway run bash watch-loading.sh
# from another service in the same project, started BEFORE you hit
# "Patch now" again. Needs REDIS_URL, or REDISHOST/REDISPORT/REDISPASSWORD
# (Railway's Redis plugin sets these).
set -euo pipefail
INTERVAL="${INTERVAL:-1}"
DURATION="${DURATION:-330}" # covers your 4m33s window with margin
if [[ -n "${REDIS_URL:-}" ]]; then
CLI=(redis-cli -u "$REDIS_URL")
elif [[ -n "${REDISHOST:-}" ]]; then
CLI=(redis-cli -h "$REDISHOST" -p "${REDISPORT:-6379}")
[[ -n "${REDISPASSWORD:-}" ]] && CLI+=(-a "$REDISPASSWORD" --no-auth-warning)
else
echo "Set REDIS_URL or REDISHOST first." >&2; exit 2
fi
echo "polling every ${INTERVAL}s for up to ${DURATION}s"
echo "time loading rdb_bgsave_status connected"
start=$(date +%s); seen_loading=0
while true; do
(( $(date +%s) - start >= DURATION )) && break
ts="$(date -u +'%Y-%m-%dT%H:%M:%SZ')"
if info="$("${CLI[@]}" info persistence 2>/dev/null)"; then
loading="$(printf '%s\n' "$info" | grep -m1 '^loading:' | cut -d: -f2 | tr -d '\r')"
bgstat="$(printf '%s\n' "$info" | grep -m1 '^rdb_last_bgsave_status:' | cut -d: -f2 | tr -d '\r')"
printf '%s %-7s %-17s yes\n' "$ts" "${loading:-?}" "${bgstat:-?}"
[[ "$loading" == "1" ]] && seen_loading=1
else
printf '%s %-7s %-17s NO\n' "$ts" "-" "-"
fi
sleep "$INTERVAL"
done
echo
if [[ "$seen_loading" == "1" ]]; then
echo "RESULT: loading:1 observed — dataset-load theory confirmed."
else
echo "RESULT: loading:1 never seen — something else is failing the healthcheck."
fiRun it, then hit "Patch now," and watch the output live. Either outcome is
useful evidence to hand Railway directly: "the healthcheck fails while
loading:1 is set" is a precise, fixable bug report, not a support thread.
Separately — and I'd file this distinctly from the version question, because
it's the more serious bug: a failed patch left the previous deployment in
COMPLETED instead of rolling back to it, so the failed patch took the whole
service down, and the dashboard kept reporting "Online" with zero running
container for three days. That's what turned a patch failure into a 4-day
outage for your dependent workers. Worth pushing on that as its own item —
no-rollback-on-failed-patch and a status indicator that reflects intent
rather than actual container state are two separate platform bugs.
Until that's addressed: monitor with something that actually issues a PING
and alerts on failure, not the dashboard status — that alone would have
turned this into a 4-minute incident instead of 4 days.