Managed Redis CVE-2025-49844 auto-patch to 8.2.9 fails healthcheck and leaves the service with no running deployment
yussefalta
HOBBYOP

a month ago

Service: Redis (managed template), US East, hobby plan, 1 replica, volume "redis-volume". Currently running redis:8.2.1, which works.

There are two issues here. The second one is the serious one.

== 1. The CVE-2025-49844 patch to redis:8.2.9 fails identically every time ==

Ten consecutive failed deployments, all with the same trace:

Deployment failed during network process (04:45)

Initialization         (00:00)  OK

Deploy                 (00:07)  OK

Network > Healthcheck  (04:33)  Healthcheck failure

Post-deploy                     Not started

Most recent attempt: deployment 012c981b-98d2-4605-88a2-b7cdf1367a21 at 2026-09-09 09:29 PDT, which I triggered manually with "Patch now" so I could watch it. Nine earlier automatic attempts ran across 2026-09-05 to 2026-09-07.

redis:8.2.1 has run on this service for about 5 months without trouble. Nothing in the service configuration changed between working and failing.

== 2. A failed patch leaves the database with NO running deployment, while still reporting Online ==

When the 8.2.9 deployment fails and rolls back, the previous redis:8.2.1 deployment is left in COMPLETED state rather than ACTIVE. No container is running. The service card and the dashboard still show "Online". Clients get connect ETIMEDOUT. It stays that way until someone manually hits Redeploy on the old deployment.

On 2026-09-05 this went unnoticed for four days:

2026-09-05 13:27:20 UTC Redis receives SIGTERM, clean shutdown, final RDB saved

                       ("User requested shutdown", "Redis is now ready to exit")

through 2026-09-06 10:46 PDT roughly 50 "Stopping Container" entries as 8.2.9

                       repeatedly fails its healthcheck

after that Redis logs are completely silent for three days,

                       while the dashboard continues to show "Online"

Our worker service could not reach Redis that entire time. It exits non-zero when Redis is unreachable, so it burned through restartPolicyMaxRetries and Railway then stopped restarting it permanently. Reports and scheduled jobs were down for four days.

I reproduced this deliberately today with "Patch now" and it behaved exactly the same way: the failure left 8.2.1 as COMPLETED and Redis offline until I redeployed it by hand.

== Questions ==

  1. Why does redis:8.2.9 fail the network healthcheck on this service when 8.2.1 passes? Is there anything on your side in the healthcheck logs for those deployments?

  2. Can a failed patch be made to leave a running deployment behind? Right now a failed security patch takes the database completely offline with no automatic recovery.

  3. Related: can the service status stop reporting "Online" when no container is actually running? That incorrect status is the main reason this went unnoticed for four days.

  4. Will this patch retry automatically in the next maintenance window (Sat 10:00 - Sun 18:00 UTC)? After clicking "Patch now" the banner disappeared, so I no longer have the Skip this week or Reschedule options. I would like to avoid a repeat while this is being investigated. Is there a way to pause it?

$10 Bounty

2 Replies

Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • 26 days ago


Dealing with the outage first, since that's the urgent half.

To get Redis running again right now: go to the service's Deployments tab,

find the last successful 8.2.1 deployment (the one that ran for five

months), and redeploy that specific deployment rather than triggering a new

one. Redeploying a known-good historical deployment restores that exact

image and config, instead of re-attempting the patch and re-entering the

failure loop. Your workers having exhausted their restart policies means

they'll need a nudge too once Redis is actually answering.

On why 8.2.9 fails where 8.2.1 passed — I can't see Railway's healthcheck

logs, so this is a hypothesis rather than an answer, but it fits your

timings unusually well and it's cheap for Railway to confirm or kill.

Your failure signature is Deploy OK at 00:07, healthcheck failure at 04:33.

That's roughly four and a half minutes of the healthcheck not getting a

satisfactory response from a Redis that, by the deploy step's account,

started fine. Redis accepts connections while it's still loading a dataset

from disk, but it answers commands with -LOADING until the load completes.

If Railway's healthcheck treats a -LOADING reply as a failure rather than

as "not ready yet, retry", then any instance whose dataset takes longer to

load than the healthcheck window will fail — deterministically, on every

attempt, exactly as yours does.

That would also explain the part that otherwise looks strange: why this

started failing after five months of working. Nothing about 8.2.9 needs to

be broken. The dataset on your volume just has to have grown past the point

where it loads inside the healthcheck budget, and the patch is simply the

first restart in months that had to load it cold.

If that's what's happening, the fix is on Railway's side — either the

healthcheck tolerating -LOADING, or a longer healthcheck timeout for

volume-backed Redis — and it would affect every managed Redis instance with

a large enough dataset, not just yours. Worth asking them directly whether

the healthcheck distinguishes -LOADING from a hard failure, and what the

timeout is. It's a yes/no question they can check quickly.

On your second issue, I'd push hard on it separately from the patch

question, because it's the more serious bug of the two. A failed patch

leaving the previous deployment in COMPLETED rather than rolling back to it

means a failed patch takes the service fully down. The dashboard continuing

to report "Online" with no running container turned what should have been a

visible failure into three days of silence — you had no signal until

downstream services had already burned through their restart budgets. Those

are arguably two separate platform bugs (no rollback-on-patch-failure, and

a status indicator that reflects intent rather than actual container state),

and they're worth filing as such rather than as footnotes to a version

question.

Practical hardening in the meantime, given the dashboard can't be trusted as

a liveness signal: monitor Redis with an external check that actually issues

a PING and alerts on failure, rather than relying on service status. That

would have turned a four-day outage into a four-minute one regardless of how

the patching question resolves.


Two parts: what to do right now to get Redis back up, and a script that turns

my -LOADING theory from a guess into evidence.

RIGHT NOW: Deployments tab → find the 8.2.1 deployment that ran for five

months → its "..." menu → Redeploy. That restores the known-good image

directly instead of re-entering the patch/fail loop. Note this has to be

done from the dashboard — Railway's CLI redeploy only re-runs the current

deployment, it can't target a specific historical one by ID, so there's no

faster CLI path here.

WHY 8.2.9 IS PROBABLY FAILING: your timing — Deploy OK at 00:07, healthcheck

failure at 04:33 — is consistent with Redis still loading its dataset from

disk when the healthcheck gives up. Redis accepts TCP connections while

loading but answers -LOADING to commands until it's done; if Railway's

healthcheck treats that as a failure instead of "not ready yet," any restart

whose dataset takes longer than the healthcheck window to load will fail

every time, deterministically — which explains both the 8.2.9 failures and

why 5 months of 8.2.1 uptime never hit this until now (nothing had forced a

cold restart on this dataset size before).

Script to confirm it — polls INFO persistence and timestamps every state

change, so you get a real record covering the failure window instead of a

guess:

#!/usr/bin/env bash

# watch-loading.sh — run via: railway run bash watch-loading.sh

# from another service in the same project, started BEFORE you hit

# "Patch now" again. Needs REDIS_URL, or REDISHOST/REDISPORT/REDISPASSWORD

# (Railway's Redis plugin sets these).

set -euo pipefail

INTERVAL="${INTERVAL:-1}"

DURATION="${DURATION:-330}"   # covers your 4m33s window with margin

if [[ -n "${REDIS_URL:-}" ]]; then

  CLI=(redis-cli -u "$REDIS_URL")

elif [[ -n "${REDISHOST:-}" ]]; then

  CLI=(redis-cli -h "$REDISHOST" -p "${REDISPORT:-6379}")

  [[ -n "${REDISPASSWORD:-}" ]] && CLI+=(-a "$REDISPASSWORD" --no-auth-warning)

else

  echo "Set REDIS_URL or REDISHOST first." >&2; exit 2

fi

echo "polling every ${INTERVAL}s for up to ${DURATION}s"

echo "time                 loading  rdb_bgsave_status  connected"

start=$(date +%s); seen_loading=0

while true; do

  (( $(date +%s) - start >= DURATION )) && break

  ts="$(date -u +'%Y-%m-%dT%H:%M:%SZ')"

  if info="$("${CLI[@]}" info persistence 2>/dev/null)"; then

    loading="$(printf '%s\n' "$info" | grep -m1 '^loading:' | cut -d: -f2 | tr -d '\r')"

    bgstat="$(printf '%s\n' "$info" | grep -m1 '^rdb_last_bgsave_status:' | cut -d: -f2 | tr -d '\r')"

    printf '%s  %-7s  %-17s  yes\n' "$ts" "${loading:-?}" "${bgstat:-?}"

    [[ "$loading" == "1" ]] && seen_loading=1

  else

    printf '%s  %-7s  %-17s  NO\n' "$ts" "-" "-"

  fi

  sleep "$INTERVAL"

done

echo

if [[ "$seen_loading" == "1" ]]; then

  echo "RESULT: loading:1 observed — dataset-load theory confirmed."

else

  echo "RESULT: loading:1 never seen — something else is failing the healthcheck."

fi

Run it, then hit "Patch now," and watch the output live. Either outcome is

useful evidence to hand Railway directly: "the healthcheck fails while

loading:1 is set" is a precise, fixable bug report, not a support thread.

Separately — and I'd file this distinctly from the version question, because

it's the more serious bug: a failed patch left the previous deployment in

COMPLETED instead of rolling back to it, so the failed patch took the whole

service down, and the dashboard kept reporting "Online" with zero running

container for three days. That's what turned a patch failure into a 4-day

outage for your dependent workers. Worth pushing on that as its own item —

no-rollback-on-failed-patch and a status indicator that reflects intent

rather than actual container state are two separate platform bugs.

Until that's addressed: monitor with something that actually issues a PING

and alerts on failure, not the dashboard status — that alone would have

turned this into a 4-minute incident instead of 4 days.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...