postgres-world frozen, volume not released
aatreviso
PROOP

2 months ago

Subject: postgres-world — frozen container holding volume after hardware failure; production DB down

Project: talented-wonder (9dddb7a2-1ddb-4adc-a780-87c273d04318)

Service: postgres-world (69d54d56-96be-4672-b491-d486f10063a3)

Volume: postgres-volume-y7Bo

Region: US West (us-west2)

Related maintenance notice: e76ea1c7-ea50-4e45-8034-b9f5c3583641

                        (hardware failure, 2026-08-20 22:53:58 UTC)

IMPACT

Production game (world.conquizzone.it) fully down for over an hour. The

app service "world" is Online and serving HTTP, but cannot open a single

connection to the database, so every request fails — game, admin panel

and health endpoint alike. All other services in the project are fine.

TIMELINE (UTC)

  • 22:33-22:34, 2026-08-20 — the postgres-world container STOPPED

    CONSUMING CPU but did not exit. Metrics show the us-west2-replica-0

    series ending there, leaving a flat 0.00 vCPU line and frozen memory.

    The dashboard kept reporting the service as "Online".

    Note this is ~20 minutes BEFORE your hardware-failure notice at

    22:53:58 — please check whether detection lagged, or whether these

    are two separate events.

  • From 22:34 onward, the app service logs a continuous stream of

    "Connection terminated due to connection timeout" (node-postgres,

    connectionTimeoutMillis = 5000), including during its own startup

    migrations.

  • ~23:00 — redeployed postgres-world. The deployment stayed in

    "Creating containers" for 8+ minutes and produced no logs at all.

  • ~23:15 — removed the pending deployment; Railway reverted to the

    previous (one-month-old) deployment, now shown ACTIVE / "Deployment

    successful", but CPU is still flat at 0.00 — nothing is running.

    Subsequent deployments also sit in "Creating containers", no logs.

  • Your own tooling now returns "Service not found" when listing

    deployments for this service.

POSTGRES DID NOT SHUT DOWN

We searched the database log for FATAL, PANIC, "out of memory" and

"shutdown": no matches on 2026-08-20 or 2026-08-21. The only FATAL

entries in the whole log are old and unrelated (2026-07-26, 2026-08-11,

2026-08-19).

Postgres never reported a reason. It did not shut down, crash, or hit a

resource limit — it stopped producing log lines in the middle of normal

activity and never exited.

RULED OUT

  • Not a deploy: this is a database service, last deployed ~1 month ago.

  • Not memory: ~900 MB at the time; the same instance ran at 1.5 GB

    earlier the same day, and there is no OOM message.

  • Not disk: the checkpoint completed at 22:33:46 UTC with sync=0.046 s

    over 174 files (1058 buffers) — I/O was healthy seconds before the

    freeze, so the volume was not stalled.

  • Not our connection pool: the app fails even for its own boot

    migrations, and restarting the app service does not help.

WHAT WE NEED

Force-release volume postgres-volume-y7Bo and clear the frozen

container so a new instance can be scheduled. Given the hardware

failure and earlier instability on this instance (two

"FATAL: canceling authentication due to timeout" on 2026-08-11 at

20:34 UTC), please migrate the service to a different host rather than

only unblocking it.

Data integrity is the priority: the database was not shut down cleanly,

so we expect WAL recovery on the next successful start.

Solved

10 Replies

Status changed to Awaiting Railway Response Railway • about 2 months ago


aatreviso
PROOP

2 months ago

Adding two things:

We have ALREADY tried redeploying postgres-world three times. Every

attempt sits in "Creating containers" and produces ZERO log lines — not

one, not even a startup message. Removing the stuck deployment just

reverts to the previous one, which is not actually running (CPU flat at

0.00 vCPU). Restarting the dependent app service has no effect either.

A redeploy cannot succeed while the host holding postgres-volume-y7Bo is

unreachable, so please don't ask us to retry one until the host is

confirmed healthy. What we need is someone with host access to

force-release the volume and clear the frozen container.


aatreviso

Adding two things: We have ALREADY tried redeploying postgres-world three times. Every attempt sits in "Creating containers" and produces ZERO log lines — not one, not even a startup message. Removing the stuck deployment just reverts to the previous one, which is not actually running (CPU flat at 0.00 vCPU). Restarting the dependent app service has no effect either. A redeploy cannot succeed while the host holding postgres-volume-y7Bo is unreachable, so please don't ask us to retry one until the host is confirmed healthy. What we need is someone with host access to force-release the volume and clear the frozen container.

ferock
HOBBY

2 months ago

i have the same issue with my postgres


aatreviso

Adding two things: We have ALREADY tried redeploying postgres-world three times. Every attempt sits in "Creating containers" and produces ZERO log lines — not one, not even a startup message. Removing the stuck deployment just reverts to the previous one, which is not actually running (CPU flat at 0.00 vCPU). Restarting the dependent app service has no effect either. A redeploy cannot succeed while the host holding postgres-volume-y7Bo is unreachable, so please don't ask us to retry one until the host is confirmed healthy. What we need is someone with host access to force-release the volume and clear the frozen container.

dizzydes90
EMPLOYEE

2 months ago

Unfortunately we're having an incident with the host you're project is currently on and fixing it is going to take about 1-2 hours.

Please accept my apologies we have several platform engineers investigating this right now. Will update here as I know more.


Status changed to Awaiting User Response Railway • about 2 months ago


aatreviso
PROOP

2 months ago

Ok thanks


Status changed to Awaiting Railway Response Railway • about 2 months ago


rdcavalheiro
HOBBY

2 months ago

Same incident here — different project, same host, same signature.

Project amusing-forgiveness — 6a7e3b0d-8324-4bf7-81cd-79893db68cc0,

environment production, service Postgres

(ghcr.io/railwayapp-templates/postgres-ssl:18, US West, 1 replica).

Timeline (UTC)

  • 22:02:10 — last checkpoint in the Postgres logs, clean: 22 buffers written,

    17 files synced, LSN advancing normally. No errors, no shutdown, checkpointer

    still the same PID it had for days. Nothing logged after this, despite

    clients retrying continuously for the next 90+ minutes.

  • ~23:04 — app service (same project) starts failing every query with

    ETIMEDOUT from pg-pool.

  • 23:28 — redeploy c6dab28d-bfd7-4402-bbc2-ccd63ffece8a hangs in

    Deploy > Creating containers for 15+ min, producing no logs.

    1f3ceb0a-8406-472a-a626-947ca5de86f1 queued behind it and never started.

    Both removed since; not retrying. The last good deployment is

    a0ba3bb0-5180-4e83-be85-de311bb4e5c0.

What is unreachable, and what is not

  • Private network from a sibling container: postgres.railway.internal resolves

    to both an IPv6 ULA and a 10.x address; TCP times out on both families.

  • Public TCP proxy: accepts the TCP connection, then ECONNRESET on the

    Postgres startup handshake.

  • railway ssh --service Postgres → "An error occurred connecting to your

    service". Volume file browser (SFTP) → timeout.

  • Meanwhile **another service in the same project, on the same private network,

    is fully reachable and healthy** — so this is not project-wide networking, it

    is scoped to the service whose volume lives on the failed host.

  • Volume postgres-volume reports 0MB/5000MB used, Status: Ready, which

    contradicts the successful checkpoints written minutes before the freeze.

Two things worth flagging

  1. I received the "Server Outage Affecting Your Service — hardware failure, no

    action is needed on your part, your service will automatically resume"

    notification, and it was marked Resolved — but the service had not

    resumed an hour later, and the staff comment in this thread (posted after

    that notification) says the fix still needs 1–2h. Should we treat the

    Resolved flag as unreliable during host incidents? It sent me chasing

    redeploys that could never work while the volume is locked to the dead host.

  2. No backups available on Hobby (backups/PITR are Pro-only), and there is

    no way to take one now since nothing can connect. My database also has a

    scheduled CVE patch (CVE-2026-15741) whose description says "a snapshot is

    taken first" — **could you trigger that snapshot before any container

    recreate?** That would be the only copy of this data in existence.

Main question: **is the volume data intact, and will it be released to a healthy

host?**


seansanjose
PRO

2 months ago

Same confirmed host incident and signature on our production project.

Project: athletic-perfection (f3edb964-5792-4033-bc20-0788410b7805)

Environment: production (e63181af-c2d4-4a37-995d-bd07c2ed1d37)

Service: Postgres (1ebf1d7d-8df3-43b7-9c15-5ffee34e34a8)

Volume: postgres-volume (f43fdd26-0aa2-45b9-babb-489c50a0af77)

Region: US West (us-west2)

Image: ghcr.io/railwayapp-templates/postgres-ssl:18

Timeline/signature:

  • Last normal Postgres log activity ended around 22:34 UTC with no FATAL, PANIC, OOM, or shutdown.
  • Dashboard continued to report Online, but postgres.railway.internal resolves and TCP 5432 times out from a healthy sibling production service.
  • railway ssh to Postgres fails; Railway Database UI remains at "Attempting to connect"; service metrics UI returns 500.
  • Dependent app deployments fail health with "Connection terminated due to connection timeout".
  • Two approved same-image/same-volume recovery attempts never started and produced zero logs:
    • 23706617-7dfe-4085-a025-e4d60b5d4fb5
    • 793ffb0d-3b03-4ad0-a2c8-b72516ec24fd
  • Persistent volume has not been deleted, replaced, detached, or modified.

Production is fully down. Please force-release the volume/frozen workload and place it on a healthy host while preserving volume data. Please confirm data integrity and host recovery before we retry the application deployment.


dizzydes90
EMPLOYEE

2 months ago

We're going through the steps for a full reboot to get everyone back up. Will continue to update.


Status changed to Awaiting User Response Railway • about 2 months ago


seansanjose
PRO

2 months ago

Fresh production update for athletic-perfection:

  • The existing postgres-volume was live-resized from 15.36 GB to 30 GB in Railway; it was not detached, deleted, replaced, or recreated.
  • Postgres still has 0/1 running replicas.
  • postgres.railway.internal now returns ENOTFOUND from a healthy sibling service.
  • A separate database in the same project, 121-platform-postgres.railway.internal:5432, resolves and accepts TCP successfully, ruling out project-wide private networking.
  • DATABASE_URL/PGHOST and the Postgres RAILWAY_PRIVATE_DOMAIN all correctly reference postgres.railway.internal:5432.
  • railway ssh to Postgres still fails at the service layer.
  • Postgres logs still end abruptly around 22:34 UTC after successful checkpoints, with no PANIC, OOM, shutdown, or no-space error.
  • 121enterprise starts but reports getaddrinfo ENOTFOUND postgres.railway.internal; it has 0/1 healthy replicas and public IDX routes remain 502.

The announced 1-2 hour recovery window has passed and the full-reboot update is now over an hour old. Please confirm whether host reboot/volume release completed for volume f43fdd26-0aa2-45b9-babb-489c50a0af77. If not, please force-release/migrate this existing volume to a healthy host while preserving its data, and tell us when one controlled Postgres redeploy is safe.


Status changed to Awaiting Railway Response Railway • about 2 months ago


Will do, as an update to everyone on the thread, we are working through the migration steps to get everyone on a healthy host.


Status changed to Awaiting User Response Railway • about 2 months ago


Update to all affected, we have performed recovery for all workloads on the machine. If you are still observing impact, reply and our on-call team will engage.

We apologize for the impact. We advise all customers to implement HA configs when possible for mission critical workloads.


Railway
BOT

a month ago

This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!

Status changed to Solved Railway • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...