2 months ago
Subject: postgres-world — frozen container holding volume after hardware failure; production DB down
Project: talented-wonder (9dddb7a2-1ddb-4adc-a780-87c273d04318)
Service: postgres-world (69d54d56-96be-4672-b491-d486f10063a3)
Volume: postgres-volume-y7Bo
Region: US West (us-west2)
Related maintenance notice: e76ea1c7-ea50-4e45-8034-b9f5c3583641
(hardware failure, 2026-08-20 22:53:58 UTC)IMPACT
Production game (world.conquizzone.it) fully down for over an hour. The
app service "world" is Online and serving HTTP, but cannot open a single
connection to the database, so every request fails — game, admin panel
and health endpoint alike. All other services in the project are fine.
TIMELINE (UTC)
-
22:33-22:34, 2026-08-20 — the postgres-world container STOPPED
CONSUMING CPU but did not exit. Metrics show the us-west2-replica-0
series ending there, leaving a flat 0.00 vCPU line and frozen memory.
The dashboard kept reporting the service as "Online".
Note this is ~20 minutes BEFORE your hardware-failure notice at
22:53:58 — please check whether detection lagged, or whether these
are two separate events.
-
From 22:34 onward, the app service logs a continuous stream of
"Connection terminated due to connection timeout" (node-postgres,
connectionTimeoutMillis = 5000), including during its own startup
migrations.
-
~23:00 — redeployed postgres-world. The deployment stayed in
"Creating containers" for 8+ minutes and produced no logs at all.
-
~23:15 — removed the pending deployment; Railway reverted to the
previous (one-month-old) deployment, now shown ACTIVE / "Deployment
successful", but CPU is still flat at 0.00 — nothing is running.
Subsequent deployments also sit in "Creating containers", no logs.
-
Your own tooling now returns "Service not found" when listing
deployments for this service.
POSTGRES DID NOT SHUT DOWN
We searched the database log for FATAL, PANIC, "out of memory" and
"shutdown": no matches on 2026-08-20 or 2026-08-21. The only FATAL
entries in the whole log are old and unrelated (2026-07-26, 2026-08-11,
2026-08-19).
Postgres never reported a reason. It did not shut down, crash, or hit a
resource limit — it stopped producing log lines in the middle of normal
activity and never exited.
RULED OUT
-
Not a deploy: this is a database service, last deployed ~1 month ago.
-
Not memory: ~900 MB at the time; the same instance ran at 1.5 GB
earlier the same day, and there is no OOM message.
-
Not disk: the checkpoint completed at 22:33:46 UTC with sync=0.046 s
over 174 files (1058 buffers) — I/O was healthy seconds before the
freeze, so the volume was not stalled.
-
Not our connection pool: the app fails even for its own boot
migrations, and restarting the app service does not help.
WHAT WE NEED
Force-release volume postgres-volume-y7Bo and clear the frozen
container so a new instance can be scheduled. Given the hardware
failure and earlier instability on this instance (two
"FATAL: canceling authentication due to timeout" on 2026-08-11 at
20:34 UTC), please migrate the service to a different host rather than
only unblocking it.
Data integrity is the priority: the database was not shut down cleanly,
so we expect WAL recovery on the next successful start.
10 Replies
Status changed to Awaiting Railway Response Railway • about 2 months ago
2 months ago
Adding two things:
We have ALREADY tried redeploying postgres-world three times. Every
attempt sits in "Creating containers" and produces ZERO log lines — not
one, not even a startup message. Removing the stuck deployment just
reverts to the previous one, which is not actually running (CPU flat at
0.00 vCPU). Restarting the dependent app service has no effect either.
A redeploy cannot succeed while the host holding postgres-volume-y7Bo is
unreachable, so please don't ask us to retry one until the host is
confirmed healthy. What we need is someone with host access to
force-release the volume and clear the frozen container.
aatreviso
Adding two things: We have ALREADY tried redeploying postgres-world three times. Every attempt sits in "Creating containers" and produces ZERO log lines — not one, not even a startup message. Removing the stuck deployment just reverts to the previous one, which is not actually running (CPU flat at 0.00 vCPU). Restarting the dependent app service has no effect either. A redeploy cannot succeed while the host holding postgres-volume-y7Bo is unreachable, so please don't ask us to retry one until the host is confirmed healthy. What we need is someone with host access to force-release the volume and clear the frozen container.
2 months ago
i have the same issue with my postgres
aatreviso
Adding two things: We have ALREADY tried redeploying postgres-world three times. Every attempt sits in "Creating containers" and produces ZERO log lines — not one, not even a startup message. Removing the stuck deployment just reverts to the previous one, which is not actually running (CPU flat at 0.00 vCPU). Restarting the dependent app service has no effect either. A redeploy cannot succeed while the host holding postgres-volume-y7Bo is unreachable, so please don't ask us to retry one until the host is confirmed healthy. What we need is someone with host access to force-release the volume and clear the frozen container.
2 months ago
Unfortunately we're having an incident with the host you're project is currently on and fixing it is going to take about 1-2 hours.
Please accept my apologies we have several platform engineers investigating this right now. Will update here as I know more.
Status changed to Awaiting User Response Railway • about 2 months ago
2 months ago
Ok thanks
Status changed to Awaiting Railway Response Railway • about 2 months ago
2 months ago
Same incident here — different project, same host, same signature.
Project amusing-forgiveness — 6a7e3b0d-8324-4bf7-81cd-79893db68cc0,
environment production, service Postgres
(ghcr.io/railwayapp-templates/postgres-ssl:18, US West, 1 replica).
Timeline (UTC)
-
22:02:10— last checkpoint in the Postgres logs, clean: 22 buffers written,17 files synced, LSN advancing normally. No errors, no shutdown, checkpointer
still the same PID it had for days. Nothing logged after this, despite
clients retrying continuously for the next 90+ minutes.
-
~23:04— app service (same project) starts failing every query withETIMEDOUTfrompg-pool. -
23:28— redeployc6dab28d-bfd7-4402-bbc2-ccd63ffece8ahangs inDeploy > Creating containersfor 15+ min, producing no logs.1f3ceb0a-8406-472a-a626-947ca5de86f1queued behind it and never started.Both removed since; not retrying. The last good deployment is
a0ba3bb0-5180-4e83-be85-de311bb4e5c0.
What is unreachable, and what is not
-
Private network from a sibling container:
postgres.railway.internalresolvesto both an IPv6 ULA and a
10.xaddress; TCP times out on both families. -
Public TCP proxy: accepts the TCP connection, then ECONNRESET on the
Postgres startup handshake.
-
railway ssh --service Postgres→ "An error occurred connecting to yourservice". Volume file browser (SFTP) → timeout.
-
Meanwhile **another service in the same project, on the same private network,
is fully reachable and healthy** — so this is not project-wide networking, it
is scoped to the service whose volume lives on the failed host.
-
Volume
postgres-volumereports0MB/5000MB used, Status: Ready, whichcontradicts the successful checkpoints written minutes before the freeze.
Two things worth flagging
-
I received the "Server Outage Affecting Your Service — hardware failure, no
action is needed on your part, your service will automatically resume"
notification, and it was marked Resolved — but the service had not
resumed an hour later, and the staff comment in this thread (posted after
that notification) says the fix still needs 1–2h. Should we treat the
Resolved flag as unreliable during host incidents? It sent me chasing
redeploys that could never work while the volume is locked to the dead host.
-
No backups available on Hobby (backups/PITR are Pro-only), and there is
no way to take one now since nothing can connect. My database also has a
scheduled CVE patch (
CVE-2026-15741) whose description says "a snapshot istaken first" — **could you trigger that snapshot before any container
recreate?** That would be the only copy of this data in existence.
Main question: **is the volume data intact, and will it be released to a healthy
host?**
2 months ago
Same confirmed host incident and signature on our production project.
Project: athletic-perfection (f3edb964-5792-4033-bc20-0788410b7805)
Environment: production (e63181af-c2d4-4a37-995d-bd07c2ed1d37)
Service: Postgres (1ebf1d7d-8df3-43b7-9c15-5ffee34e34a8)
Volume: postgres-volume (f43fdd26-0aa2-45b9-babb-489c50a0af77)
Region: US West (us-west2)
Image: ghcr.io/railwayapp-templates/postgres-ssl:18
Timeline/signature:
- Last normal Postgres log activity ended around 22:34 UTC with no FATAL, PANIC, OOM, or shutdown.
- Dashboard continued to report Online, but postgres.railway.internal resolves and TCP 5432 times out from a healthy sibling production service.
- railway ssh to Postgres fails; Railway Database UI remains at "Attempting to connect"; service metrics UI returns 500.
- Dependent app deployments fail health with "Connection terminated due to connection timeout".
- Two approved same-image/same-volume recovery attempts never started and produced zero logs:
- 23706617-7dfe-4085-a025-e4d60b5d4fb5
- 793ffb0d-3b03-4ad0-a2c8-b72516ec24fd
- Persistent volume has not been deleted, replaced, detached, or modified.
Production is fully down. Please force-release the volume/frozen workload and place it on a healthy host while preserving volume data. Please confirm data integrity and host recovery before we retry the application deployment.
2 months ago
We're going through the steps for a full reboot to get everyone back up. Will continue to update.
Status changed to Awaiting User Response Railway • about 2 months ago
2 months ago
Fresh production update for athletic-perfection:
- The existing postgres-volume was live-resized from 15.36 GB to 30 GB in Railway; it was not detached, deleted, replaced, or recreated.
- Postgres still has 0/1 running replicas.
- postgres.railway.internal now returns ENOTFOUND from a healthy sibling service.
- A separate database in the same project, 121-platform-postgres.railway.internal:5432, resolves and accepts TCP successfully, ruling out project-wide private networking.
- DATABASE_URL/PGHOST and the Postgres RAILWAY_PRIVATE_DOMAIN all correctly reference postgres.railway.internal:5432.
- railway ssh to Postgres still fails at the service layer.
- Postgres logs still end abruptly around 22:34 UTC after successful checkpoints, with no PANIC, OOM, shutdown, or no-space error.
- 121enterprise starts but reports getaddrinfo ENOTFOUND postgres.railway.internal; it has 0/1 healthy replicas and public IDX routes remain 502.
The announced 1-2 hour recovery window has passed and the full-reboot update is now over an hour old. Please confirm whether host reboot/volume release completed for volume f43fdd26-0aa2-45b9-babb-489c50a0af77. If not, please force-release/migrate this existing volume to a healthy host while preserving its data, and tell us when one controlled Postgres redeploy is safe.
Status changed to Awaiting Railway Response Railway • about 2 months ago
2 months ago
Will do, as an update to everyone on the thread, we are working through the migration steps to get everyone on a healthy host.
Status changed to Awaiting User Response Railway • about 2 months ago
2 months ago
Update to all affected, we have performed recovery for all workloads on the machine. If you are still observing impact, reply and our on-call team will engage.
We apologize for the impact. We advise all customers to implement HA configs when possible for mission critical workloads.
a month ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • about 1 month ago