2 months ago
Project: QuizPilot (fe3d7b8d-f9f9-40e2-b740-cf04d28df285)
Environment: production (bb806758-43ba-44b0-aaa7-294beb4b3d2a)
Service: Postgres (0ccd334e-4d38-48a5-8c90-85c4fc2de3a8), image ghcr.io/railwayapp-templates/postgres-ssl:17
Volume: postgres-volume (3dbd00e8-97f9-4e69-8cd7-1547bd5aab5d), 19660MB / 50000MB used
Stuck deployment: 47b5ffe7-da13-4e79-a572-13c3913602da
TIMELINE (all UTC)
16:42:38 Last log line the Postgres container has EVER written (a normal checkpoint complete).
16:47 / 16:52 / 16:57 ... scheduled 5-minute checkpoints never fired. Container went silent mid-service.
16:43-16:44 All client services begin failing. web-api / telegram-bot / Worker Service all raise the
identical trace: asyncpg _create_ssl_connection -> CancelledError -> TimeoutError,
surfacing through SQLAlchemy pool _do_get. Every NEW connection dies in the handshake.
External psql over the TCP proxy: 0 of 12+ attempts succeeded, each 7-8s then either
"server closed the connection unexpectedly" or "timeout expired".
TCP to switchyard.proxy.rlwy.net:51258 connects fine, so the proxy is up - whatever is
behind it never completes the Postgres handshake.17:03:25 I issued railway redeploy -s Postgres.
17:14 Still status DEPLOYING, 11 minutes later. The new deployment has produced ZERO log lines -
not one, not even a startup message.WHAT THIS LOOKS LIKE FROM OUTSIDE
The new container cannot start because the old one has not released postgres-volume, and the old one
is not dying. A process that stops writing its own log, stops checkpointing, hangs every new backend
at startup, and then cannot be terminated by a redeploy reads as blocked in uninterruptible I/O wait
on the underlying storage.
2 Replies
2 months ago
The deployment you triggered at 17:03 UTC failed at the container-creation step because the host it targeted was unreachable at that time, and the deployment never produced any logs. That host event has since resolved and the deploy backlog for your workspace is now clear. Your postgres-volume is intact and in READY state with ~19660 MB used. The Postgres service's latest deployment is still the failed one, so it is not currently running. To bring it back, open the Postgres service, press Cmd+K (or Ctrl+K), and select "Redeploy source image" to pull a fresh image onto the now-available infrastructure.
Status changed to Awaiting User Response Railway • about 2 months ago
Status changed to Solved tommyvahabov • about 2 months ago
2 months ago
Hi Rahmonberdi, human here, following up properly on this.
Your Postgres is back up as of 18:20 UTC. It ran automatic crash recovery cleanly on startup (recovery from WAL completed, "database system is ready to accept connections"), and your data is intact. Please verify your services have reconnected.
To be straight with you about what happened: the physical host holding your postgres-volume became unreachable at around 16:43 UTC and did not fully recover until roughly 18:15 UTC. Because your postgres-volume lives on that host, none of the redeploys could start until it recovered, which is why they sat in DEPLOYING with no logs. The earlier automated reply told you the host issue was resolved before it actually was, and the redeploys it suggested could not have worked at that point. I'm sorry for that, and for the fact that a production outage didn't reach a human sooner.
On our side, the host has been taken out of rotation for new workloads while we investigate.
Again, apologies for the outage and for the poor first response. If you see anything abnormal in the database since recovery, reply here and I'll dig in immediately.
Kind regards,
Sam
Status changed to Awaiting User Response Railway • about 2 months ago
Status changed to Solved sam-a • about 2 months ago