Production database is down and a restart is stuck. I need someone with host access.
tommyvahabov
PROOP

2 months ago

Project: QuizPilot (fe3d7b8d-f9f9-40e2-b740-cf04d28df285)

Environment: production (bb806758-43ba-44b0-aaa7-294beb4b3d2a)

Service: Postgres (0ccd334e-4d38-48a5-8c90-85c4fc2de3a8), image ghcr.io/railwayapp-templates/postgres-ssl:17

Volume: postgres-volume (3dbd00e8-97f9-4e69-8cd7-1547bd5aab5d), 19660MB / 50000MB used

Stuck deployment: 47b5ffe7-da13-4e79-a572-13c3913602da

TIMELINE (all UTC)

16:42:38 Last log line the Postgres container has EVER written (a normal checkpoint complete).

16:47 / 16:52 / 16:57 ... scheduled 5-minute checkpoints never fired. Container went silent mid-service.

16:43-16:44 All client services begin failing. web-api / telegram-bot / Worker Service all raise the

         identical trace: asyncpg _create_ssl_connection -> CancelledError -> TimeoutError,

         surfacing through SQLAlchemy pool _do_get. Every NEW connection dies in the handshake.

         External psql over the TCP proxy: 0 of 12+ attempts succeeded, each 7-8s then either

         "server closed the connection unexpectedly" or "timeout expired".

         TCP to switchyard.proxy.rlwy.net:51258 connects fine, so the proxy is up - whatever is

         behind it never completes the Postgres handshake.

17:03:25 I issued railway redeploy -s Postgres.

17:14 Still status DEPLOYING, 11 minutes later. The new deployment has produced ZERO log lines -

      not one, not even a startup message.

WHAT THIS LOOKS LIKE FROM OUTSIDE

The new container cannot start because the old one has not released postgres-volume, and the old one

is not dying. A process that stops writing its own log, stops checkpointing, hangs every new backend

at startup, and then cannot be terminated by a redeploy reads as blocked in uninterruptible I/O wait

on the underlying storage.

Solved

2 Replies

Railway
BOT

2 months ago

The deployment you triggered at 17:03 UTC failed at the container-creation step because the host it targeted was unreachable at that time, and the deployment never produced any logs. That host event has since resolved and the deploy backlog for your workspace is now clear. Your postgres-volume is intact and in READY state with ~19660 MB used. The Postgres service's latest deployment is still the failed one, so it is not currently running. To bring it back, open the Postgres service, press Cmd+K (or Ctrl+K), and select "Redeploy source image" to pull a fresh image onto the now-available infrastructure.


Status changed to Awaiting User Response Railway • about 2 months ago


Status changed to Solved tommyvahabov • about 2 months ago


sam-a
EMPLOYEE

2 months ago

Hi Rahmonberdi, human here, following up properly on this.

Your Postgres is back up as of 18:20 UTC. It ran automatic crash recovery cleanly on startup (recovery from WAL completed, "database system is ready to accept connections"), and your data is intact. Please verify your services have reconnected.

To be straight with you about what happened: the physical host holding your postgres-volume became unreachable at around 16:43 UTC and did not fully recover until roughly 18:15 UTC. Because your postgres-volume lives on that host, none of the redeploys could start until it recovered, which is why they sat in DEPLOYING with no logs. The earlier automated reply told you the host issue was resolved before it actually was, and the redeploys it suggested could not have worked at that point. I'm sorry for that, and for the fact that a production outage didn't reach a human sooner.

On our side, the host has been taken out of rotation for new workloads while we investigate.

Again, apologies for the outage and for the poor first response. If you see anything abnormal in the database since recovery, reply here and I'll dig in immediately.

Kind regards,

Sam


Status changed to Awaiting User Response Railway • about 2 months ago


Status changed to Solved sam-a • about 2 months ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...