2 months ago
For following project, my postgres deployment is not connecting to postgres disk.
Project: 8231d9a6-da27-400e-87f9-da5b32dbc993
Service (Postgres): be413d86-484b-4207-b580-749658d61565
Environment: a43e30a8-a478-4c55-a023-e8097bc5e8b0
Deployment: 56b8c36d-f432-499e-b1b9-0f9d603cc891
I saw the hardware-failure notification, which is marked Resolved, but the service did not automatically resume — Restart returns "Problem processing request", the console returns "routing failed", and a new deployment hung at "Creating containers" for 11 minutes.
Probably something is stuck!.
8 Replies
Status changed to Awaiting Railway Response Railway • about 2 months ago
2 months ago
Error logs from backend
[2026-08-18T17:22:49.980994662Z] [error] npm warn config production Use --omit=dev instead.
[2026-08-18T17:22:50.681615522Z] [info] Prisma schema loaded from prisma/schema.prisma
[2026-08-18T17:22:50.691939375Z] [info] Datasource "db": PostgreSQL database "railway", schema "public" at "postgres.railway.internal:5432"
[2026-08-18T17:22:55.754508745Z] [info]
[2026-08-18T17:22:55.754512335Z] [error] Error: P1001: Can't reach database server at postgres.railway.internal:5432
[2026-08-18T17:22:55.754563295Z] [error]
[2026-08-18T17:22:55.754566875Z] [error] Please make sure your database server is running at postgres.railway.internal:5432.
[2026-08-18T17:22:56.712906334Z] [error] npm warn config production Use --omit=dev instead.
[2026-08-18T17:22:56.745894827Z] [info]
[2026-08-18T17:22:56.745901837Z] [info] > blogapp-backend@1.0.0 start
2 months ago
There was a service outage, which tell that it was resolved, but it isn't
Attachments
2 months ago
Recent logs from postgres container
2026-08-18 16:52:33.444 UTC [8361] LOG: could not accept SSL connection: EOF detected
2026-08-18 16:52:38.028 UTC [8364] LOG: failed to send SSL negotiation response: Broken pipe
2026-08-18 16:52:39.570 UTC [8365] LOG: failed to send SSL negotiation response: Broken pipe
2026-08-18 16:52:44.980 UTC [8367] LOG: could not accept SSL connection: Connection reset by peer
2026-08-18 16:52:48.075 UTC [8366] LOG: could not accept SSL connection: Connection reset by peer
2026-08-18 16:56:21.601 UTC [8377] LOG: could not accept SSL connection: EOF detected
2026-08-18 16:53:10.029 UTC [8370] LOG: could not accept SSL connection: Connection reset by peer
2026-08-18 16:53:11.941 UTC [8368] LOG: could not accept SSL connection: Connection reset by peer
2026-08-18 16:53:18.360 UTC [8371] LOG: could not accept SSL connection: Connection reset by peer
2026-08-18 16:54:13.966 UTC [8372] LOG: could not accept SSL connection: Connection reset by peer
2026-08-18 16:56:09.809 UTC [8375] LOG: could not accept SSL connection: Connection reset by peer
2026-08-18 16:56:14.821 UTC [8376] LOG: could not accept SSL connection: Connection reset by peer
2026-08-18 16:56:27.929 UTC [8378] LOG: could not accept SSL connection: EOF detected
2026-08-18 16:56:34.488 UTC [8380] LOG: could not accept SSL connection: EOF detected
2026-08-18 16:56:44.835 UTC [8385] LOG: could not accept SSL connection: EOF detected
2026-08-18 16:56:44.827 UTC [8383] LOG: could not accept SSL connection: EOF detected
2026-08-18 16:56:44.838 UTC [8389] LOG: could not accept SSL connection: EOF detected
2026-08-18 16:56:45.007 UTC [8386] LOG: could not accept SSL connection: EOF detected
2026-08-18 16:56:45.118 UTC [8384] LOG: could not accept SSL connection: EOF detected
2026-08-18 16:56:45.504 UTC [8388] LOG: could not accept SSL connection: EOF detected
2026-08-18 16:56:45.642 UTC [8390] LOG: could not accept SSL connection: EOF detected
2 months ago
Backup creation also failed. Please check and help here @railway team
2 months ago
I have the same issue, effecting hundreds of thousand users
2 months ago
You're right that it isn't actually back yet, despite the resolved notice. One of our servers hit a hardware fault and went offline, which is what took your Postgres down and is why the restarts and new deploys are hanging at creating containers. Your volume and its data are safe and persist through this. We're moving the workload onto a healthy machine now and are actively on it.
Status changed to Awaiting User Response Railway • about 2 months ago
2 months ago
The reattach is fully complete and it is safe to resume use now. No data has been affected. There is unfortunately a small risk of 2-5 mins downtime in the next 24h as we drain the stacker completely to prevent recurrence. My apologies for this.
a month ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • about 1 month ago