2 days ago
Since 20 minutes ago, the postgres is not working, all my system is down, the status page everything is green
Error: read ECONNRESET
at TCP.onStreamRead (node:internal/stream_base_commons:216:20)
at TCP.callbackTrampoline (node:internal/async_hooks:130:17)40 Replies
2 days ago
We haven't established what's causing this yet. The community can help investigate your setup.
That's exactly what the Railway community is good at, so we'd like to open your thread as a community bounty. Railway pays a bounty to the community member who solves it, and threads like this usually get picked up quickly.
Opening it makes this entire thread public, including everything already posted. Nothing becomes public until you decide. Use the buttons below.
- Open to the community - Before you click, take a moment to edit or remove anything you'd rather not share. The thread becomes publicly visible right away.
- Keep it private and close the thread - Nothing becomes public and the thread closes.
Status changed to Awaiting User Response Railway • 2 days ago
2 days ago
This thread has been opened as a public bounty so the community can help solve it. The thread and any further activity are now visible to everyone.
Status changed to Open Railway • 2 days ago
2 days ago
still not being able to connect, even on the railway dashboard
2 days ago
Try selecting your service, press Cmd/Ctrl + K, then select deploy source image.
0x5b62656e5d
Try selecting your service, press Cmd/Ctrl + K, then select deploy source image.
2 days ago
not sure if that's, because I didn't any deploy recently
0x5b62656e5d
Which region is the database deployed in?
2 days ago
US WEST
2 days ago
Us west
2 days ago
backup and restart
2 days ago
I have 2 databases on us west and one is working and the other one is not
deboracgs
I have 2 databases on us west and one is working and the other one is not
2 days ago
Yes the one not working backup, update and restart
waspdrey
backup and restart
2 days ago
backup is not even working
waspdrey
Yes the one not working backup, update and restart
2 days ago
Attachments
zelta
 
2 days ago
same here
2 days ago
Region: us-west2. Our hypothesis: the host backing our Postgres volume degraded around 14:00 UTC. Checkpoints writing only a few hundred buffers took 140–270 s, and shared memory hit "No space left on device" at 14:02. The container went silent at 14:15 with no shutdown or crash logs, yet Railway still lists it as SUCCESS. We suspect the volume is still attached to that unresponsive host, which would explain why two redeploys stall or fail at "Create container" with no logs. The volume itself is at 8.8/50 GB.
2 days ago
I believe it's the region but i have a template posted on railway templates you can try that it has everything you'll need automatically set up with pgbouncer https://railway.com/deploy/postgrespostgispgbouncer?referralCode=953jk4&utm_medium=integration&utm_source=template&utm_campaign=generic
zelta
Region: us-west2. Our hypothesis: the host backing our Postgres volume degraded around 14:00 UTC. Checkpoints writing only a few hundred buffers took 140–270 s, and shared memory hit "No space left on device" at 14:02. The container went silent at 14:15 with no shutdown or crash logs, yet Railway still lists it as SUCCESS. We suspect the volume is still attached to that unresponsive host, which would explain why two redeploys stall or fail at "Create container" with no logs. The volume itself is at 8.8/50 GB.
2 days ago
Right it's a region problem
deboracgs
same here
2 days ago
Is your problem solved?
waspdrey
Is your problem solved?
2 days ago
nope, still the same
2 days ago
Is there no one from Railway who can provide support on this? This clearly seems to be an issue with the US-West region.
I don’t know if other users are experiencing the same problem, but in my case, I didn’t make any deployment, update, or configuration change. PostgreSQL simply went down unexpectedly, and now I have a large number of users who can’t access the app.
This is affecting production and needs to be addressed as soon as possible. Could someone from Railway please take a look at what’s happening with the PostgreSQL service in the US-West region?
2 days ago
Moderator said they were checking with the team. I assume or hope this is happening.
deboracgs
still not being able to connect, even on the railway dashboard
2 days ago
i'm getting the same exact issue since morning
0x5b62656e5d
Let me check with the team.
2 days ago
do you have any updates? it's almost 2 hours like this
2 days ago
2 hours if the same problem, I'm losing money and no answers
2 days ago
lol crazy how they haven't responded yet, neither updated status page
2 days ago
Hi all, apologies for the delay in communication here. We are seeing some hardware failures on isolated hosts, and have notified affected customers via email. Some of these failure notices have failed to be sent, and we will be retroactively notifying affected customers.
This is not the standard we aim to uphold, and we're sorry for the lack of transparency. If you have not received a notification and your service is experiencing downtime, please open your a new support thread.
Status changed to Awaiting User Response Railway • 2 days ago
2 days ago
I'm not sure how you are evaluating affected customers or diagnosing. I do not have an email, I have a private thread open, I'm responding here. Private thread does not have a response. No email response.
Status changed to Awaiting Railway Response Railway • 2 days ago
2 days ago
@nico thanks for confirming. Our production Postgres (us-west2, service 67b1d65f-bed2-4812-ae95-62ba47c32f92) has been down since 14:15 UTC, almost 2.5 hours now. Two redeploys (d01ed640 and 12b1c19a) both failed at "Create container" with no logs, and we have not received any email.
Our company's entire operation is stopped because of this. Could you please tell us:
- Is our service on one of the affected hosts, and is there an ETA for recovery?
- Is there anything we can do right now to get back online without losing data? For example, can you move the volume to a healthy host, or should we restore a backup into the same service?
We'll avoid more redeploys until you confirm. Honestly, this isn't the first reliability issue we've had with Railway, so a clear ETA and plan would help a lot.
2 days ago
@zelta your Postgres is on one of the affected hosts. It stopped responding at about 14:15 UTC and was restarted at about 16:30 UTC. Your two redeploys failed because a database has to start on the same host as its volume, and that host was down. It's back now, so please redeploy the Postgres service once. Your volume and its data are still on that host, and the redeploy reattaches it, so don't restore a backup. If the redeploy fails, reply here and I'll look at it directly.
@jb55goodman-ui @deboracgs your databases were on affected hosts too, which have also been restarted. I'm checking both directly and will reply in your threads.
Status changed to Awaiting User Response Railway • 2 days ago
zelta
@nico thanks for confirming. Our production Postgres (us-west2, service 67b1d65f-bed2-4812-ae95-62ba47c32f92) has been down since 14:15 UTC, almost 2.5 hours now. Two redeploys (d01ed640 and 12b1c19a) both failed at "Create container" with no logs, and we have not received any email. Our company's entire operation is stopped because of this. Could you please tell us: 1. Is our service on one of the affected hosts, and is there an ETA for recovery? 2. Is there anything we can do right now to get back online without losing data? For example, can you move the volume to a healthy host, or should we restore a backup into the same service? We'll avoid more redeploys until you confirm. Honestly, this isn't the first reliability issue we've had with Railway, so a clear ETA and plan would help a lot.
2 days ago
I've experienced at least a handful of severe outages this year. My general issue with Railway is not getting to recovery, it's a continued failure to diagnose the affected customers and severity of the issue in a timely fashion. Too retroactively waiting for enough support tickets to roll in to escalate. I get that it doesn't help that its a weekend, but we are paying for uptime.
Status changed to Awaiting Railway Response Railway • 2 days ago
nico
@zelta your Postgres is on one of the affected hosts. It stopped responding at about 14:15 UTC and was restarted at about 16:30 UTC. Your two redeploys failed because a database has to start on the same host as its volume, and that host was down. It's back now, so please redeploy the Postgres service once. Your volume and its data are still on that host, and the redeploy reattaches it, so don't restore a backup. If the redeploy fails, reply here and I'll look at it directly. @jb55goodman-ui @deboracgs your databases were on affected hosts too, which have also been restarted. I'm checking both directly and will reply in your threads.
2 days ago
Ty Nico, will monitor and report back here.
2 days ago
@nico thanks, the redeploy (c8ee34a7) succeeded and Postgres logged "database system is ready to accept connections" at 17:01:16 UTC after a clean crash recovery. But it is still unreachable: our services time out on postgres.railway.internal:5432 (10.181.195.128 and fd12:e1c3:480e:1:2000:71:feb5:c380), and the TCP proxy closes connections. Postgres logs show no incoming connection attempts. One connection succeeded around 17:01:36, then it dropped again. Could you check networking to that host? We're not redeploying again.
nico
@zelta your Postgres is on one of the affected hosts. It stopped responding at about 14:15 UTC and was restarted at about 16:30 UTC. Your two redeploys failed because a database has to start on the same host as its volume, and that host was down. It's back now, so please redeploy the Postgres service once. Your volume and its data are still on that host, and the redeploy reattaches it, so don't restore a backup. If the redeploy fails, reply here and I'll look at it directly. @jb55goodman-ui @deboracgs your databases were on affected hosts too, which have also been restarted. I'm checking both directly and will reply in your threads.
2 days ago
Thanks, it's working now
Status changed to Solved Railway • 2 days ago


