cannot connect to postgres
deboracgs
PROOP

2 days ago

Since 20 minutes ago, the postgres is not working, all my system is down, the status page everything is green

Error: read ECONNRESET

at TCP.onStreamRead (node:internal/stream_base_commons:216:20)

at TCP.callbackTrampoline (node:internal/async_hooks:130:17)

image.png

image.png

Solved$20 Bounty

40 Replies

Railway
BOT

2 days ago

We haven't established what's causing this yet. The community can help investigate your setup.

That's exactly what the Railway community is good at, so we'd like to open your thread as a community bounty. Railway pays a bounty to the community member who solves it, and threads like this usually get picked up quickly.

Opening it makes this entire thread public, including everything already posted. Nothing becomes public until you decide. Use the buttons below.

  • Open to the community - Before you click, take a moment to edit or remove anything you'd rather not share. The thread becomes publicly visible right away.
  • Keep it private and close the thread - Nothing becomes public and the thread closes.

Status changed to Awaiting User Response Railway • 2 days ago


Railway
BOT

2 days ago

This thread has been opened as a public bounty so the community can help solve it. The thread and any further activity are now visible to everyone.

Status changed to Open Railway • 2 days ago


deboracgs
PROOP

2 days ago

still not being able to connect, even on the railway dashboard


Try selecting your service, press Cmd/Ctrl + K, then select deploy source image.


0x5b62656e5d

Try selecting your service, press Cmd/Ctrl + K, then select deploy source image.

deboracgs
PROOP

2 days ago

not sure if that's, because I didn't any deploy recently


jb55goodman-ui
PRO

2 days ago

This is a production outage. I'm having the same issue.


zelta
PRO

2 days ago

same issue


Which region is the database deployed in?


0x5b62656e5d

Which region is the database deployed in?

deboracgs
PROOP

2 days ago

US WEST


jb55goodman-ui
PRO

2 days ago

Us west


waspdrey
PRO

2 days ago

backup and restart


deboracgs
PROOP

2 days ago

I have 2 databases on us west and one is working and the other one is not


deboracgs

I have 2 databases on us west and one is working and the other one is not

waspdrey
PRO

2 days ago

Yes the one not working backup, update and restart


waspdrey

backup and restart

deboracgs
PROOP

2 days ago

backup is not even working


zelta
PRO

2 days ago

image.png

image.png


waspdrey

Yes the one not working backup, update and restart

deboracgs
PROOP

2 days ago

image.png

Attachments


zelta

![image.png](https://station-server.railway.com/attachments/att_01m4162ggxenf97h4463ajt42y) ![image.png](https://station-server.railway.com/attachments/att_01m4163bsvexw9kxdf49b2wv7j)

deboracgs
PROOP

2 days ago

same here


Let me check with the team.


zelta
PRO

2 days ago

Region: us-west2. Our hypothesis: the host backing our Postgres volume degraded around 14:00 UTC. Checkpoints writing only a few hundred buffers took 140–270 s, and shared memory hit "No space left on device" at 14:02. The container went silent at 14:15 with no shutdown or crash logs, yet Railway still lists it as SUCCESS. We suspect the volume is still attached to that unresponsive host, which would explain why two redeploys stall or fail at "Create container" with no logs. The volume itself is at 8.8/50 GB.


waspdrey
PRO

2 days ago

I believe it's the region but i have a template posted on railway templates you can try that it has everything you'll need automatically set up with pgbouncer https://railway.com/deploy/postgrespostgispgbouncer?referralCode=953jk4&utm_medium=integration&utm_source=template&utm_campaign=generic


zelta

Region: us-west2. Our hypothesis: the host backing our Postgres volume degraded around 14:00 UTC. Checkpoints writing only a few hundred buffers took 140–270 s, and shared memory hit "No space left on device" at 14:02. The container went silent at 14:15 with no shutdown or crash logs, yet Railway still lists it as SUCCESS. We suspect the volume is still attached to that unresponsive host, which would explain why two redeploys stall or fail at "Create container" with no logs. The volume itself is at 8.8/50 GB.

waspdrey
PRO

2 days ago

Right it's a region problem


waspdrey
PRO

2 days ago

You should also check out https://github.com/waspdrey/swarmauth


deboracgs

same here

waspdrey
PRO

2 days ago

Is your problem solved?


waspdrey

Is your problem solved?

deboracgs
PROOP

2 days ago

nope, still the same


deboracgs
PROOP

2 days ago

Is there no one from Railway who can provide support on this? This clearly seems to be an issue with the US-West region.

I don’t know if other users are experiencing the same problem, but in my case, I didn’t make any deployment, update, or configuration change. PostgreSQL simply went down unexpectedly, and now I have a large number of users who can’t access the app.

This is affecting production and needs to be addressed as soon as possible. Could someone from Railway please take a look at what’s happening with the PostgreSQL service in the US-West region?


jb55goodman-ui
PRO

2 days ago

Moderator said they were checking with the team. I assume or hope this is happening.


deboracgs

still not being able to connect, even on the railway dashboard

krishsoni2
PRO

2 days ago

i'm getting the same exact issue since morning


keshavsoni17
PRO

2 days ago

Same issue, but for Redis


0x5b62656e5d

Let me check with the team.

deboracgs
PROOP

2 days ago

do you have any updates? it's almost 2 hours like this


nevo-david
PRO

2 days ago

Hosting company and nobody answers. Nice.


2 days ago

same issue for singapore


deboracgs
PROOP

2 days ago

2 hours if the same problem, I'm losing money and no answers


keshavsoni17
PRO

2 days ago

lol crazy how they haven't responded yet, neither updated status page


2 days ago

Hi all, apologies for the delay in communication here. We are seeing some hardware failures on isolated hosts, and have notified affected customers via email. Some of these failure notices have failed to be sent, and we will be retroactively notifying affected customers.

This is not the standard we aim to uphold, and we're sorry for the lack of transparency. If you have not received a notification and your service is experiencing downtime, please open your a new support thread.


Status changed to Awaiting User Response Railway • 2 days ago


jb55goodman-ui
PRO

2 days ago

I'm not sure how you are evaluating affected customers or diagnosing. I do not have an email, I have a private thread open, I'm responding here. Private thread does not have a response. No email response.


Status changed to Awaiting Railway Response Railway • 2 days ago


zelta
PRO

2 days ago

@nico thanks for confirming. Our production Postgres (us-west2, service 67b1d65f-bed2-4812-ae95-62ba47c32f92) has been down since 14:15 UTC, almost 2.5 hours now. Two redeploys (d01ed640 and 12b1c19a) both failed at "Create container" with no logs, and we have not received any email.

Our company's entire operation is stopped because of this. Could you please tell us:

  1. Is our service on one of the affected hosts, and is there an ETA for recovery?
  2. Is there anything we can do right now to get back online without losing data? For example, can you move the volume to a healthy host, or should we restore a backup into the same service?

We'll avoid more redeploys until you confirm. Honestly, this isn't the first reliability issue we've had with Railway, so a clear ETA and plan would help a lot.


2 days ago

@zelta your Postgres is on one of the affected hosts. It stopped responding at about 14:15 UTC and was restarted at about 16:30 UTC. Your two redeploys failed because a database has to start on the same host as its volume, and that host was down. It's back now, so please redeploy the Postgres service once. Your volume and its data are still on that host, and the redeploy reattaches it, so don't restore a backup. If the redeploy fails, reply here and I'll look at it directly.

@jb55goodman-ui @deboracgs your databases were on affected hosts too, which have also been restarted. I'm checking both directly and will reply in your threads.


Status changed to Awaiting User Response Railway • 2 days ago


zelta

@nico thanks for confirming. Our production Postgres (us-west2, service 67b1d65f-bed2-4812-ae95-62ba47c32f92) has been down since 14:15 UTC, almost 2.5 hours now. Two redeploys (d01ed640 and 12b1c19a) both failed at "Create container" with no logs, and we have not received any email. Our company's entire operation is stopped because of this. Could you please tell us: 1. Is our service on one of the affected hosts, and is there an ETA for recovery? 2. Is there anything we can do right now to get back online without losing data? For example, can you move the volume to a healthy host, or should we restore a backup into the same service? We'll avoid more redeploys until you confirm. Honestly, this isn't the first reliability issue we've had with Railway, so a clear ETA and plan would help a lot.

jb55goodman-ui
PRO

2 days ago

I've experienced at least a handful of severe outages this year. My general issue with Railway is not getting to recovery, it's a continued failure to diagnose the affected customers and severity of the issue in a timely fashion. Too retroactively waiting for enough support tickets to roll in to escalate. I get that it doesn't help that its a weekend, but we are paying for uptime.


Status changed to Awaiting Railway Response Railway • 2 days ago


nico

@zelta your Postgres is on one of the affected hosts. It stopped responding at about 14:15 UTC and was restarted at about 16:30 UTC. Your two redeploys failed because a database has to start on the same host as its volume, and that host was down. It's back now, so please redeploy the Postgres service once. Your volume and its data are still on that host, and the redeploy reattaches it, so don't restore a backup. If the redeploy fails, reply here and I'll look at it directly. @jb55goodman-ui @deboracgs your databases were on affected hosts too, which have also been restarted. I'm checking both directly and will reply in your threads.

jb55goodman-ui
PRO

2 days ago

Ty Nico, will monitor and report back here.


zelta
PRO

2 days ago

@nico thanks, the redeploy (c8ee34a7) succeeded and Postgres logged "database system is ready to accept connections" at 17:01:16 UTC after a clean crash recovery. But it is still unreachable: our services time out on postgres.railway.internal:5432 (10.181.195.128 and fd12:e1c3:480e:1:2000:71:feb5:c380), and the TCP proxy closes connections. Postgres logs show no incoming connection attempts. One connection succeeded around 17:01:36, then it dropped again. Could you check networking to that host? We're not redeploying again.


nico

@zelta your Postgres is on one of the affected hosts. It stopped responding at about 14:15 UTC and was restarted at about 16:30 UTC. Your two redeploys failed because a database has to start on the same host as its volume, and that host was down. It's back now, so please redeploy the Postgres service once. Your volume and its data are still on that host, and the redeploy reattaches it, so don't restore a backup. If the redeploy fails, reply here and I'll look at it directly. @jb55goodman-ui @deboracgs your databases were on affected hosts too, which have also been restarted. I'm checking both directly and will reply in your threads.

deboracgs
PROOP

2 days ago

Thanks, it's working now


Status changed to Solved Railway • 2 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...