2 months ago
I have no idea why but yesterday all apps went down in production for almost 1 hour. It seems like I didn't have nothing to do with it no bug or something like that but something that was initiated by railway. Please check and tell me why this happened and who decided it I've got a lot of complaints from customers.
3 Replies
2 months ago
A server hosting your Postgres database in production experienced a hardware failure yesterday at around 17:14 UTC, and the database was automatically migrated to a healthy server. That migration restarted Postgres, which would have caused all services depending on it to lose their database connections and go down until the migration completed (roughly 60 minutes). This was not initiated by anyone on our team - it was an automated response to a hardware event, and your services should have recovered on their own once the database came back online. All your production services are currently running normally.
Status changed to Awaiting User Response Railway • about 2 months ago
2 months ago
Thank you for the explanation. I understand that hardware failures can happen, but I’m concerned that a failure on a single server hosting our production Postgres database caused our entire production environment to remain unavailable for roughly 60 minutes.
Since Railway is managing the infrastructure and database hosting, could you please clarify:
- Is an outage of this duration considered normal or expected for a Railway-managed Postgres database after a hardware failure?
- Why was there no automatic failover to a standby database that could keep our applications available?
- Does our current Railway plan include any high-availability or redundancy protection for Postgres?
- What architecture, configuration, or Railway service should we use to prevent a single database host failure from taking down all our production applications in the future?
- Is there an SLA or documented recovery-time target for this type of incident?
- Can you confirm whether the approximately 60-minute migration and recovery time was expected, or whether something delayed the recovery?
We are expecting significant growth over the next few months and plan to expand our customer base substantially. As Bumpy.fm becomes responsible for more customers and live events, an hour-long outage could have a serious operational and financial impact.
We therefore need to ensure that this type of incident cannot take down our entire platform again. Please provide a clear recommendation for the production architecture, failover mechanisms, redundancy, backups, recovery process, and Railway plan required to achieve the highest possible availability, including any necessary changes and their associated costs.
Status changed to Awaiting Railway Response Railway • about 2 months ago
2 months ago
This is a question the Railway community is better placed to answer than support: people who have already worked this out on their own projects and can tell you what actually worked.
So we'd like to open your thread as a community bounty. Railway pays a bounty to the community member who answers it, and threads like this usually get picked up quickly.
Opening it makes this entire thread public, including everything already posted. Nothing becomes public until you decide. Use the buttons below.
- Open to the community - Before you click, take a moment to edit or remove anything you'd rather not share. The thread becomes publicly visible right away.
- Keep it private and close the thread - Nothing becomes public and the thread closes.
Status changed to Awaiting User Response Railway • about 2 months ago
2 months ago
This thread has been opened as a public bounty so the community can help solve it. The thread and any further activity are now visible to everyone.
Status changed to Open Railway • about 2 months ago