downtime.- "technical failure"
startupsyndicate
HOBBYOP

a month ago

I received the below email but no further details provided. Why would my service by down for 30 minutes! The server shows as online in the dashboard!

Server Outage Affecting Your Service

A server running your service has experienced a hardware failure. We are actively working to restore it. No action is needed on your part — your service will automatically resume once the issue is resolved.

Solved

7 Replies

Railway
BOT

a month ago

The email you received is a targeted notification for a hardware failure on the server hosting your "qaflah" service, which is still being resolved. The dashboard shows the service as online because it reflects the deployment's own status, not the health of the underlying server. No action is needed on your part while recovery is in progress, and your data remains intact on the affected server's disk. For mission-critical workloads, we advise implementing a high-availability configuration so that a single-server failure does not take production down.


Status changed to Awaiting User Response Railway • 29 days ago


Railway

The email you received is a targeted notification for a hardware failure on the server hosting your "qaflah" service, which is still being resolved. The dashboard shows the service as online because it reflects the deployment's own status, not the health of the underlying server. No action is needed on your part while recovery is in progress, and your data remains intact on the affected server's disk. For mission-critical workloads, we advise implementing a high-availability configuration so that a single-server failure does not take production down.

startupsyndicate
HOBBYOP

a month ago

Understood, but be it a high-availability or otherwise, a random "technical failure" without further details isnt something that should be okay right? Please suggest way of conversion of the current deployment to a high-availability one as well as if it should be done right now or after the resolution


Status changed to Awaiting Railway Response Railway • 29 days ago


Railway
BOT

a month ago

You're right that a bare "technical failure" notice is not ideal, and we appreciate the feedback on that. The hardware failure affecting your service is still being resolved, and your service will resume automatically once it is, with your volume data intact. For the HA question: multi-region replicas place each replica on a physically separate server by construction, so a single-host failure only takes down one replica while the others keep serving traffic. You can configure regions from the Scale section in your service settings. Wait until the current issue is resolved before making changes, since no configuration change on your side can take effect while the underlying host is down.


Status changed to Awaiting User Response Railway • 29 days ago


Railway

You're right that a bare "technical failure" notice is not ideal, and we appreciate the feedback on that. The hardware failure affecting your service is still being resolved, and your service will resume automatically once it is, with your volume data intact. For the HA question: [multi-region replicas](https://docs.railway.com/deployments/scaling#multi-region-replicas) place each replica on a physically separate server by construction, so a single-host failure only takes down one replica while the others keep serving traffic. You can configure regions from the Scale section in your service settings. Wait until the current issue is resolved before making changes, since no configuration change on your side can take effect while the underlying host is down.

startupsyndicate
HOBBYOP

a month ago

1 hour has passed, service still down, need an update and resolution please


Status changed to Awaiting Railway Response Railway • 29 days ago


Status changed to Awaiting User Response Railway • 29 days ago


Railway

The email you received is a targeted notification for a hardware failure on the server hosting your "qaflah" service, which is still being resolved. The dashboard shows the service as online because it reflects the deployment's own status, not the health of the underlying server. No action is needed on your part while recovery is in progress, and your data remains intact on the affected server's disk. For mission-critical workloads, we advise implementing a high-availability configuration so that a single-server failure does not take production down.

startupsyndicate
HOBBYOP

a month ago

ETA on recovery please? 30 minutes quoted (i understand its an estimate) but over 2 hours have passed now


Status changed to Awaiting Railway Response Railway • 29 days ago


Railway
BOT

a month ago

The server-side recovery finished a while ago, but your service did not fully restart with it, which is why it is still down. A fresh deploy from your side will bring it back immediately: open the command palette (Cmd+K or Ctrl+K) and select the deploy option for the service. Your volume data is intact and unaffected by the redeploy.


Status changed to Awaiting User Response Railway • 29 days ago


Railway

The server-side recovery finished a while ago, but your service did not fully restart with it, which is why it is still down. A fresh deploy from your side will bring it back immediately: open the command palette (Cmd+K or Ctrl+K) and select the deploy option for the service. Your volume data is intact and unaffected by the redeploy.

startupsyndicate
HOBBYOP

a month ago

thank you, received NO notification of it at all hence had to ask.


Status changed to Awaiting Railway Response Railway • 29 days ago


Status changed to Solved startupsyndicate • 29 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...