a month ago
I received the below email but no further details provided. Why would my service by down for 30 minutes! The server shows as online in the dashboard!
Server Outage Affecting Your Service
A server running your service has experienced a hardware failure. We are actively working to restore it. No action is needed on your part — your service will automatically resume once the issue is resolved.
7 Replies
a month ago
The email you received is a targeted notification for a hardware failure on the server hosting your "qaflah" service, which is still being resolved. The dashboard shows the service as online because it reflects the deployment's own status, not the health of the underlying server. No action is needed on your part while recovery is in progress, and your data remains intact on the affected server's disk. For mission-critical workloads, we advise implementing a high-availability configuration so that a single-server failure does not take production down.
Status changed to Awaiting User Response Railway • 29 days ago
Railway
The email you received is a targeted notification for a hardware failure on the server hosting your "qaflah" service, which is still being resolved. The dashboard shows the service as online because it reflects the deployment's own status, not the health of the underlying server. No action is needed on your part while recovery is in progress, and your data remains intact on the affected server's disk. For mission-critical workloads, we advise implementing a high-availability configuration so that a single-server failure does not take production down.
a month ago
Understood, but be it a high-availability or otherwise, a random "technical failure" without further details isnt something that should be okay right? Please suggest way of conversion of the current deployment to a high-availability one as well as if it should be done right now or after the resolution
Status changed to Awaiting Railway Response Railway • 29 days ago
a month ago
You're right that a bare "technical failure" notice is not ideal, and we appreciate the feedback on that. The hardware failure affecting your service is still being resolved, and your service will resume automatically once it is, with your volume data intact. For the HA question: multi-region replicas place each replica on a physically separate server by construction, so a single-host failure only takes down one replica while the others keep serving traffic. You can configure regions from the Scale section in your service settings. Wait until the current issue is resolved before making changes, since no configuration change on your side can take effect while the underlying host is down.
Status changed to Awaiting User Response Railway • 29 days ago
Railway
You're right that a bare "technical failure" notice is not ideal, and we appreciate the feedback on that. The hardware failure affecting your service is still being resolved, and your service will resume automatically once it is, with your volume data intact. For the HA question: [multi-region replicas](https://docs.railway.com/deployments/scaling#multi-region-replicas) place each replica on a physically separate server by construction, so a single-host failure only takes down one replica while the others keep serving traffic. You can configure regions from the Scale section in your service settings. Wait until the current issue is resolved before making changes, since no configuration change on your side can take effect while the underlying host is down.
a month ago
1 hour has passed, service still down, need an update and resolution please
Status changed to Awaiting Railway Response Railway • 29 days ago
Status changed to Awaiting User Response Railway • 29 days ago
Railway
The email you received is a targeted notification for a hardware failure on the server hosting your "qaflah" service, which is still being resolved. The dashboard shows the service as online because it reflects the deployment's own status, not the health of the underlying server. No action is needed on your part while recovery is in progress, and your data remains intact on the affected server's disk. For mission-critical workloads, we advise implementing a high-availability configuration so that a single-server failure does not take production down.
a month ago
ETA on recovery please? 30 minutes quoted (i understand its an estimate) but over 2 hours have passed now
Status changed to Awaiting Railway Response Railway • 29 days ago
a month ago
The server-side recovery finished a while ago, but your service did not fully restart with it, which is why it is still down. A fresh deploy from your side will bring it back immediately: open the command palette (Cmd+K or Ctrl+K) and select the deploy option for the service. Your volume data is intact and unaffected by the redeploy.
Status changed to Awaiting User Response Railway • 29 days ago
Railway
The server-side recovery finished a while ago, but your service did not fully restart with it, which is why it is still down. A fresh deploy from your side will bring it back immediately: open the command palette (Cmd+K or Ctrl+K) and select the deploy option for the service. Your volume data is intact and unaffected by the redeploy.
a month ago
thank you, received NO notification of it at all hence had to ask.
Status changed to Awaiting Railway Response Railway • 29 days ago
Status changed to Solved startupsyndicate • 29 days ago