a month ago
new deployments with a documentation only change started failing health checks on our staging environment. Prior healthy and still active deployment using a xxx.up.railway.com domain is no longer reachable and gets a 502 Gateway error when trying to open it.
====================
Starting Healthcheck
====================
Path: /
Retry window: 5m0s
Attempt #1 failed with service unavailable. Continuing to retry for 4m56s
Attempt #2 failed with service unavailable. Continuing to retry for 4m55s
Attempt #3 failed with service unavailable. Continuing to retry for 4m53s
Attempt #4 failed with service unavailable. Continuing to retry for 4m49s
Attempt #5 failed with service unavailable. Continuing to retry for 4m41s
Attempt #6 failed with service unavailable. Continuing to retry for 4m25s
Attempt #7 failed with service unavailable. Continuing to retry for 3m55s
Attempt #8 failed with service unavailable. Continuing to retry for 3m25s
Attempt #9 failed with service unavailable. Continuing to retry for 2m55s
Attempt #10 failed with service unavailable. Continuing to retry for 2m25s
Attempt #11 failed with service unavailable. Continuing to retry for 1m55s
Attempt #12 failed with service unavailable. Continuing to retry for 1m25s
Pinned Solution
a month ago
well, I figured it out so I'll leave it here for posterity.
Our health check endpoint has been '/' for the lifetime of the project and it's always returned a 303 since our service requires login.
However, starting earlier today the health check requires 200 as a response, so changing just the route to an endpoint that returns 200 instead instantly passes.
The errors in the railway health check need to indicate if a response is received, but with the wrong code vs just saying 'unavailable'...
11 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 1 month ago
a month ago
I haven't tried doing that yet as that too takes a PR to strike it -- but the 502 gateway error while still having an active 'good' deployment on staging has been what I've been investigating. My assumption has been that the server is fine, but nothing can resolve to it...
a month ago
Was there any change made to the network configuration in the service settings, like changing the target port, or modifying the PORT variable of your service? Also, removing the healthcheck doesn't require a new PR, you just have to remove it from service settings, and hit apply changes, which will redeploy the service.
a month ago
This error is caused when Railway's edge routes the traffic to your service, but your app doesn't receive it, because it has an incorrect port/host binding in most cases. Also, healthchecks require the PORT env variable to be set, so it knows which port to route the request to. Hopefully, this helps you in narrowing down what the issue is.
a month ago
looking at the config, none of the network settings have changed in over 6 months... it is set to the correct port.
our config is set with a railway file, the web UI will not let you modify it in that case it needs to be a new code change.
a month ago
I've disabled the health check and the system does come up fine now.
It doesn't really answer why the health checks are failing when the server is actually up. Any ideas or things do debug?
a month ago
That means the app is binding correctly. I assume the PORT variable is set to the same port your app is listening on?
a month ago
We've never had the PORT variable needing to be set (we use 8080), i tried setting per your suggestion while debugging but it made no difference either way
a month ago
currently the only change made to the server was disabling the health check -- which has worked, and the route set for the health check works fine after being deployed...
a month ago
well, I figured it out so I'll leave it here for posterity.
Our health check endpoint has been '/' for the lifetime of the project and it's always returned a 303 since our service requires login.
However, starting earlier today the health check requires 200 as a response, so changing just the route to an endpoint that returns 200 instead instantly passes.
The errors in the railway health check need to indicate if a response is received, but with the wrong code vs just saying 'unavailable'...
a month ago
Thank you for your suggestions darseen!
Status changed to Solved dev • about 1 month ago