Getting 502 Gateway errors on existing deployment and all new deployments fail health checks
benpoieszhv
PROOP

a month ago

new deployments with a documentation only change started failing health checks on our staging environment. Prior healthy and still active deployment using a xxx.up.railway.com domain is no longer reachable and gets a 502 Gateway error when trying to open it.

====================

Starting Healthcheck

====================

Path: /

Retry window: 5m0s

Attempt #1 failed with service unavailable. Continuing to retry for 4m56s

Attempt #2 failed with service unavailable. Continuing to retry for 4m55s

Attempt #3 failed with service unavailable. Continuing to retry for 4m53s

Attempt #4 failed with service unavailable. Continuing to retry for 4m49s

Attempt #5 failed with service unavailable. Continuing to retry for 4m41s

Attempt #6 failed with service unavailable. Continuing to retry for 4m25s

Attempt #7 failed with service unavailable. Continuing to retry for 3m55s

Attempt #8 failed with service unavailable. Continuing to retry for 3m25s

Attempt #9 failed with service unavailable. Continuing to retry for 2m55s

Attempt #10 failed with service unavailable. Continuing to retry for 2m25s

Attempt #11 failed with service unavailable. Continuing to retry for 1m55s

Attempt #12 failed with service unavailable. Continuing to retry for 1m25s

Solved$20 Bounty

Pinned Solution

benpoieszhv
PROOP

a month ago

well, I figured it out so I'll leave it here for posterity.

Our health check endpoint has been '/' for the lifetime of the project and it's always returned a 303 since our service requires login.

However, starting earlier today the health check requires 200 as a response, so changing just the route to an endpoint that returns 200 instead instantly passes.

The errors in the railway health check need to indicate if a response is received, but with the wrong code vs just saying 'unavailable'...

11 Replies

Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 1 month ago


Does it start if the healthcheck is removed?


benpoieszhv
PROOP

a month ago

I haven't tried doing that yet as that too takes a PR to strike it -- but the 502 gateway error while still having an active 'good' deployment on staging has been what I've been investigating. My assumption has been that the server is fine, but nothing can resolve to it...

image.png

image.png


Was there any change made to the network configuration in the service settings, like changing the target port, or modifying the PORT variable of your service? Also, removing the healthcheck doesn't require a new PR, you just have to remove it from service settings, and hit apply changes, which will redeploy the service.


This error is caused when Railway's edge routes the traffic to your service, but your app doesn't receive it, because it has an incorrect port/host binding in most cases. Also, healthchecks require the PORT env variable to be set, so it knows which port to route the request to. Hopefully, this helps you in narrowing down what the issue is.


benpoieszhv
PROOP

a month ago

looking at the config, none of the network settings have changed in over 6 months... it is set to the correct port.

our config is set with a railway file, the web UI will not let you modify it in that case it needs to be a new code change.


benpoieszhv
PROOP

a month ago

I've disabled the health check and the system does come up fine now.

It doesn't really answer why the health checks are failing when the server is actually up. Any ideas or things do debug?


That means the app is binding correctly. I assume the PORT variable is set to the same port your app is listening on?


benpoieszhv
PROOP

a month ago

We've never had the PORT variable needing to be set (we use 8080), i tried setting per your suggestion while debugging but it made no difference either way


benpoieszhv
PROOP

a month ago

currently the only change made to the server was disabling the health check -- which has worked, and the route set for the health check works fine after being deployed...


benpoieszhv
PROOP

a month ago

well, I figured it out so I'll leave it here for posterity.

Our health check endpoint has been '/' for the lifetime of the project and it's always returned a 303 since our service requires login.

However, starting earlier today the health check requires 200 as a response, so changing just the route to an endpoint that returns 200 instead instantly passes.

The errors in the railway health check need to indicate if a response is received, but with the wrong code vs just saying 'unavailable'...


benpoieszhv
PROOP

a month ago

Thank you for your suggestions darseen!


Status changed to Solved dev • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...