Database hanging randomly
lain67
PROOP

a month ago

Out of nowhere database connection starts hanging and we hit request timeout errors. This is happening hours after the redeployment with non-predictive patterns and no redeployment or restart seems to be helping it. This happens trough our backend service of which we can't even reach /health. We suspect that its internal networking problem since the domains itself become unreachable. We aren't sure wether the problem is itself in the database or where the backend is hosted but after weeks of security audits, the problem isn't in our source code.

image.png

As i have mentioned this happened several times in last few months and it then gets "fixed" automatically after hours of downtime

image.png

*5xx errors are time periods where this problem ocures

$20 Bounty

4 Replies

Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 1 month ago


ericknextlink-cmd
PRO

a month ago

Can you provide a link to the deployed project or is it purely internal and no outbound access?

Your database is it also hosted on Railway or you use supabase?

How were you using your database locally, did you spin a postgres with docker locally to use or a shared deployed remote db?

Have you happened to enable the "Under Attack Mode" DDoS protection on your backend service, if yes then likely disable it might be the source, faced this same problem

For networking unless your railway generated url does not work and you use a domain then it might be a DNS configuration, also you might need to check the port you are enabling for your server in code and the one for railway. Typically node backends railway uses 8080 and if you use something like 4000 it might break, its not a usual thing but sometimes happens.

I doubt this will be an issue but also check the route exports, I had a backend where locally using /health was fine but then the exported routes were under /api/v1 which made it that the deployed endpoint was /api/v1/health and not /health, I am very sure this might not be the case on your end. But you could explore these to see what works and what doesn't. Glad to had some feedback to follow up.

If you are using php laravel then thats another thing but let me know what it is then lets see


ericknextlink-cmd

Can you provide a link to the deployed project or is it purely internal and no outbound access? Your database is it also hosted on Railway or you use supabase? How were you using your database locally, did you spin a postgres with docker locally to use or a shared deployed remote db? Have you happened to enable the "Under Attack Mode" DDoS protection on your backend service, if yes then likely disable it might be the source, faced this same problem For networking unless your railway generated url does not work and you use a domain then it might be a DNS configuration, also you might need to check the port you are enabling for your server in code and the one for railway. Typically node backends railway uses 8080 and if you use something like 4000 it might break, its not a usual thing but sometimes happens. I doubt this will be an issue but also check the route exports, I had a backend where locally using /health was fine but then the exported routes were under /api/v1 which made it that the deployed endpoint was /api/v1/health and not /health, I am very sure this might not be the case on your end. But you could explore these to see what works and what doesn't. Glad to had some feedback to follow up. If you are using php laravel then thats another thing but let me know what it is then lets see

lain67
PROOP

a month ago

  • database is hosted on railway
  • deployed project has outbound access but during such indicidents we are unable to communicate to it
  • no problem when running in local environment, with our without docker
  • UAM DDoS proection isn't enabled
  • neither does railway generated domain nor custom domain works
  • port is correctly set up
  • it's not just health, its almost every route
  • it's rust not php

and main headache is that this problem occurs randomly, for hours and it can happen almost a day or half a day after a previous deployment


ericknextlink-cmd
PRO

a month ago

can you share the url and currently is it working or its down again?

and you mentioned its after a previous deployment is it always after a previous deployment or there are times that it happens randomly as well

Can you connect to the backend using Railway's internal/private networking during the incident?

Can you reproduce the problem from another Railway service

During an incident, have a tiny throwaway Railway service continuously run something like:

every 5 seconds:

resolve backend hostname

TCP connect to backend

HTTP GET /health/live

record latency/error

Do the same for Postgres.

This gives you evidence instead of relying on whether your own client can reach the service.

During the next incident, can you run a test from another Railway service in the same environment that simultaneously checks DNS resolution,TCP connectivity to the backend, HTTP /health/live and TCP/Postgres connectivity to the database

if possible, please capture the results every few seconds until the incident recovers.


lain67
PROOP

a month ago

  • it's currently working, it's api.poly.inc
  • it happens randomly, it hasnt started from the previous deployment, what i mean is that it happens hours after the previous deployment during that time, so it's not caused by the deployment itself or the code
  • i am unable to connect to railway's internal/private networking during the indicdent neither
  • haven't noticed anything similar on other deployed services on the railway

I'm gonna capture the evidence when the next incident occurs and report it back, thanks for the help


Welcome!

Sign in to your Railway account to join the conversation.

Loading...