worker ram usage at idle is stuck at high usage after scaling down.
sbiefeni
PROOP

7 days ago

I have a service (image worker) which has not changed since July 27. It has always scaled down after a storm to ~40mb.

today I was testing cost efficiency of scaling to more replicas during an ingest - ie: less compute more time, vs more compute less time.

this worker usually scaled to 6 replicas. today I tested 12, then 24 replicas. Typically the ram usage after scaling back down to 1 would sit about 40mb. Now after 24 replicas, scaled back down to 1, the ram usage is sitting at a constant 1.5 - 2 GB!!

once again the cost is ticking. redeploying the image worker also did NOT fix the issue.

similar to what happened to me a few weeks ago, stuck invisible processes running. Please fix.

Solved

8 Replies

Railway
BOT

7 days ago

We've looked into this from our side and haven't found anything on the Railway platform that explains what you're seeing, so working it out means digging into your specific setup.

That's exactly what the Railway community is good at, so we'd like to open your thread as a community bounty. Railway pays a bounty to the community member who solves it, and threads like this usually get picked up quickly.

Opening it makes this entire thread public, including everything already posted. Nothing becomes public until you decide. Use the buttons below.

  • Open to the community - Before you click, take a moment to edit or remove anything you'd rather not share. The thread becomes publicly visible right away.
  • Keep it private and close the thread - Nothing becomes public and the thread closes.

Status changed to Awaiting User Response Railway 7 days ago


sbiefeni
PROOP

7 days ago

please look again. the code in that service deploy has not changed in more than 2 weeks. the only difference is today I was testing scaling it to more replicas during an ingest. something is stuck.

NOTE that this EXACT thing happened to me some time back, to which I got a similar response.. and eventually someone ELSE from railway looked at it and said yes, the problem is on the railway side.. and I was refunded some usage.

I suggest checking this again.


Status changed to Awaiting Railway Response Railway 7 days ago


Railway
BOT

7 days ago

We still haven't found anything on the Railway side behind this, so the community is the best next step. The buttons below are still live: open the thread up as a public bounty after editing out anything sensitive, or keep it private and close it.


Status changed to Awaiting User Response Railway 7 days ago


Railway
BOT

7 days ago

This thread has been opened as a public bounty so the community can help solve it. The thread and any further activity are now visible to everyone.

Status changed to Open Railway 7 days ago


sbiefeni
PROOP

7 days ago

I just realized I was talking to a bot too. Can I get a support person to look at this?


sbiefeni
PROOP

7 days ago

Thanks mr-h. is this a known issue on railway? happened to me once before, they insisted it wasn't their fault until someone looked at it and said oops, it was our fault..


sbiefeni
PROOP

7 days ago

Also, I did check both metrics - the project metrics, and the service worker metrics. the service metrics still shows all the replicas, even though it scaled back to 1. after some time at least, but it's been several hours, the service metrics still shows replicas, and the project metrics still shows very high RAM usage for that worker..


sbiefeni
PROOP

7 days ago

no. and I restarted it, no effect, and then redeployed it.. also no effect. that service normally sits about 40mb when scaled to 1 and idle. but after testing 24 instead of 6 replicas during ingest, now it stuck at high ram..


Status changed to Awaiting Railway Response Railway 7 days ago


7 days ago

I checked this at the container level rather than through the metrics dashboard.

There is one container running for the image worker right now, sitting at about 20 MB. That matches the single replica you're scaled to. Nothing else is running against that service.


Status changed to Awaiting User Response Railway 7 days ago


Railway
BOT

7 hours ago

This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!

Status changed to Solved Railway about 7 hours ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...