13 days ago
Hi Railway team,
We run a video rendering service on Railway (Pro plan) and we are hitting the per-replica process limit, not CPU or memory. We would like to understand how that limit is set and whether it can be raised.
Current layout: 3 replicas
What the container reports:
/sys/fs/cgroup/pids.max = 1000
/sys/fs/cgroup/cpu.max = 3200000 100000 (i.e. a 32 vCPU quota)
nproc = 48 (host cores visible through the container)
The workload is an ffmpeg-based video pipeline. A single reel is assembled from many short segments, so one job spawns a large number of short-lived ffmpeg processes and their threads rather than a few long-running ones.
What we observed:
With 4 concurrent jobs per replica we hit the 1000-PID ceiling and ffmpeg started failing with exit code 245 (EAGAIN on fork).
We worked around it by cutting to 2 concurrent jobs per replica and adding a third replica. With that layout, 8 concurrent jobs peaked at 704 of 1000 PIDs and completed successfully.
The side effect is that we now use roughly 15 of the 96 vCPU we are paying for across the three replicas. The process ceiling, not the CPU quota, is what caps our throughput, so we are paying for cores we cannot reach.
Our questions:
Is pids.max=1000 a fixed platform value, or is it derived from the service's vCPU/memory allocation? If it scales, what does it scale with?
Can it be raised for our service — for example to 8000 or 16000 PIDs per replica? We are not asking for more CPU or memory, only for the ability to use the allocation we already have.
Our container reports a 32 vCPU cgroup quota, while the Pro plan page states up to 24 vCPU per replica. Which value is authoritative for billing and for scheduling?
If the PID limit cannot be raised on Pro, what is the supported path — a higher plan, a different service configuration, or something else?
Happy to provide logs, metric exports, or a reproduction if that helps.
1 Replies
13 days ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 13 days ago
13 days ago
PID limits won’t be raised.
For the 24/32 discrepancy, legacy Pro workspaces have 32 vCPU limits, while new subscribers now have 24.
I’m not sure how your service is set up but it’s possible that you could split the video into parts then delegate the job to several replicas, or upgrade to an enterprise plan.