Replica scale-up stopped starting extra containers (desired count updates, still one box)
sbiefeni
PROOP

a month ago

The problem (plain English)

We run several workers on this project (image processing, web, face detection). When a lot of photos are uploaded, we ask Railway to run more copies of those workers (for example 1 → 24 for image-worker). When the work is done, we ask it to go back to one copy.

This used to work. On the Metrics tab I could see extra replicas appear, and RAM/CPU went up. Extra boxes were actually running.

It does not work now. We still ask for 24. Railway saves 24 (Settings / API). Only one container stays running. Metrics still shows a single replica. RAM stays at one small process (~70 MB for image-worker). Scale-down and restart of that one box still look fine — we never had extra boxes to remove.

We have not changed how we ask for 24 since 14 August 2026. That method was working after we switched to it. Same services, same region (sfo), same kind of upload storm. Please look at why extra replicas are no longer being started.

───

Technical details

Project: 5be56683-2ad3-4519-8774-78750b12d5eb (production)

Environment ID: b98a0f8d-6eab-4a13-a712-51c1c7be1213

Region: sfo

Services

• image-worker 1e4198b2-eea9-45d5-a781-ca60d8ba21d6

• web 59c607f3-b945-4270-b978-45cc07c1a395

• detect-worker 55670fa1-58f5-480a-88a3-5a267b52f7b5

Scale-up call (no redeploy):

mutation($serviceId: String!, $environmentId: String!, $input: ServiceInstanceUpdateInput!) {

serviceInstanceUpdate(serviceId: $serviceId, environmentId: $environmentId, input: $input)

}

input: { numReplicas: N }

This storm: image-worker 1 → 24, web 1 → 4, detect-worker 1 → 6. Mutation returns OK.

On 14 Aug 2026 we switched to this official numReplicas field (no multiRegionConfig, no serviceInstanceRedeploy on scale). After that, Metrics showed extra replica series. We still idle-redeploy image/web/detect at 1 replica after a storm (RAM flush only). We did not change the 1→24 mutation after 14 Aug.

Now (26 Aug 2026, two storms, ~22:20–23:20 UTC):

• serviceInstance.numReplicas = 24 / 4 / 6

• Running instances = 1

• Latest deployment serviceManifest.deploy.numReplicas and multiRegionConfig.sfo still 1

• image-worker ~70 MB / ~0.6 vCPU (one PHP process)

• detect-worker ~0.7 GB (one InsightFace child, not six)

Please check

  1. Why updating serviceInstance.numReplicas is not scheduling more instances.
  2. Whether a redeploy at 1 replica pins the running deployment so later live numReplicas updates are ignored.
  3. Whether we should also set region replicas (sfo).
  4. Any platform change after ~14 Aug that would make config-only replica updates stop applying.

NOTE: I selected image worker, but it's all workers that have the same issue.

Solved$20 Bounty

Pinned Solution

numReplicas is deprecated. Use the multiRegionConfig field instead.

2 Replies

Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 1 month ago


numReplicas is deprecated. Use the multiRegionConfig field instead.


sbiefeni
PROOP

a month ago

super. testing it out now..


Status changed to Solved 0x5b62656e5d • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...