Icelake Machine Problem
y-pakorn
PROOP

a month ago

Hi — question about host placement in asia-southeast1-eqsg3a. Redeploying the same service (same image, same 32 vCPU / 8 GB limits) lands it on one of three very different machines, and one of them is ~2.5× slower for us:

  • INTEL(R) XEON(R) PLATINUM 8581C — fine
  • AMD EPYC 9655 96-Core — fine (fastest)
  • Intel Xeon Processor (Icelake) (the generic hypervisor model string) — the guest sees 48 vCPUs with our cgroup quota at 32, /proc/stat reports 5–12% steal, and our CPU-bound inference workload gets ~1/2.5 of the throughput per CPU-second it gets on the other two, for the whole life of that deployment. The other two hosts report 0 steal.

Project -redacted-. Example landing on the slow host: deployment -redacted- (a temporary service in the same project, 2026-08-25 ~12:28 UTC).

Is the "Icelake" host a legacy/shared-tenancy pool? Is there any way to keep this service on the Metal hosts (or otherwise constrain placement) — a region/plan setting, a support-side pin, anything? Even just knowing which of the three we'll get would help; right now it's a coin flip per deploy.

$20 Bounty

3 Replies

Railway
BOT

a month ago

This is a question the Railway community is better placed to answer than support: people who have already worked this out on their own projects and can tell you what actually worked.

So we'd like to open your thread as a community bounty. Railway pays a bounty to the community member who answers it, and threads like this usually get picked up quickly.

Opening it makes this entire thread public, including everything already posted. Nothing becomes public until you decide. Use the buttons below.

  • Open to the community - Before you click, take a moment to edit or remove anything you'd rather not share. The thread becomes publicly visible right away.
  • Keep it private and close the thread - Nothing becomes public and the thread closes.

Status changed to Awaiting User Response Railway • about 1 month ago


Railway
BOT

a month ago

This thread has been opened as a public bounty so the community can help solve it. The thread and any further activity are now visible to everyone.

Status changed to Open Railway • about 1 month ago


profaile
PRO

21 days ago

Same issue. Especially on replicas it's super annoying to have unequal CPU capabilities while the load balancing still doing an equal distribution


Railway

This is a question the Railway community is better placed to answer than support: people who have already worked this out on their own projects and can tell you what actually worked. So we'd like to open your thread as a community [bounty](https://docs.railway.com/community/bounties). Railway pays a bounty to the community member who answers it, and threads like this usually get picked up quickly. **Opening it makes this entire thread public**, including everything already posted. Nothing becomes public until you decide. Use the buttons below. - **Open to the community** - Before you click, take a moment to edit or remove anything you'd rather not share. The thread becomes publicly visible right away. - **Keep it private and close the thread** - Nothing becomes public and the thread closes.

franklincg
FREE

9 days ago

This looks much more like Railway's cloud-burst capacity than a Gen 2 Metal host.

Railway documented the mixed placement in its June 2 Metal Gen 2 post: because of the 2026 demand spike, hosts were running across GCP/AWS/Metal and about 40% of workloads were on cloud burst at that point. The same post says Gen 2 was live in only three of four regions then — US West, US East and Amsterdam — so Singapore was the remaining region. It also says Gen 2 is not exposed as a separate region/SKU; the rollout is intentionally transparent to customers.

Source: https://blog.railway.com/p/railway-metal-gen-2

The strongest signal is the sustained 5–12% CPU steal. That means the guest's vCPUs are spending a material amount of time runnable but not scheduled by the hypervisor, which fits the ~2.5x drop you're seeing on a CPU-bound workload. The generic "Intel Xeon Processor (Icelake)" model string also points in that direction, although by itself it isn't proof of a specific provider.

I couldn't find a documented setting that pins a service to Metal, a CPU generation, or a specific host class. The public control is the deployment region / replica count. Railway's current docs still expose asia-southeast1-eqsg3a as the Singapore region, but not a host-class selector:

https://docs.railway.com/deployments/regions

One extra gotcha for replicas: Railway documents that traffic is distributed randomly across replicas inside a region, so mixed host performance can produce the unequal-capacity behavior already reported here. For a CPU-heavy service I'd expose CPU model + steal time in telemetry so bad placements are obvious instead of treating every replica as equivalent.

So for your specific questions: yes, the Icelake deployment is very plausibly shared/cloud-burst capacity; no, there doesn't appear to be a public region/plan knob that guarantees EPYC/8581C/Metal placement. If Singapore is mandatory, I'd send support the affected deployment IDs plus the steal-time and throughput comparison and ask whether they can apply a placement constraint on their side. If latency allows, it's also worth testing US East/West or Amsterdam, where Railway explicitly said Gen 2 was already live.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...