Recurring ~100 s process stalls, eu-west, flat metrics — host pause/migration?
1clickcomputers
HOBBYOP

21 days ago

  • Service: tourpilot-workers · Environment: production · Region: eu-west · Replicas: 1

  • Runtime: Node.js (Express), long-running single process

  • Question: during the windows below, was the container for this service **paused, live-migrated, or

    CPU-throttled at host level**? If platform/host events exist for these windows, could you share them?

What we observe

The service normally answers availability requests in well under a second (median ~0.6 s). Occasionally it

stops making progress for ~90–100 seconds and then resumes, completing all in-flight requests successfully

in a single burst. Nothing in the application fails: no exceptions, no retries, no error responses.

Occurrences (UTC)

| # | Window (UTC) | Deployment | What we have |

|---|---|---|---|

| 1 | 2026-08-10, ~11:22 | aacd27e1 (short id) | Identified retrospectively: one request timed by the app at 47,396 ms (a second one at 19,531 ms), both completed 200 OK. Partial data only — our slow-request warning was not yet in place, and no metrics were captured for this window. |

| 2 | 2026-08-15, 15:10:22 → 15:11:49 (~87 s) | 9c77808b-1192-4ee8-b850-459fae246ef5 | 43 requests timed by the app at 27.6–64.6 s, all 200 OK; 43 HTTP 499 in the Railway HTTP log at 24.9–25.0 s. |

| 3 | 2026-09-13, 11:11:10 → 11:12:53 (~102 s) | 2464ef3e (short id) | 2 requests timed by the app at 101,990 ms and 38,532 ms, both 200 OK, both completing at the same millisecond (11:12:52.899); 2 HTTP 499 at 24.7–25.0 s. |

Signature (occurrences 2 and 3)

  1. App-measured durations far above normal, all successful. The durations come from the application's

    own clock (start/end deltas inside the process), not from log ingestion times. Work was not failing — it

    stopped and resumed.

  2. HTTP 499 at ~25 s in the Railway HTTP log. Our upstream proxy has a 25-second timeout: it cancels the

    request and falls back to a secondary backend. The 499 is the client closing, not a server error.

  3. Nothing else in the window: zero application errors, zero stale-cache fallbacks, zero circuit-breaker

    openings, zero upstream retries, zero 5xx responses.

  4. Flat resource metrics. 13/9, from railway metrics --raw at 30 s resolution: CPU ≤ 0.022 vCPU

    (limit 8), memory 0.375 → 0.390 GB (limit 8 GB), network within its normal range, no gaps in the series.

    15/8: CPU and memory flat on the dashboard. We understand that metric sampling continues while a process is

    paused, so this rules out load and memory pressure but does not prove a pause by itself. Flat CPU also

    argues against a CPU-bound event-loop block on our side.

Supporting hint (not conclusive)

On 2026-08-15 the 43 warnings written as each request completed have ingestion timestamps within 10 ms

(15:11:52.342 → 15:11:52.352), although the operations started over a 37-second span — as if stdout was

flushed once on resume. This could also be an artifact of log delivery, so we do not rely on it.

Ruled out on our side

  • A restart: the same deployment kept running through each window (boot timestamp unchanged, in-flight

    requests completed rather than dropped).

  • CPU load, memory pressure or GC (metrics flat).

  • Upstream API slowness: that failure mode has a different signature for us (retries, errors, stale-cache

    fallbacks), and none of it appears in these windows.

  • The Railway status page was clean for 2026-08-15; the "logs/metrics" incident of 2026-08-14 16:22 UTC does

    not cover that window.

Impact

No data loss so far: the proxy fallback covered all three occurrences. Each stall does degrade availability

responses to a third-party marketplace for up to ~100 s, and the pattern has now repeated three times, which is

why we are asking.

What would help

  • Any host or container events (pause/freeze, live migration, CPU throttling, noisy neighbour, host

    maintenance) for the three windows above.

  • Whether a single-replica service in eu-west can be paused or migrated without a restart, and whether such

    events are visible to us anywhere (logs, metrics, activity feed).

We can provide redacted extracts of the application and HTTP logs for these windows on request.

$10 Bounty

1 Replies

Railway
BOT

21 days ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • 21 days ago


townifiedapp-dot
FREE

20 days ago

This strongly indicates a platform-level container freeze/suspension or host migration, not application load; ask Railway to inspect host/container events


Welcome!

Sign in to your Railway account to join the conversation.

Loading...