21 days ago
-
Service:
tourpilot-workers· Environment: production · Region: eu-west · Replicas: 1 -
Runtime: Node.js (Express), long-running single process
-
Question: during the windows below, was the container for this service **paused, live-migrated, or
CPU-throttled at host level**? If platform/host events exist for these windows, could you share them?
What we observe
The service normally answers availability requests in well under a second (median ~0.6 s). Occasionally it
stops making progress for ~90–100 seconds and then resumes, completing all in-flight requests successfully
in a single burst. Nothing in the application fails: no exceptions, no retries, no error responses.
Occurrences (UTC)
| # | Window (UTC) | Deployment | What we have |
|---|---|---|---|
| 1 | 2026-08-10, ~11:22 | aacd27e1 (short id) | Identified retrospectively: one request timed by the app at 47,396 ms (a second one at 19,531 ms), both completed 200 OK. Partial data only — our slow-request warning was not yet in place, and no metrics were captured for this window. |
| 2 | 2026-08-15, 15:10:22 → 15:11:49 (~87 s) | 9c77808b-1192-4ee8-b850-459fae246ef5 | 43 requests timed by the app at 27.6–64.6 s, all 200 OK; 43 HTTP 499 in the Railway HTTP log at 24.9–25.0 s. |
| 3 | 2026-09-13, 11:11:10 → 11:12:53 (~102 s) | 2464ef3e (short id) | 2 requests timed by the app at 101,990 ms and 38,532 ms, both 200 OK, both completing at the same millisecond (11:12:52.899); 2 HTTP 499 at 24.7–25.0 s. |
Signature (occurrences 2 and 3)
-
App-measured durations far above normal, all successful. The durations come from the application's
own clock (start/end deltas inside the process), not from log ingestion times. Work was not failing — it
stopped and resumed.
-
HTTP 499 at ~25 s in the Railway HTTP log. Our upstream proxy has a 25-second timeout: it cancels the
request and falls back to a secondary backend. The 499 is the client closing, not a server error.
-
Nothing else in the window: zero application errors, zero stale-cache fallbacks, zero circuit-breaker
openings, zero upstream retries, zero 5xx responses.
-
Flat resource metrics. 13/9, from
railway metrics --rawat 30 s resolution: CPU ≤ 0.022 vCPU(limit 8), memory 0.375 → 0.390 GB (limit 8 GB), network within its normal range, no gaps in the series.
15/8: CPU and memory flat on the dashboard. We understand that metric sampling continues while a process is
paused, so this rules out load and memory pressure but does not prove a pause by itself. Flat CPU also
argues against a CPU-bound event-loop block on our side.
Supporting hint (not conclusive)
On 2026-08-15 the 43 warnings written as each request completed have ingestion timestamps within 10 ms
(15:11:52.342 → 15:11:52.352), although the operations started over a 37-second span — as if stdout was
flushed once on resume. This could also be an artifact of log delivery, so we do not rely on it.
Ruled out on our side
-
A restart: the same deployment kept running through each window (boot timestamp unchanged, in-flight
requests completed rather than dropped).
-
CPU load, memory pressure or GC (metrics flat).
-
Upstream API slowness: that failure mode has a different signature for us (retries, errors, stale-cache
fallbacks), and none of it appears in these windows.
-
The Railway status page was clean for 2026-08-15; the "logs/metrics" incident of 2026-08-14 16:22 UTC does
not cover that window.
Impact
No data loss so far: the proxy fallback covered all three occurrences. Each stall does degrade availability
responses to a third-party marketplace for up to ~100 s, and the pattern has now repeated three times, which is
why we are asking.
What would help
-
Any host or container events (pause/freeze, live migration, CPU throttling, noisy neighbour, host
maintenance) for the three windows above.
-
Whether a single-replica service in eu-west can be paused or migrated without a restart, and whether such
events are visible to us anywhere (logs, metrics, activity feed).
We can provide redacted extracts of the application and HTTP logs for these windows on request.
1 Replies
21 days ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 21 days ago
20 days ago
This strongly indicates a platform-level container freeze/suspension or host migration, not application load; ask Railway to inspect host/container events