Recurring ~17-minute execution stalls on Primary (n8n, queue mode) — EU West, project pleasant-prosperity
izamora-hub
PROOP

a month ago

Project: pleasant-prosperity

Project ID: cfc0d89a-ced1-429f-8216-bdfa111211a2

Environment: production (9aefda70-4c2d-4c35-bffa-d3facd4d118e)

Service: Primary (service ID b41b5eda-1090-4923-a41c-183fca898310), n8n in queue mode with a separate Worker service

Region: EU West

Current deployment: 957a6281-84a5-4958-aae1-c94d28108d8a

Since yesterday (2026-08-20) we've had four separate n8n workflow executions on this Worker service get stuck for almost exactly ~17 minutes before completing on their own (they all eventually succeed — no data loss — but the delay is user-facing, since these executions send WhatsApp replies to end users):

  • 2026-08-20 16:20:37 UTC — 17m 9.077s
  • 2026-08-21 06:26:02 UTC — 17m 7.464s
  • 2026-08-21 07:17:32 UTC — 17m 8.386s
  • 2026-08-21 08:16:01 UTC — 17m 0.0s

The near-identical duration across four independent occurrences (different triggers: one right after a manual redeploy of Primary, one right after reimporting/saving a workflow with no redeploy involved, one with no obvious trigger at all) suggests a fixed-interval recovery/timeout mechanism somewhere in the platform rather than anything in our own workflow logic. Normal executions on this same service complete in 2-10 seconds.

We also saw this at container boot today (2026-08-21 06:36:42 UTC, deploy 957a6281): "Cluster check warning Detected 2 instances claiming leader ro9664425f097, deff4006-bf52-41e3-8818-fd0fc1bd09d8"

— only once, not sustained, but it coincides with the start of this pattern.

We're aware of the resolved incident at status.railway.com/incident/VVL3A03V (Google Cloud infra issue, deploy pipeline congestion, resolved 2026-08-20 18:20 UTC) — but our stalls continued well after that was markk it's the same root cause, or the underlying issue wasn'tfully cleared for our project/region.

Could you check EU West infra health for this project around these four timestamps, and let us know if there's a known issue (queue/cluster leader election, worker-primary coordination, etc.) causing this ~17-

Happy to provide execution ID 3668 on the Worker

Solved$10 Bounty

0 Replies

Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 2 months ago


Status changed to Solved izamora-hub • about 2 months ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...