shifttracker-backend: recurring europe-west4-drams3a edge issue, possible regression of 2026-06-09 fix
classebasse
PROOP

4 days ago

Follow-up to our earlier private ticket (station.railway.com/support/intermittent-silent-response-loss-on-shi-6abb3a12), which was auto-closed as "Solved" after 21 days of inactivity without ever being reviewed by an engineer - the two replies we got were both the automated assistant offering a community bounty, twice, despite us explaining why that wasn't the right fit here.

WHY WE'RE FILING THIS PUBLICLY THIS TIME

While researching this we found station.railway.com/questions/intermittent-522-connection-timed-out-556e8e65 - a very close match (same edge region, same "healthy container, edge just doesn't respond" signature). It was resolved by Railway engineers directly in that public thread (not a paid community bounty) once given exact request-ID / X-Railway-Edge evidence, and multiple independent customers corroborating the same symptom in the same window helped get real traction. One Railway engineer (noahd) said on that thread that a fix had "recently" been rolled out around 2026-06-09. We think we may be looking at a regression of the same underlying edge-layer issue, about 3 months later, and would appreciate someone from the infra/networking team checking internal telemetry for europe-west4-drams3a at the timestamps below - this isn't something community members can diagnose, since it requires access to Railway's own edge logs, but we're posting publicly in case other affected customers see this and corroborate (as happened in the referenced thread).

SERVICE DETAILS

Project: Shifttracker (fb059ffd-d5fe-4898-8d23-dcd63db78a0d)

Service: shifttracker-backend, production (744bcf9e-0439-4147-a488-a5565698b3da), environment ef7214fd-2ca8-4af7-86af-b83a16ab8bb0

Public domain: api.shifttracker.se / shifttracker-backend-production.up.railway.app

Every traced incident has edgeRegion: europe-west4-drams3a. Backend container shows zero restarts/redeploys across all investigated windows - this is not an app crash.

SYMPTOM

Our mobile app (React Native / OkHttp on Android) intermittently gets a generic network error with no response at all, which self-heals on retry. We've ruled out our own client-side connection pooling as the sole cause (shipped a native OkHttp keep-alive tuning fix in app v1.3.0 on 2026-09-01 - failure rate afterward was unchanged), and ruled out device/platform specifics (reproduces on different Android hardware, and in a controlled synthetic test over both WiFi and LTE). Two distinct failure signatures show up in our own server-side HTTP logs:

Signature A - response lost in transit: our server receives the request and returns 200, but the client never sees it, retries, and gets a duplicate 200 about 0.6-0.7s later. Example: POST /shift/entry, first attempt 2026-09-02 06:56:02.868 UTC -> 200 in 80ms (requestId eEueigrrRw-U5SLHs_GTAg), retry 2026-09-02 06:56:03.543 UTC -> 200 in 65ms (requestId qnRpbQKvTSeq8FIZs_GTAg). We confirmed via our own idempotency keys that only one row was ever persisted, so no data corruption - but the first response clearly never reached the client. 7 confirmed instances of this pattern so far, all slower POST calls (74-85ms server-side): 2026-09-01 13:56:58, 18:39:03, 19:19:49, 19:34:25, 19:40:20 UTC; 2026-09-02 06:56:01 and 07:42:29 UTC.

Signature B - request never arrives: no corresponding entry appears anywhere in our server-side HTTP logs for the failed attempt - the request appears to die before ever reaching your edge (the same absent-request-id signal that was the key diagnostic clue in the referenced thread). One confirmed instance: GET /tasks/last-logged, 2026-09-02 05:38:10 UTC.

ONGOING RECURRENCE

This is not a one-off: our client-side logs show a steady background rate of 5-9 affected requests/week since mid-August, with a spike to 81 and 63 in the weeks of 2026-08-31 and 2026-09-07 respectively, and it is still happening as recently as 2026-09-30 07:41 UTC.

Happy to share further request IDs, exact timestamps, or anything else if useful - we have a full log of every incident since mid-August.

Thanks!

$20 Bounty

4 Replies

Railway
BOT

4 days ago

This thread has been marked as private.

Status changed to Awaiting Railway Response Railway • 4 days ago


Railway
BOT

4 days ago

We can't reach the shifttracker-backend service or its environment from the account this thread is tied to, so we can't look at its edge traffic or request logs from here.

Those resources sit under a different account from the one you're writing from. The request needs to come from the account that owns the Shifttracker project before its HTTP logs and the request IDs you listed can be checked against that service.


Status changed to Awaiting User Response Railway • 4 days ago


classebasse
PROOP

3 days ago

The account posting this thread (classebasse) does have access to this service - it's exactly the one selected from this ticket's own Service picker dropdown when filing it (Shifttracker / shifttracker-backend / production), which only lists services the account can see, so access isn't the issue.

To clarify further: our account has two separate Railway workspaces - a Hobby plan ("classesbasse's projects") and the Pro plan workspace "Classes Grus och Gnejs AB", where Shifttracker actually lives. If your support tooling defaults to one workspace context rather than the workspace the selected service belongs to, that would explain why it can't find it. Could you check whether this ticket's lookup is scoped to the right workspace?

Also noting this thread appears to have been automatically converted to private just now, even though we selected "Ask the Community" when creating it - could that be restored, since the public visibility was intentional?


Status changed to Awaiting Railway Response Railway • 3 days ago


Railway
BOT

3 days ago

We haven't established what's causing this yet. The community can help investigate your setup.

That's exactly what the Railway community is good at, so we'd like to open your thread as a community bounty. Railway pays a bounty to the community member who solves it, and threads like this usually get picked up quickly.

Opening it makes this entire thread public, including everything already posted. Nothing becomes public until you decide. Use the buttons below.

  • Open to the community - Before you click, take a moment to edit or remove anything you'd rather not share. The thread becomes publicly visible right away.
  • Keep it private and close the thread - Nothing becomes public and the thread closes.

Status changed to Awaiting User Response Railway • 3 days ago


Railway
BOT

3 days ago

This thread has been opened as a public bounty so the community can help solve it. The thread and any further activity are now visible to everyone.

Status changed to Open Railway • 3 days ago


classebasse
PROOP

3 days ago

Opened to the community as suggested, though to be clear: the open question here isn't something community members can diagnose (it requires Railway's internal edge telemetry for europe-west4-drams3a), so we'd still appreciate an engineer taking a look, same as happened in the referenced precedent thread.

Also, the previous bot reply didn't address the workspace-scoping question from our last message: our account has two Railway workspaces (Hobby "classesbasse's projects" and Pro "Classes Grus och Gnejs AB"), and Shifttracker lives under the Pro one. The service was selectable from this ticket's own Service picker, so the account clearly has access - if your lookup tool defaults to the wrong workspace, that would explain the earlier "can't reach it" response.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...