Frontend service down 24+ hrs after hardware failure notice — deployments show "successful" but site 502s
gigasouk
PROOP

19 days ago

Description of the issue:

We received a Service Notice email ~24 hours ago stating a server running our "frontend" service

(project: Gigasouk, environment: production) experienced a hardware failure, with an expected

resolution "within 1 hour." It has now been down for approximately 24 hours.

Current state:

  • gigasouk.in and gigasouk.com/manufacturer return a 502 Bad Gateway (Cloudflare shows the error

    originating from the Host, not Cloudflare).

  • The Railway dashboard shows the latest deployment as "Deployment successful" and the service

    status as "Active" / "Online" for both the frontend and backend services.

  • This mismatch (deploy succeeded, service reported online, but the live site is unreachable) has

    persisted since the outage email.

Error messages:

502 Bad Gateway (Cloudflare error page), Host: Error — as shown in Cloudflare's diagnostic panel.

What I need:

Please check whether our frontend service is still running on the failed hardware mentioned in the

outage notice, and manually reassign/restart it on a healthy host if so. Third-party status monitors

show Railway as operational platform-wide, so this appears isolated to our specific service/host.

Project: Gigasouk

Service: frontend

Environment: production

Domain: gigasouk.in (custom domain also mapped: gigasouk.com)

Region: US West

Solved$20 Bounty

5 Replies

Railway
BOT

19 days ago

The hardware issue from yesterday has been resolved, but the frontend service did not recover with it. The latest deploy's Next.js build completed without errors, but the application never started responding on its port, so no container is currently serving traffic. A fresh deploy will bring the service back: open the frontend service, press Cmd+K (Ctrl+K on Windows/Linux) to open the command palette, and run "Deploy latest commit." This will not affect any of your data.

Additionally, HTTP requests to gigasouk.in are being redirected to gigasouk.com before reaching us, which is blocking TLS certificate issuance for that domain. You will need to check your DNS or redirect configuration so that gigasouk.in resolves to Railway directly instead of redirecting.


Status changed to Awaiting User Response Railway • 19 days ago


gigasouk
PROOP

18 days ago

Update on the frontend outage (service: frontend, project: Gigasouk):

We've done extensive debugging and can rule out an application-level cause:

  • Reproduced the exact production build (Next.js 16.3.1) locally — server starts and responds

    correctly (HTTP 200) to requests on the same port within milliseconds.

  • Added diagnostic logging to our startup script: the child process starts, logs "Ready," and stays

    alive (confirmed via exitCode checks) — it is not crashing.

  • A self-check fetch to the app's own port from within the same container consistently returns

    ECONNREFUSED, on every deploy, despite the app reporting "Ready."

  • Ruled out port mismatch (Railway target port 8080 matches app config exactly).

  • Ruled out DNS/IPv6 resolution (tested with explicit 127.0.0.1, same failure).

  • Ruled out healthcheck timeout (failed across a full 2-minute retry window, 6 attempts).

  • Ruled out memory (24GB/24vCPU allocated, far above what a Next.js frontend needs).

Additional evidence from the Metrics tab: CPU and Memory both show flat 0 usage for the entire

deployment window, despite our container-level logs showing the process started, printed "Ready,"

and stayed alive per our own diagnostics above. This means the process is not actually running or

attached to the host from Railway's own resource accounting, even though it appears healthy from

inside the container. The Requests panel for the same window shows 9 total requests, all 5xx —

consistent with your edge reaching the host but failing to reach our container.

Given this started immediately after the hardware failure notice we received, a redeploy did not

resolve it, and there's now a direct contradiction between our container-level logs and your

host-level metrics, we suspect this specific service/replica is either still running on degraded

infrastructure or is in a broken state left over from the migration. Could you investigate at the

host/container level, or force-migrate this service to a fresh host?


Status changed to Awaiting Railway Response Railway • 18 days ago


Railway
BOT

18 days ago

Having looked into this, the issue appears to be in your application code or configuration rather than the Railway platform itself, which puts it outside what Railway support can resolve directly.

This is exactly the kind of problem the Railway community is good at, so we'd like to open your thread as a community bounty. Railway pays a bounty to the community member who solves it, and threads like this usually get picked up quickly.

Opening it makes this entire thread public, including everything already posted. Nothing becomes public until you decide. Use the buttons below.

  • Open to the community - Before you click, take a moment to edit or remove anything you'd rather not share. The thread becomes publicly visible right away.
  • Keep it private and close the thread - Nothing becomes public and the thread closes.

Status changed to Awaiting User Response Railway • 18 days ago


gigasouk
PROOP

18 days ago

We've now tested with your platform's actual production environment variables (via railway run,

pulling the real Variables from this service — including SUPABASE_SERVICE_ROLE_KEY and

RESEND_API_KEY) run locally against the identical production build. The server starts and serves

a fully working page with live data in milliseconds — no errors, no hang, no missing config.

This directly rules out an application code or configuration issue: the exact same code, with the

exact same production secrets, works everywhere except inside a container running on your

infrastructure. Given this, we don't believe closing this as a code/config issue is accurate, and would

like it escalated for a proper infrastructure-level look rather than redirected to the community.

Happy to share the reproduction steps or logs again if useful. This is a live production outage for a

paying (Pro plan) customer that has now lasted over 24 hours.


Status changed to Awaiting Railway Response Railway • 18 days ago


Railway
BOT

18 days ago

This still looks like an application-level problem, so Railway support can't take it further, but the community can. The buttons below are still live: open the thread up as a public bounty after editing out anything sensitive, or keep it private and close it.


Status changed to Awaiting User Response Railway • 18 days ago


Railway
BOT

18 days ago

This thread has been opened as a public bounty so the community can help solve it. The thread and any further activity are now visible to everyone.

Status changed to Open Railway • 18 days ago


Status changed to Solved gigasouk • 17 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...