Replica served production traffic for five hours while producing no logs (service 0a02ddfc)
imchikachirag
HOBBYOP

a month ago

Hello,

One of our replicas served production traffic for over five hours while producing no deploy logs at all. It has since been superseded by a new deployment, which logs normally. We would like to understand what happened, because nothing we changed explains it and nothing prevents it recurring.

IDENTIFIERS

• Project: 49e739bd-587c-4d35-8be7-2c985e1018c8

• Environment: 1629afb7-63f3-4f8c-aa85-6424ab8dc47c

• Service: 0a02ddfc-5913-4e84-961f-dc8888af5ca8

• AFFECTED: deployment d5c5c8db-0f37-4d91-afe4-0c1b7119deed, replica 0833e6df, started 2026-08-24T12:04:39Z

• WORKING, BEFORE: deployment f7df7006-7525-4ffe-a146-80c9cbe4772a, replica c13eee8e

• WORKING, AFTER: deployment 8c264509-cd25-4de7-bcf4-d2348152a793, replica 15d242b3, started 2026-08-24T13:42:04Z

WHAT WE SAW

• The affected deployment shows exactly one line for its entire life: "Starting Container" at 12:04:39Z. Nothing else, across more than five hours.

• The deployment before it has a complete unbroken log from 08:24:51Z to a clean "Stopping Container" at 12:04:45Z, including a keep-alive line every 4 minutes and a database latency summary every 5 minutes.

• The deployment after it logged all four of our startup lines within 0.63 seconds of container start.

• So the log store retains and renders correctly. Only that one replica was affected.

THE APPLICATION WAS RUNNING NORMALLY THROUGHOUT

• Your healthcheck on /api/health passed at 12:04:40Z.

• At 12:41Z it served a 2.2 GB HTTP download to completion and committed the resulting database writes correctly.

• It continued answering requests for the rest of its life.

WHAT WE ELIMINATED BEFORE CONTACTING YOU

• Not a filter, time range, or wrong deployment. The surrounding deployments render correctly in the same view, and your dashboard marked the affected one Active.

• Not transient. A restart produced a second container from the same image, equally silent.

• Not stdout-only. We deliberately triggered application errors at approximately 13:24Z that write to stderr. They returned HTTP 500, so our error handler definitely executed. Those lines did not appear either. Both streams were affected.

• Not our code, and this is the strongest evidence we have: the deployment that restored logging changed only client-side files. No server file changed between the silent deployment and the working one. The same server code was silent on one container and logged normally on the next.

• Not a dependency change. Both lockfiles are tracked in git and unchanged across all three deployments.

• Start command is a plain "node index.js" with no redirection. Our railway.json is unchanged.

WHAT WE ARE ASKING

• Whether log ingestion for replica 0833e6df failed at container start, and whether that output was dropped or is retained but not being served to us.

• Whether anything on your side can cause a single replica to be excluded from log collection while continuing to receive routed traffic.

• Whether there is any way to detect that collection has failed, rather than discovering hours later that a window is empty.

CONTEXT

We had a different incident on 17 August where roughly 90 to 95 percent of our output was dropped over several hours before recovering on its own, leaving survivors at a regular cadence. This one was total from the moment of boot, so we are reporting it as a separate problem rather than a recurrence.

The affected replica has been superseded, so there is no live container to inspect. If anyone has seen a single replica drop out of log collection while still receiving traffic, we would like to hear how it resolved. Railway team: the retained record should still show this from your side.

Thank you.

Solved

3 Replies

Status changed to Awaiting Railway Response Railway • about 1 month ago


sam-a
EMPLOYEE

a month ago

Your findings are correct. We confirmed that log ingestion failed for the affected replica from the moment it started. Our own systems show zero runtime log entries for that deployment, so the output was dropped, not retained somewhere inaccessible. There is currently no mechanism to detect that log collection has stopped for a running replica while traffic continues to route to it.


Status changed to Awaiting User Response Railway • about 1 month ago


Status changed to Solved sam-a • about 1 month ago


imchikachirag
HOBBYOP

a month ago

Excellent response. You confirmed the fault, confirmed the output was dropped rather than retained, and were straight about there being no detection mechanism, which is more than most vendors would say. The one thing that would help us most is exactly that gap: any signal at all that log collection has stopped for a running replica. We lost five hours to it because a silent log and a healthy system are indistinguishable from our side. Even a metric or a status flag we could poll would turn this from hours into a minute.


Status changed to Awaiting Railway Response Railway • about 1 month ago


a month ago

That feedback is noted and well-framed. A log-collection health signal visible to the customer is not something we expose today, and your thread here is the right place for it to be tracked.


Status changed to Awaiting User Response Railway • about 1 month ago


Railway
BOT

a month ago

This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!

Status changed to Solved Railway • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...