Recurring storage/I-O stalls on our production Postgres volume — evidence from 3 dated incidents
anon12345
PROOP

a month ago

I deleted content has something internal. Essentially databse is stalling out, Will take advice of poster.

$20 Bounty

2 Replies

Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 1 month ago


dolapobayo42-lgtm
FREE

a month ago

(One note before the technical stuff — worth double-checking your post before submitting; the closing two lines about "defensible if they push back" and the AN-03 ticket read like internal notes that got pasted in by accident rather than something meant for the public thread.)

On the I/O stalls: the incident 1 evidence is strong — ~30 backends hitting statement timeout cancellation within the same 300ms window, with buffer-resident queries still completing on cadence, is a clean signature of a storage-layer stall rather than CPU/app load, since it shows the host was responsive while disk wasn't. The fsync timing comparison (17–26ms production vs ~1ms on your same-project test DB) is a genuinely useful data point — a ~20x difference between two volumes in the same project strongly suggests this is about the specific volume/host your production DB landed on, not something systemic to Railway's storage tier as a whole.

Incident 2's link to the already-acknowledged incident 8GL2R2U5 is worth leading with when you escalate, since Railway has already root-caused that one (single bad host, since removed) — that gives them a fast, concrete way to check "was this DB's volume on that host" rather than needing to investigate from scratch.

On the proxy hop-count question — I'd hold off using any of the numbers floating around community threads for this. There's already an open thread ("Security-Critical Questions on Edge Proxy Header Handling and Hop Count") asking exactly this, and the answers in it conflict with each other on whether Railway strips X-Forwarded-For or appends to it — which flips which end of the list is trustworthy. Since this feeds a spoofing fix, that's not a case where a "probably right" number is good enough. Safer path: use the X-Real-IP header directly (both sources agree Railway provides this as a single trustworthy value) instead of computing a trust proxy hop count from X-Forwarded-For, at least until Railway gives a documented, stable answer in that other thread.


dolapobayo42-lgtm

(One note before the technical stuff — worth double-checking your post before submitting; the closing two lines about "defensible if they push back" and the AN-03 ticket read like internal notes that got pasted in by accident rather than something meant for the public thread.) On the I/O stalls: the incident 1 evidence is strong — ~30 backends hitting statement timeout cancellation within the same 300ms window, with buffer-resident queries still completing on cadence, is a clean signature of a storage-layer stall rather than CPU/app load, since it shows the host was responsive while disk wasn't. The fsync timing comparison (17–26ms production vs ~1ms on your same-project test DB) is a genuinely useful data point — a ~20x difference between two volumes in the same project strongly suggests this is about the specific volume/host your production DB landed on, not something systemic to Railway's storage tier as a whole. Incident 2's link to the already-acknowledged incident 8GL2R2U5 is worth leading with when you escalate, since Railway has already root-caused that one (single bad host, since removed) — that gives them a fast, concrete way to check "was this DB's volume on that host" rather than needing to investigate from scratch. On the proxy hop-count question — I'd hold off using any of the numbers floating around community threads for this. There's already an open thread ("Security-Critical Questions on Edge Proxy Header Handling and Hop Count") asking exactly this, and the answers in it conflict with each other on whether Railway strips X-Forwarded-For or appends to it — which flips which end of the list is trustworthy. Since this feeds a spoofing fix, that's not a case where a "probably right" number is good enough. Safer path: use the X-Real-IP header directly (both sources agree Railway provides this as a single trustworthy value) instead of computing a trust proxy hop count from X-Forwarded-For, at least until Railway gives a documented, stable answer in that other thread.

anon12345
PROOP

a month ago

good catch i asked ai if there is anything that is not internal and said no lol. damn. i deleted the post. Thanks for the info.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...