MySQL service unreachable after infra maintenance — production down, no backups
asoranz
PROOP

2 months ago

MySQL service unreachable after infra maintenance — production down, no backups

Project: deixa-comigo (e7460b50-ddb9-4d70-a071-44e4b4f53596)

Service (MySQL): 204f1955-ad75-4676-ac8f-e5842e7a4a44

Region: us-west2 Volume: mysql-volume

Maintenance event ID: 47d47cbd-6180-412e-a448-edc53522e1ae (active since 22:53 UTC)

Stuck deployment ID: 8169c971-92e5-4ad0-9437-8ab15daa025e

Stacker: production-stacker-cssv9a-metal-052-02 (RPC timeouts)

The MySQL service shows Online but is unreachable from every path:

  • App service in the same project, over the private network: connect ETIMEDOUT

  • railway connect MySQL --tunnel-only: "channel 1: open failed: unknown channel

    type: unsupported", then "Connection closed by remote host"

  • Dashboard > Database tab: "We are unable to connect to the database via SSH —

    Connection lost: The server closed the connection."

Container logs show no output since 2026-08-18 03:28 UTC (no crash, no restart).

A redeploy triggered at 23:08 UTC has been stuck at "Deploy > Creating containers"

for over 15 minutes, producing no container logs.

Impact: production is DOWN. All users are locked out, because authentication

reads the database. This volume has NO backups (PITR is Pro-only).

Please either bring the stacker back online or migrate the service to healthy

infrastructure. Happy to accept a volume migration if that is the faster path. Additional impact: three other production services of mine are also FAILED

right now — Lex Social / Backend, Criapost / criapost-backend, and

Memovoz / MemoVoz. This looks account-wide or stacker-wide, not limited to

one service.

Related past incident with identical symptoms: RL8FRJE6 (thread

"MySQL volume unresponsive after US West incident") — in that case only

Railway staff could recover it, by triggering a redeploy after checking the

storage layer. My own redeploys have no effect, same as reported there.

Solved

7 Replies

Status changed to Awaiting Railway Response Railway • about 2 months ago


asoranz
PROOP

2 months ago

Update: I just upgraded this workspace to the Pro plan. Service is still stuck

at "Creating containers" 28+ minutes after the redeploy. Production remains down.


asoranz
PROOP

2 months ago

Update: the redeploy I triggered has now FAILED after 31 minutes, never

producing a container or any logs. Service state is "Deploy failed"; the

database remains unreachable from the app (connect ETIMEDOUT), from the CLI

tunnel, and from the dashboard's own Database tab.

I also upgraded this workspace to the Pro plan a few minutes ago.

I am not going to trigger further redeploys — per incident RL8FRJE6, only

staff-side action recovered it. Please advise or act from your side.


dizzydes90
EMPLOYEE

2 months ago

Unfortunately we're having an incident with the host you're project is currently on and fixing it is going to take about 1-2 hours.

Please accept my apologies we have several platform engineers investigating this right now. Will update here as I know more.


Status changed to Awaiting User Response Railway • about 2 months ago


asoranz
PROOP

2 months ago

Thank you that's helpful. Standing by, not touching the service.

Please ping this thread when the host is recovered so I can verify the

database comes back and restart my app service.


Status changed to Awaiting Railway Response Railway • about 2 months ago


Quick update we have began to host recovery process what happened was the host that your machine was on was unhealthy and we had to perform a migration. The infrastructure team has now begin this process and will be giving you updates as it goes along.


Status changed to Awaiting User Response Railway • about 2 months ago


dizzydes90
EMPLOYEE

a month ago

Good news: your MySQL is back up. It redeployed successfully at 01:26 UTC and the volume shows your data in place (about 305 MB), and your app service came back right after at 01:28. After any host recovery it's worth giving your most recent data a quick check to confirm everything's where you expect it.

On the other services you mentioned: those failures predate this incident. Criapost's last deploy failed on July 25, and the Lex Social and MemoVoz deployments were removed months ago, so they weren't part of tonight's issue.

Now that you're on Pro, enabling backups on this volume from its settings will cover you if anything like this happens again.


asoranz
PROOP

a month ago

Confirmed recovered on my side, thank you both.

Data check done as suggested everything intact: 14,119 sales, 678 clients,

18,115 chat messages, and the most recent sale timestamped 22:18:16 UTC-3,

minutes before the outage. SHOW ENGINE INNODB STATUS shows no corruption

signals. Nothing was lost.

Thanks for the correction on the other services you're right, those

failures predate this incident and I misread them as related.

Backups are now enabled on the volume, plus a daily off-platform logical dump

on my side. Closing this out.


Status changed to Awaiting Railway Response Railway • about 2 months ago


Status changed to Solved asoranz • about 2 months ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...