2 months ago
My MySQL service has been unreachable since today's US West incident (RL8FRJE6). The status page now shows "Fully Operational", but my database never recovered, so I don't think this project came back with the rest of the region.
PROJECT DETAILS
Project: sincere-expression (c361a60d-b21a-4aa7-adb7-64175498f54e)
Environment: production (0f8441b6-6e77-4350-8f36-155bbedf45e9)
MySQL service: 2ca0cd14-7071-4981-926a-807cf9b29b24 (mysql:9.4, US West), deployment 374a0749
Volume: mysql-volume -> vol_4pi6943khquovy5v
App service: 7fe4e22c-0160-4dad-be1f-4d25b40723b5
TIMELINE (times GMT+8)
- 11:33 - InnoDB started reporting long semaphore waits. Multiple threads blocked 305-886 seconds on the same LOG_WRITER mutex:
--Thread 139943866226240 has waited at log0write.cc line 971 for 886 seconds the semaphore:
Mutex at 0x7f47b7c2e6c0, Mutex LOG_WRITER created log0log.cc:625, locked by 139943867283008
This looks like log writes to the volume stopped completing.
- 11:34:38 - InnoDB's long-semaphore-wait watchdog aborted the server:
InnoDB: about forcing recovery.
2026-08-10T03:34:38Z UTC - mysqld got signal 6 ;
Thread pointer: 0x0
BuildID[sha1]=c4e0cd59af2a4ab7eb21afe2a4c2c6c7e745a603
-
12:20:29 - stack trace written, followed by the standard MySQL crash notice.
-
12:22:14 - container restarted and logged:
Mounting volume on: /var/lib/containers/railwayapp/bind-mounts/a228d6ae-2d8b-41a5-ac14-690c0b289cbd/vol_4pi6943khquovy5v
After this line there is NO further log output at all. mysqld never prints its usual startup or InnoDB recovery messages, so it appears to block immediately after mounting the volume.
CURRENT SYMPTOMS
- Service shows "Online" in the dashboard, but no MySQL output since 12:22:14.
-
- Railway's own Database > Data tab cannot connect either: "We are unable to connect to the database via SSH - Connection lost: The server closed the connection."
-
- My app connects over private networking (mysql.railway.internal:3306) and every query fails with ETIMEDOUT. Requests consistently take ~10.2s, matching the mysql2 client's 10s connect timeout, i.e. nothing is listening on 3306.
-
- I restarted the APP service once (no effect, as expected). I have deliberately NOT restarted the MySQL service, to avoid interrupting any InnoDB crash recovery in progress.
WHAT I THINK IS WRONG
Volume vol_4pi6943khquovy5v appears to be unresponsive at the storage layer. The original crash was InnoDB protecting itself after writes stopped completing, and the restarted container now hangs at exactly the same point - right after mounting that volume.
WHAT I'M ASKING FOR
-
Could you check the health of volume vol_4pi6943khquovy5v and its underlying host? I suspect it is still hung even though the regional incident is marked resolved.
-
If the volume is healthy and the data needs InnoDB recovery, please advise the safest path. I would rather not set innodb_force_recovery blindly.
-
IMPORTANT: this is a production system for a small manufacturing company and there are no backups (backups/PITR require the Pro plan). Please prioritise data preservation over restoring service quickly - please do not reinitialise or wipe the volume.
Happy to provide more logs or run diagnostics. Thanks for your help.
5 Replies
Status changed to Awaiting Railway Response Railway • about 2 months ago
2 months ago
I am facing the same issue, did you find a resolution?
2 months ago
Apologies for the trouble. Your MySQL service was affected by the connectivity incident we had in US West earlier today, which has since been resolved. You can see the full timeline here: https://status.railway.com/incident/RL8FRJE6
Your volume is attached and in READY state in our systems. We have looked into the storage layer for you and triggered a fresh redeploy of your MySQL service. Your volume data is preserved through this process. Please check whether MySQL comes back up and let us know if the issue persists.
Status changed to Awaiting User Response Railway • about 2 months ago
2 months ago
Thanks Noah, I appreciate you looking into the storage layer. Unfortunately the service is still down, and from my side it doesn't look like the redeploy actually took effect.
What I'm seeing right now:
-
No new deployment was created. The MySQL service still shows the same deployment 374a0749 as ACTIVE, originally created 2026-05-25. There is nothing newer in the deployment history.
-
No new log output at all. The last line in Deploy Logs is still: "2026-08-10 12:22:14 Mounting volume on: /var/lib/containers/railwayapp/bind-mounts/a228d6ae-2d8b-41a5-ac14-690c0b289cbd/vol_4pi6943khquovy5v". That is now over 16 hours old, and mysqld has never printed a single startup or InnoDB recovery line since.
-
Still nothing listening on 3306. My app on the private network keeps failing with ETIMEDOUT, and every request takes almost exactly 10.2s (mysql2's default 10s connect timeout). I just re-ran it three times: 10179ms, 10276ms, 10175ms.
-
Railway's own Database > Data tab still cannot connect either.
So the container still appears to be blocked immediately after mounting the volume, exactly as before. If the volume is attached and READY on your side, could you please check whether the container is actually stuck on I/O against it, and confirm the redeploy you triggered was really applied? It may have failed to schedule.
If a plain redeploy can't get past this, either of these would help:
- mount the volume on a temporary service so I can pull a mysqldump off it, or
-
- advise the safest innodb_force_recovery level to get mysqld up far enough to dump the data.
As mentioned there are no backups on this project, so I'd rather take the slow, safe route than risk the data. Thanks again.
Status changed to Awaiting Railway Response Railway • about 2 months ago
2 months ago
Apologies for the trouble. Your MySQL service was affected by the connectivity incident we had in US West earlier today, which has since been resolved. You can see the full timeline here: https://status.railway.com/incident/RL8FRJE6
We triggered a redeploy of your database and your volume is attached and in READY state in our systems. Your volume data is preserved through this process. Please let us know if you still have any issues.
Status changed to Awaiting User Response Railway • about 2 months ago
sam-a
Apologies for the trouble. Your MySQL service was affected by the connectivity incident we had in US West earlier today, which has since been resolved. You can see the full timeline here: https://status.railway.com/incident/RL8FRJE6 We triggered a redeploy of your database and your volume is attached and in READY state in our systems. Your volume data is preserved through this process. Please let us know if you still have any issues.
2 months ago
Confirmed — MySQL is back up and all data is intact. Thanks for the help!
Status changed to Awaiting Railway Response Railway • about 2 months ago
Status changed to Solved nickwinwinwin • about 2 months ago