9 days ago
Hi Railway team/community,
I’m seeing repeated SSH/SFTP access failures for a production service while the deployment itself remains healthy and shows SUCCESS.
Observed behavior:
-
SSH access to ssh.railway.com was initially able to connect and authenticate successfully.
-
Later SSH attempts began failing intermittently with errors such as:
Connection closed by 66.33.22.3 port 22
and:
ssh_dispatch_run_fatal: Connection to 66.33.22.3 port 22: Connection timed out
- I also tried Railway’s read-only service files command to inspect an ephemeral /tmp path:
railway service files ... list /tmp --json
The CLI reported:
Failed to connect to Railway SFTP at ssh.railway.com
with Windows socket error:
os error 10060
- The production deployment continues to show SUCCESS throughout these failures.
I have stopped further SSH/SFTP retries to avoid making the situation worse.
Could a Railway team member please clarify:
- Are Railway SSH and
railway service filesusing the same SSH/SFTP relay infrastructure or failure domain? - Can a production service remain healthy while the SSH/SFTP access plane is unavailable?
- Is there currently any known issue affecting ssh.railway.com or the SFTP/service-files access path?
- Is there any provider-supported read-only method that does not use SSH/SFTP for checking the ephemeral filesystem of a running service?
- If this requires any restart, redeploy, or replica recycle, please tell me first. I do not want any state-changing action performed before I understand the impact.
Important:
The service is currently running successfully, so I do not want to restart or redeploy it just for troubleshooting.
If a Railway staff member needs project/service/deployment identifiers for investigation, I can provide them privately rather than posting them publicly.
Thanks.
6 Replies
9 days ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 9 days ago
9 days ago
Thanks — this is helpful and consistent with what I observed.
For now I’m keeping the service untouched because I still need the current ephemeral /tmp state preserved, so I will not restart or redeploy it.
Before I resume any connectivity tests or SSH/SFTP access attempts, I’d like confirmation from a Railway staff member on two points:
-
Does
railway service filesofficially share the same ssh.railway.com SSH/SFTP relay failure domain asrailway ssh? -
Is there any provider-side way for Railway staff to inspect or restore access to the current running replica without restarting, redeploying, or recycling it?
I’m intentionally not running further SSH/SFTP retries until this is clarified.
If Railway staff needs the project/service/deployment identifiers, I can provide them through a private channel.
9 days ago
railway service files uses SFTP.
Also, have you tried using a VPN or a different network to establish SSH/SFTP connection?
0x5b62656e5d
`railway service files` uses SFTP. Also, have you tried using a VPN or a different network to establish SSH/SFTP connection?
9 days ago
Thanks for confirming that railway service files uses SFTP.
I have not tried a VPN or a different network yet.
At the moment I have intentionally paused all SSH/SFTP retries because I need to preserve the current running replica and its ephemeral /tmp state.
Before I resume any connectivity test from another network, could you please confirm:
-
Would trying the same read-only SSH/SFTP access from a VPN or different network be considered non-disruptive to the currently running replica?
-
If access still fails from a different network, is there any Railway-side way to restore the SSH/SFTP access path without restarting, redeploying, or recycling the current replica?
I do not want any action that could replace the current container or wipe its ephemeral /tmp state.
9 days ago
Trying to connect via SSH/SFTP won't cause the replica to randomly restart, unless your code crashes somehow.
0x5b62656e5d
Trying to connect via SSH/SFTP won't cause the replica to randomly restart, unless your code crashes somehow.
9 days ago
Thanks for the earlier guidance. I have a new data point that changes the diagnosis.
I repeated the access-path test from an independent mobile hotspot rather than the original network.
On the alternate mobile hotspot:
-
TCP/22 to ssh.railway.com succeeded.
-
railway service files ... list /tmp --jsonsucceeded. -
A strict-pinned SSH read-only gate succeeded against the exact production service/environment/deployment.
-
The existing executor was read back successfully with its expected SHA-256, byte length, and mode.
-
The ephemeral LOADER preparation also succeeded: the chunk namespace was created, the loader chunk transferred successfully, the loader was assembled successfully, and the same-command final readback verified:
bytes: 2047
SHA-256:
810c91ac46f985c18f3c4466dad1f9b9af9dde8813bab4b48fcd686033fcfb39
mode:
0400
inode:
351010830
However, the immediately following independent SSH read-only proof failed after authentication with:
Connection closed by 66.33.22.3 port 22
This means the failure is no longer isolated to the original network path.
The same alternate mobile-hotspot path successfully handled TCP/22, Railway service-files/SFTP, multiple strict-pinned SSH sessions, and the loader preparation, but then a subsequent independent SSH session was closed by the Railway SSH endpoint.
The service itself remains healthy.
We have intentionally stopped all further access attempts.
We have NOT:
- restarted the service
- redeployed
- recycled the replica
- changed variables
- changed the start command
- changed the volume attachment
- changed the SSH key
- changed known_hosts
- deleted or modified the prepared LOADER
- written to the production /data volume
Could you please check the Railway-side access path for this service, especially:
-
ssh.railway.com relay / gateway logs around the failed session.
-
Any gateway/session/rate-limit behavior that could explain successful SSH/SFTP sessions followed by an immediate session close.
-
Any service-specific routing or session-affinity issue for this service or running replica.
-
Whether
railway service filesand direct SSH share any relay/session/rate-limit layer that could produce this pattern. -
Whether Railway exposes any non-SSH/SFTP read-only method to inspect a running service's ephemeral
/tmpfilesystem. -
Whether the SSH/SFTP access path can be restored without restarting, redeploying, recycling the replica, or mutating the attached production volume.
Important constraint:
Please do not trigger or recommend a restart, redeploy, or replica recycle as an implicit diagnostic step without explicitly calling it out first.
We currently have ephemeral forensic state in /tmp that must be preserved.
Current exact blocker:
REQUIRED_STRICT_PINNED_SSH_CHANNEL_CLOSED_DURING_LOADER_FINAL_INDEPENDENT_PROOF_ON_SELECTED_ALTERNATE_NETWORK
Observed endpoint message:
Connection closed by 66.33.22.3 port 22
Thank you.
freddy-tan
Thanks for the earlier guidance. I have a new data point that changes the diagnosis. I repeated the access-path test from an independent mobile hotspot rather than the original network. On the alternate mobile hotspot: - TCP/22 to ssh.railway.com succeeded. - `railway service files ... list /tmp --json` succeeded. - A strict-pinned SSH read-only gate succeeded against the exact production service/environment/deployment. - The existing executor was read back successfully with its expected SHA-256, byte length, and mode. - The ephemeral LOADER preparation also succeeded: the chunk namespace was created, the loader chunk transferred successfully, the loader was assembled successfully, and the same-command final readback verified: bytes: 2047 SHA-256: 810c91ac46f985c18f3c4466dad1f9b9af9dde8813bab4b48fcd686033fcfb39 mode: 0400 inode: 351010830 However, the immediately following independent SSH read-only proof failed after authentication with: Connection closed by 66.33.22.3 port 22 This means the failure is no longer isolated to the original network path. The same alternate mobile-hotspot path successfully handled TCP/22, Railway service-files/SFTP, multiple strict-pinned SSH sessions, and the loader preparation, but then a subsequent independent SSH session was closed by the Railway SSH endpoint. The service itself remains healthy. We have intentionally stopped all further access attempts. We have NOT: - restarted the service - redeployed - recycled the replica - changed variables - changed the start command - changed the volume attachment - changed the SSH key - changed known_hosts - deleted or modified the prepared LOADER - written to the production /data volume Could you please check the Railway-side access path for this service, especially: 1. ssh.railway.com relay / gateway logs around the failed session. 2. Any gateway/session/rate-limit behavior that could explain successful SSH/SFTP sessions followed by an immediate session close. 3. Any service-specific routing or session-affinity issue for this service or running replica. 4. Whether `railway service files` and direct SSH share any relay/session/rate-limit layer that could produce this pattern. 5. Whether Railway exposes any non-SSH/SFTP read-only method to inspect a running service's ephemeral `/tmp` filesystem. 6. Whether the SSH/SFTP access path can be restored without restarting, redeploying, recycling the replica, or mutating the attached production volume. Important constraint: Please do not trigger or recommend a restart, redeploy, or replica recycle as an implicit diagnostic step without explicitly calling it out first. We currently have ephemeral forensic state in `/tmp` that must be preserved. Current exact blocker: REQUIRED_STRICT_PINNED_SSH_CHANNEL_CLOSED_DURING_LOADER_FINAL_INDEPENDENT_PROOF_ON_SELECTED_ALTERNATE_NETWORK Observed endpoint message: Connection closed by 66.33.22.3 port 22 Thank you.
8 days ago
Following up after 24 hours on the updated alternate-network evidence above.
The issue remains blocked at the same point:
REQUIRED_STRICT_PINNED_SSH_CHANNEL_CLOSED_DURING_LOADER_FINAL_INDEPENDENT_PROOF_ON_SELECTED_ALTERNATE_NETWORK
We have kept the running replica and all current ephemeral /tmp state untouched since the failure.
No further SSH/SFTP attempts, restart, redeploy, replica recycle, volume mutation, or /data write have been performed.
Could a Railway staff member please review the SSH relay/gateway/session side of this case, or advise the next non-disruptive diagnostic step that preserves the current replica and /tmp state?
Thank you.