SSH/SFTP access-plane failures on healthy production service
freddy-tan
FREEOP

9 days ago

Hi Railway team/community,

I’m seeing repeated SSH/SFTP access failures for a production service while the deployment itself remains healthy and shows SUCCESS.

Observed behavior:

  1. SSH access to ssh.railway.com was initially able to connect and authenticate successfully.

  2. Later SSH attempts began failing intermittently with errors such as:

Connection closed by 66.33.22.3 port 22

and:

ssh_dispatch_run_fatal: Connection to 66.33.22.3 port 22: Connection timed out

  1. I also tried Railway’s read-only service files command to inspect an ephemeral /tmp path:

railway service files ... list /tmp --json

The CLI reported:

Failed to connect to Railway SFTP at ssh.railway.com

with Windows socket error:

os error 10060

  1. The production deployment continues to show SUCCESS throughout these failures.

I have stopped further SSH/SFTP retries to avoid making the situation worse.

Could a Railway team member please clarify:

  • Are Railway SSH and railway service files using the same SSH/SFTP relay infrastructure or failure domain?
  • Can a production service remain healthy while the SSH/SFTP access plane is unavailable?
  • Is there currently any known issue affecting ssh.railway.com or the SFTP/service-files access path?
  • Is there any provider-supported read-only method that does not use SSH/SFTP for checking the ephemeral filesystem of a running service?
  • If this requires any restart, redeploy, or replica recycle, please tell me first. I do not want any state-changing action performed before I understand the impact.

Important:

The service is currently running successfully, so I do not want to restart or redeploy it just for troubleshooting.

If a Railway staff member needs project/service/deployment identifiers for investigation, I can provide them privately rather than posting them publicly.

Thanks.

$10 Bounty

6 Replies

Railway
BOT

9 days ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • 9 days ago


freddy-tan
FREEOP

9 days ago

Thanks — this is helpful and consistent with what I observed.

For now I’m keeping the service untouched because I still need the current ephemeral /tmp state preserved, so I will not restart or redeploy it.

Before I resume any connectivity tests or SSH/SFTP access attempts, I’d like confirmation from a Railway staff member on two points:

  1. Does railway service files officially share the same ssh.railway.com SSH/SFTP relay failure domain as railway ssh?

  2. Is there any provider-side way for Railway staff to inspect or restore access to the current running replica without restarting, redeploying, or recycling it?

I’m intentionally not running further SSH/SFTP retries until this is clarified.

If Railway staff needs the project/service/deployment identifiers, I can provide them through a private channel.


railway service files uses SFTP.

Also, have you tried using a VPN or a different network to establish SSH/SFTP connection?


0x5b62656e5d

`railway service files` uses SFTP. Also, have you tried using a VPN or a different network to establish SSH/SFTP connection?

freddy-tan
FREEOP

9 days ago

Thanks for confirming that railway service files uses SFTP.

I have not tried a VPN or a different network yet.

At the moment I have intentionally paused all SSH/SFTP retries because I need to preserve the current running replica and its ephemeral /tmp state.

Before I resume any connectivity test from another network, could you please confirm:

  1. Would trying the same read-only SSH/SFTP access from a VPN or different network be considered non-disruptive to the currently running replica?

  2. If access still fails from a different network, is there any Railway-side way to restore the SSH/SFTP access path without restarting, redeploying, or recycling the current replica?

I do not want any action that could replace the current container or wipe its ephemeral /tmp state.


Trying to connect via SSH/SFTP won't cause the replica to randomly restart, unless your code crashes somehow.


0x5b62656e5d

Trying to connect via SSH/SFTP won't cause the replica to randomly restart, unless your code crashes somehow.

freddy-tan
FREEOP

9 days ago

Thanks for the earlier guidance. I have a new data point that changes the diagnosis.

I repeated the access-path test from an independent mobile hotspot rather than the original network.

On the alternate mobile hotspot:

  • TCP/22 to ssh.railway.com succeeded.

  • railway service files ... list /tmp --json succeeded.

  • A strict-pinned SSH read-only gate succeeded against the exact production service/environment/deployment.

  • The existing executor was read back successfully with its expected SHA-256, byte length, and mode.

  • The ephemeral LOADER preparation also succeeded: the chunk namespace was created, the loader chunk transferred successfully, the loader was assembled successfully, and the same-command final readback verified:

    bytes: 2047

    SHA-256:

    810c91ac46f985c18f3c4466dad1f9b9af9dde8813bab4b48fcd686033fcfb39

    mode:

    0400

    inode:

    351010830

However, the immediately following independent SSH read-only proof failed after authentication with:

Connection closed by 66.33.22.3 port 22

This means the failure is no longer isolated to the original network path.

The same alternate mobile-hotspot path successfully handled TCP/22, Railway service-files/SFTP, multiple strict-pinned SSH sessions, and the loader preparation, but then a subsequent independent SSH session was closed by the Railway SSH endpoint.

The service itself remains healthy.

We have intentionally stopped all further access attempts.

We have NOT:

  • restarted the service
  • redeployed
  • recycled the replica
  • changed variables
  • changed the start command
  • changed the volume attachment
  • changed the SSH key
  • changed known_hosts
  • deleted or modified the prepared LOADER
  • written to the production /data volume

Could you please check the Railway-side access path for this service, especially:

  1. ssh.railway.com relay / gateway logs around the failed session.

  2. Any gateway/session/rate-limit behavior that could explain successful SSH/SFTP sessions followed by an immediate session close.

  3. Any service-specific routing or session-affinity issue for this service or running replica.

  4. Whether railway service files and direct SSH share any relay/session/rate-limit layer that could produce this pattern.

  5. Whether Railway exposes any non-SSH/SFTP read-only method to inspect a running service's ephemeral /tmp filesystem.

  6. Whether the SSH/SFTP access path can be restored without restarting, redeploying, recycling the replica, or mutating the attached production volume.

Important constraint:

Please do not trigger or recommend a restart, redeploy, or replica recycle as an implicit diagnostic step without explicitly calling it out first.

We currently have ephemeral forensic state in /tmp that must be preserved.

Current exact blocker:

REQUIRED_STRICT_PINNED_SSH_CHANNEL_CLOSED_DURING_LOADER_FINAL_INDEPENDENT_PROOF_ON_SELECTED_ALTERNATE_NETWORK

Observed endpoint message:

Connection closed by 66.33.22.3 port 22

Thank you.


freddy-tan

Thanks for the earlier guidance. I have a new data point that changes the diagnosis. I repeated the access-path test from an independent mobile hotspot rather than the original network. On the alternate mobile hotspot: - TCP/22 to ssh.railway.com succeeded. - `railway service files ... list /tmp --json` succeeded. - A strict-pinned SSH read-only gate succeeded against the exact production service/environment/deployment. - The existing executor was read back successfully with its expected SHA-256, byte length, and mode. - The ephemeral LOADER preparation also succeeded: the chunk namespace was created, the loader chunk transferred successfully, the loader was assembled successfully, and the same-command final readback verified: bytes: 2047 SHA-256: 810c91ac46f985c18f3c4466dad1f9b9af9dde8813bab4b48fcd686033fcfb39 mode: 0400 inode: 351010830 However, the immediately following independent SSH read-only proof failed after authentication with: Connection closed by 66.33.22.3 port 22 This means the failure is no longer isolated to the original network path. The same alternate mobile-hotspot path successfully handled TCP/22, Railway service-files/SFTP, multiple strict-pinned SSH sessions, and the loader preparation, but then a subsequent independent SSH session was closed by the Railway SSH endpoint. The service itself remains healthy. We have intentionally stopped all further access attempts. We have NOT: - restarted the service - redeployed - recycled the replica - changed variables - changed the start command - changed the volume attachment - changed the SSH key - changed known_hosts - deleted or modified the prepared LOADER - written to the production /data volume Could you please check the Railway-side access path for this service, especially: 1. ssh.railway.com relay / gateway logs around the failed session. 2. Any gateway/session/rate-limit behavior that could explain successful SSH/SFTP sessions followed by an immediate session close. 3. Any service-specific routing or session-affinity issue for this service or running replica. 4. Whether `railway service files` and direct SSH share any relay/session/rate-limit layer that could produce this pattern. 5. Whether Railway exposes any non-SSH/SFTP read-only method to inspect a running service's ephemeral `/tmp` filesystem. 6. Whether the SSH/SFTP access path can be restored without restarting, redeploying, recycling the replica, or mutating the attached production volume. Important constraint: Please do not trigger or recommend a restart, redeploy, or replica recycle as an implicit diagnostic step without explicitly calling it out first. We currently have ephemeral forensic state in `/tmp` that must be preserved. Current exact blocker: REQUIRED_STRICT_PINNED_SSH_CHANNEL_CLOSED_DURING_LOADER_FINAL_INDEPENDENT_PROOF_ON_SELECTED_ALTERNATE_NETWORK Observed endpoint message: Connection closed by 66.33.22.3 port 22 Thank you.

freddy-tan
FREEOP

8 days ago

Following up after 24 hours on the updated alternate-network evidence above.

The issue remains blocked at the same point:

REQUIRED_STRICT_PINNED_SSH_CHANNEL_CLOSED_DURING_LOADER_FINAL_INDEPENDENT_PROOF_ON_SELECTED_ALTERNATE_NETWORK

We have kept the running replica and all current ephemeral /tmp state untouched since the failure.

No further SSH/SFTP attempts, restart, redeploy, replica recycle, volume mutation, or /data write have been performed.

Could a Railway staff member please review the SSH relay/gateway/session side of this case, or advise the next non-disruptive diagnostic step that preserves the current replica and /tmp state?

Thank you.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...