Sandboxes:
docker exec
enters the wrong filesystem — breaks every container HEALTHCHECK
atx365
HOBBYOP

a month ago

Summary

Inside a Railway sandbox, docker exec does not place the process in the target

container's mount namespace — it lands in an unrelated filesystem. A process

exec'd into an alpine:3.20 container reports Debian 12.

Docker implements HEALTHCHECK as an exec, so every healthcheck that runs a

binary the image provides fails with exit 127 and the container is marked

unhealthy forever, no matter how healthy the service is. That also means

depends_on: condition: service_healthy never releases, so dependent services

never start.

docker run is unaffected — a container's own main process gets the correct

filesystem. It is specifically entering an already-running container that breaks.

Environment

  • Docker 29.1.2, build 890dcca
  • Kernel 6.18.46-railway
  • Sandbox created with railway sandbox create --checkpoint <name>
  • dockerd: /opt/railway/ext/docker/usr/local/bin/dockerd --host=unix:///var/run/docker.sock --data-root=/var/lib/docker --ipv6=true --firewall-backend=nftables --cgroup-parent=/workload/docker
  • Reproduces on both storage drivers: overlayfs (containerd snapshotter, the default) and overlay2

Steps to reproduce

Run inside any sandbox:

printf 'FROM alpine:3.20\nRUN touch /file-added-by-a-layer\n' > Dockerfile
docker build -q -t exec-repro .
CID="$(docker run -d exec-repro sleep 120)"

docker run --rm exec-repro ls -la /file-added-by-a-layer   # exists
docker exec "$CID"        ls -la /file-added-by-a-layer    # missing

Expected

Both commands list the file. docker exec sees the container's filesystem.

Actual

main process : -rw-r--r-- 1 root root 0 /file-added-by-a-layer
docker exec  : ls: cannot access '/file-added-by-a-layer': No such file or directory

The exec'd process is not in a degraded view of the container — it is in a

different OS:

docker run --rm exec-repro cat /etc/alpine-release   # 3.20.10
docker exec "$CID" cat /etc/alpine-release           # No such file or directory
docker exec "$CID" head -1 /etc/os-release           # PRETTY_NAME="Debian GNU/Linux 12 (bookworm)"

Root mount differs:

docker run --rm exec-repro sh -c 'grep " / " /proc/self/mountinfo'
#   overlay rw,lowerdir=/var/lib/docker/overlay2/l/AC2M2EFB...,upperdir=...

docker exec "$CID" sh -c 'grep " / " /proc/self/mountinfo'
#   /dev/root ro

The exec'd process also sees the sandbox's entire mount table, including

/var/lib/docker/overlay2/... mounts belonging to containers — so it is outside

the container's mount namespace rather than inside a broken one. It is in the

container's UTS namespace: hostname correctly returns the container id.

Real-world impact (stock postgres:16)

services:
  db:
    image: postgres:16
    environment: { POSTGRES_PASSWORD: x }
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U postgres"]
db-1 | database system is ready to accept connections   <- the server is serving
healthcheck -> exit 127, "pg_isready: not found"        <- exec cannot see it
docker inspect .State.Health.Status -> unhealthy

Any orchestration waiting on docker compose ps db | grep "(healthy)" waits

forever.

Ruled out

  • Permissions. CapEff/CapPrm are identical between the container's main

    process and the exec'd process (00000000a80425fb). docker exec -u 0 and

    docker exec --privileged change nothing. The errno is ENOENT, not EACCES.

  • Storage driver. Setting {"features":{"containerd-snapshotter":false}} in

    /etc/docker/daemon.json and restarting dockerd successfully switches the

    driver to overlay2 (Backing Filesystem extfs, Native Overlay Diff true).

    Behaviour is unchanged.

  • The sandbox host. A marker file written at / in the sandbox is not

    visible from the exec'd process, and the sandbox is Debian 13 while the exec'd

    process reports Debian 12.

  • The image. docker image inspect and a --entrypoint sh container both

    show the full, correct filesystem for the same image id.

Severity

High for anyone running docker-compose stacks in sandboxes. Healthchecks and

depends_on: service_healthy are standard Compose practice, and there is no

application-side fix — no exec-based probe can work while exec is in the wrong

filesystem.

Solved

1 Replies

Railway
BOT

a month ago

Sandboxes are currently in Priority Boarding (beta), and the behavior you've documented with docker exec landing in the wrong mount namespace is not something we have a workaround for on the sandbox side. Your report and reproduction steps here on Central Station are exactly how Priority Boarding issues get visibility with the team that owns sandbox infrastructure.


Status changed to Awaiting User Response Railway • about 1 month ago


Railway
BOT

a month ago

This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!

Status changed to Solved Railway • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...