docker execenters the wrong filesystem — breaks every container HEALTHCHECK
a month ago
Summary
Inside a Railway sandbox, docker exec does not place the process in the target
container's mount namespace — it lands in an unrelated filesystem. A process
exec'd into an alpine:3.20 container reports Debian 12.
Docker implements HEALTHCHECK as an exec, so every healthcheck that runs a
binary the image provides fails with exit 127 and the container is marked
unhealthy forever, no matter how healthy the service is. That also means
depends_on: condition: service_healthy never releases, so dependent services
never start.
docker run is unaffected — a container's own main process gets the correct
filesystem. It is specifically entering an already-running container that breaks.
Environment
- Docker 29.1.2, build 890dcca
- Kernel 6.18.46-railway
- Sandbox created with
railway sandbox create --checkpoint <name> - dockerd:
/opt/railway/ext/docker/usr/local/bin/dockerd --host=unix:///var/run/docker.sock --data-root=/var/lib/docker --ipv6=true --firewall-backend=nftables --cgroup-parent=/workload/docker - Reproduces on both storage drivers:
overlayfs(containerd snapshotter, the default) andoverlay2
Steps to reproduce
Run inside any sandbox:
printf 'FROM alpine:3.20\nRUN touch /file-added-by-a-layer\n' > Dockerfile
docker build -q -t exec-repro .
CID="$(docker run -d exec-repro sleep 120)"
docker run --rm exec-repro ls -la /file-added-by-a-layer # exists
docker exec "$CID" ls -la /file-added-by-a-layer # missingExpected
Both commands list the file. docker exec sees the container's filesystem.
Actual
main process : -rw-r--r-- 1 root root 0 /file-added-by-a-layer
docker exec : ls: cannot access '/file-added-by-a-layer': No such file or directoryThe exec'd process is not in a degraded view of the container — it is in a
different OS:
docker run --rm exec-repro cat /etc/alpine-release # 3.20.10
docker exec "$CID" cat /etc/alpine-release # No such file or directory
docker exec "$CID" head -1 /etc/os-release # PRETTY_NAME="Debian GNU/Linux 12 (bookworm)"Root mount differs:
docker run --rm exec-repro sh -c 'grep " / " /proc/self/mountinfo'
# overlay rw,lowerdir=/var/lib/docker/overlay2/l/AC2M2EFB...,upperdir=...
docker exec "$CID" sh -c 'grep " / " /proc/self/mountinfo'
# /dev/root roThe exec'd process also sees the sandbox's entire mount table, including
/var/lib/docker/overlay2/... mounts belonging to containers — so it is outside
the container's mount namespace rather than inside a broken one. It is in the
container's UTS namespace: hostname correctly returns the container id.
Real-world impact (stock postgres:16)
services:
db:
image: postgres:16
environment: { POSTGRES_PASSWORD: x }
healthcheck:
test: ["CMD-SHELL", "pg_isready -U postgres"]db-1 | database system is ready to accept connections <- the server is serving
healthcheck -> exit 127, "pg_isready: not found" <- exec cannot see it
docker inspect .State.Health.Status -> unhealthyAny orchestration waiting on docker compose ps db | grep "(healthy)" waits
forever.
Ruled out
-
Permissions.
CapEff/CapPrmare identical between the container's mainprocess and the exec'd process (
00000000a80425fb).docker exec -u 0anddocker exec --privilegedchange nothing. The errno isENOENT, notEACCES. -
Storage driver. Setting
{"features":{"containerd-snapshotter":false}}in/etc/docker/daemon.jsonand restarting dockerd successfully switches thedriver to
overlay2(Backing Filesystem extfs,Native Overlay Diff true).Behaviour is unchanged.
-
The sandbox host. A marker file written at
/in the sandbox is notvisible from the exec'd process, and the sandbox is Debian 13 while the exec'd
process reports Debian 12.
-
The image.
docker image inspectand a--entrypoint shcontainer bothshow the full, correct filesystem for the same image id.
Severity
High for anyone running docker-compose stacks in sandboxes. Healthchecks and
depends_on: service_healthy are standard Compose practice, and there is no
application-side fix — no exec-based probe can work while exec is in the wrong
filesystem.
1 Replies
a month ago
Sandboxes are currently in Priority Boarding (beta), and the behavior you've documented with docker exec landing in the wrong mount namespace is not something we have a workaround for on the sandbox side. Your report and reproduction steps here on Central Station are exactly how Priority Boarding issues get visibility with the team that owns sandbox infrastructure.
Status changed to Awaiting User Response Railway • about 1 month ago
a month ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • about 1 month ago