20 days ago
Description of the issue:
SQL Server crashes on every startup attempt with "Stack Overflow" (Reason: 0x00000006)
right after "Starting up database 'master'". Last errno before crash: 11
(EAGAIN / Resource temporarily unavailable).
The crash log shows the container detecting the full physical host resources instead
of my plan's allocated limits:
Processors: 48
Total Memory: 346488942592 bytes (~346 GB)
My actual plan limit (Settings > Scale) is 8 vCPU / 8 GB RAM.
What I've already tried (none fixed it):
- Custom Dockerfile with USER root (fixed an earlier /.system permission error)
- MSSQL_MEMORY_LIMIT_MB=2048
- Downgrading from SQL Server 2022 to 2019 image (avoided a separate "persistent hive" error)
- Startup trace flag -T8015 (force single NUMA node)
It seems SQL Server's thread/worker initialization scales based on the 48 detected
processors and hits a container-level thread limit, causing EAGAIN then Stack Overflow.
Question: is there a way to make cgroup CPU limits visible to /proc/cpuinfo inside
the container (so SQL Server detects the actual 8 vCPU allocation instead of the
full host), or another way to cap the container's visible processor count?
Error messages:
Reason: 0x00000006
Message: Stack Overflow
Last errno: 11 (Resource temporarily unavailable)
Processors: 48
Total Memory: 346488942592 bytes
Docker image: mcr.microsoft.com/mssql/server:2019-latest (custom Dockerfile, USER root, CMD with -T8015)
1 Replies
20 days ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 20 days ago
11 days ago
The 48 CPUs / 346 GB in the SQL Server log are not, by themselves, evidence that Railway's 8 CPU / 8 GB limits are missing. Microsoft's cgroup v2 documentation says SQL Server may still log the host CPU count while the kernel enforces the container quota. A CPU quota does not rewrite /proc/cpuinfo. SQL Server 2019 also lacks the cgroup v2 awareness added in SQL Server 2022 CU20. -T8015 changes NUMA behavior; it does not cap CPU count.
I would test this in two stages:
-
In the same container, before starting sqlservr, record:
stat -fc %T /sys/fs/cgroup
cat /sys/fs/cgroup/cpu.max /sys/fs/cgroup/memory.max /sys/fs/cgroup/pids.max /sys/fs/cgroup/pids.current /sys/fs/cgroup/pids.events /sys/fs/cgroup/memory.events
cat /proc/self/status
ulimit -u; ulimit -s
If it is cgroup v2, compare pids.events and memory.events before and immediately after the failed start. An increment of pids.events max supports a task/thread limit as the cause of EAGAIN; an OOM event points elsewhere. "Last errno 11" alone does not prove which resource failed. Please share these values and the lines immediately preceding the stack overflow (redact secrets).
- If the allowed CPU list includes 0-7, try a one-deployment experiment with the 2019 image by installing util-linux in the image and launching: taskset -c 0-7 /opt/mssql/bin/sqlservr -T8015. If the allowed IDs differ, use eight IDs from the allowed CPU list in /proc/self/status. This constrains process affinity before SQL Server starts. It is a diagnostic experiment, not a guaranteed fix.
For a durable path, test SQL Server 2022 CU20 or newer with a fresh, separate test volume, so the earlier persistent-hive error can be diagnosed separately from the existing database. If that starts, Microsoft's documented PROCESS AFFINITY setting and trace flag 8002 can align schedulers with the CPU quota. That SQL statement cannot help until the instance starts.
Microsoft reference: https://learn.microsoft.com/en-us/sql/linux/sql-server-linux-editions-and-components-2022#control-group-cgroup-v2-support