6 days ago
Body:
We run a FastAPI app on Railway (Nixpacks builder) that uses the Python package weasyprint to render PDFs. It has been failing to import at runtime with a native-library-loading error; we've since added the apt packages weasyprint's own dependency list calls for, explicitly fixed LD_LIBRARY_PATH, and re-run ldconfig at build time, but the exact same failure persists across all of these changes. Diagnostics we run from inside the same container, at the moment of failure, show the specific library it can't find is present on disk, registered in ldconfig's cache, and on LD_LIBRARY_PATH — every cause we know how to check for is ruled out. The app itself deploys and runs fine; only the request path that imports weasyprint fails, and we catch that and fall back to a lower-quality PDF renderer, so nothing is currently broken for users. We'd like weasyprint working as intended, and would appreciate help understanding why dlopen() keeps failing here.
This was investigated over several iterations and are raising it with Railway after the standard causes (missing package, wrong search path, stale cache) were each individually ruled out via runtime diagnostics, without finding an application-level explanation.
Resources
Currently running deployment: deployed 2026-09-29 21:23 GMT+5:30 — this deployment builds and starts successfully; the issue reproduces on every PDF-generation request against it, it is not a deploy or build failure.
Builder image: ghcr.io/railwayapp/nixpacks:ubuntu-1745885067
Exact error (from our application logs — caught and handled, not a crash):
OSError: cannot load library 'libgobject-2.0-0': libgobject-2.0-0: cannot open shared object file: No such file or directory. Additionally, ctypes.util.find_library() did not manage to locate a library called 'libgobject-2.0-0'
and, loading the known absolute path directly, from a diagnostic we added ourselves:
ctypes.CDLL('/usr/lib/x86_64-linux-gnu/libgobject-2.0.so.0')
→ OSError('libglib-2.0.so.0: cannot open shared object file: No such file or directory')
What we checked, and the result of each — all run as a diagnostic inside the failing request, on deployment ef578caa:
dpkg -l shows libglib2.0-0t64 installed (via our aptPkgs entry libglib2.0-0), and the build log shows apt-get install completing without error.
find /usr/lib/x86_64-linux-gnu -iname 'glib' returns both the bare soname libglib-2.0.so.0 and the versioned libglib-2.0.so.0.8000.0.
ldconfig -p | grep glib returns libglib-2.0.so.0 (libc6,x86-64) => /lib/x86_64-linux-gnu/libglib-2.0.so.0, matching the file above.
LD_LIBRARY_PATH (we set this ourselves at Python import time, prepending /usr/lib/x86_64-linux-gnu) is: /usr/lib/x86_64-linux-gnu:/nix/store/.../gcc-13.3.0-lib/lib:/nix/store/.../zlib-1.3.1/lib:/usr/lib.
We added cmds = ["sudo ldconfig"] to the Nixpacks setup phase, after our aptPkgs, to rule out a stale or incomplete cache from apt's own postinst trigger. No change in behavior — the ldconfig output above is from after this ran.
Every one of these matches what should be needed for dlopen() to succeed, and the failure is identical across several redeploys with these changes in place. We suspect this could be a runtime restriction on dynamic library loading (the build log says "scheduling build on Metal builder") that wouldn't surface in ordinary file-existence or cache checks — but this is a guess on our part, not something we've confirmed, and we have no visibility into the runtime sandbox to check it ourselves.
What we're asking: is there anything on Railway's runtime side — a sandbox policy, an isolated filesystem view, or similar — that could cause dlopen()/ctypes.CDLL() to fail on a shared library that file-level checks confirm exists, is cached, and is on the search path? We're not asking for any change to our project, just help understanding why this specific call keeps failing.
We're not blocked: the app currently falls back to a pure-Python PDF renderer with functional but visually rougher output, and our production environment (a separate branch, unaffected by any of this work) is unaffected.
4 Replies
6 days ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 6 days ago
6 days ago
update: Following the earlier ldd result (both libgobject-2.0.so.0 and libglib-2.0.so.0 resolve every dependency cleanly via ldd, while ctypes.CDLL() of the exact same absolute path fails from inside our Python process), we checked the process's security state directly from /proc/self/status:
uid/euid/suid: 0/0/- gid/egid: 0/0
CapEff: 00000000800405fb CapBnd: 00000000800405fb
NoNewPrivs: 0 Seccomp: 2 Seccomp_filters: 3
UID/GID mismatch (glibc "secure execution" mode) is ruled out: uid == euid == 0, gid == egid. No mismatch, so glibc wouldn't be dropping LD_LIBRARY_PATH for this reason.
Seccomp restriction is also ruled out as the differentiator: a filter is active (mode 2, 3 stacked filters), but the ldd subprocess that succeeded was spawned as a child of this exact process — seccomp filters are inherited and can only get stricter in children, never looser. So ldd ran under the identical (or stricter) filter set and still succeeded at the equivalent operation (loading and resolving the library). If a seccomp rule were blocking a syscall dlopen() needs, ldd would have hit it too.
With both of those cleared, the only remaining difference between "resolves fine via ldd" and "fails via ctypes.CDLL/weasyprint's cffi-based loader" is something specific to how cffi itself invokes the native dynamic loader from inside an already-running Python process — not the filesystem, not the library, not the search path, and not a sandbox/security restriction we can detect. This is likely at the level of cffi's internal loading mechanism (e.g. an isolated loader namespace) rather than anything in our own configuration.
Nothing really for us to do I think. Any pointers appreciated
6 days ago
ldd succeeded because it's a child process. That's the difference.
You set LD_LIBRARY_PATH from Python at import time. That does nothing to the running process — glibc reads it once at exec and caches the search list. But subprocess passes os.environ to children, so ldd started with your prepended path and found the library. Your Python process never had it.
That's why ldd resolves and CDLL on the same absolute path fails.
Check it from inside the failing request:
open('/proc/self/environ','rb').read().split(b'\0') # env at exec
os.environ['LD_LIBRARY_PATH'] # your edited copy
/proc/self/environ doesn't change when Python calls putenv. If those disagree, the loader used the first one.
And:
env -u LD_LIBRARY_PATH ldd /usr/lib/x86_64-linux-gnu/libgobject-2.0.so.0
Should now fail on libglib-2.0.so.0, same as your process does.
Two other things:
ctypes.CDLL and cffi both call plain dlopen(). Isolated namespaces need dlmopen(), which neither uses — so that's not it.
ldconfig -p was never relevant. Nixpacks gives you a Nix-built Python (your path has /nix/store/…gcc-13.3.0-lib/lib), and nixpkgs patches glibc to read a store-local ld.so.cache instead of /etc/ld.so.cache.
Fix: set LD_LIBRARY_PATH as a Railway service variable so it exists before the process starts. Or install glib/pango/cairo/gdk-pixbuf via nixPkgs instead of aptPkgs so they're RPATH-resolved. Or use a Dockerfile.
6 days ago
Thank you for the guidance. Problem resolved.
The library file existing on disk, being correctly cached by ldconfig, and being on LD_LIBRARY_PATH were all real, but none of them were the actual problem. The real cause: python311 was installed via Nixpacks' nixPkgs, so the venv's python binary is a Nix-built binary linked against Nix's own patched glibc via RUNPATH — not Ubuntu's system glibc, even though the apt-installed native libraries (pango/cairo/glib) live in the normal /usr/lib/x86_64-linux-gnu.
Putting that multiarch directory on LD_LIBRARY_PATH looked like the obvious fix, and it even correctly resolved weasyprint's own libraries — but LD_LIBRARY_PATH takes priority over RUNPATH for every dynamic load in the process, including the interpreter's own dependency on glibc itself. That broke the Nix-built python binary before it even started (error while loading shared libraries: __vdso_gettimeofday: invalid mode for dlopen()) — a worse failure than the original problem, since it took the whole process down instead of leaving a working fallback.
The fix that worked, with zero LD_LIBRARY_PATH risk: preload each native library by absolute path via ctypes.CDLL(path, mode=ctypes.RTLD_GLOBAL) before the library that needs them (weasyprint, via cffi) ever imports. glibc's dlopen() always checks whether a library with a matching soname is already resident in the process before doing any path search — so the later bare-name dlopen() calls find them already loaded and never search at all. This never touches the environment, so it can't affect the interpreter's own resolution.
One more thing worth flagging for anyone trying this: don't hand-list which libraries to preload. A hand-written list only ever surfaces one missing dependency per attempt (fixed libgobject → failure moved to libglib → fixed that → failure moved to libpango, which needed libfribidi/libthai that were never in the list). Running ldd against each top-level library and preloading the full union of what it resolves — since ldd correctly resolves the whole transitive closure in one pass in this environment — fixed it completely in one shot.
Best practice for Nixpacks + a language runtime from nixPkgs + native libs from aptPkgs: never add the apt library directory to LD_LIBRARY_PATH if your runtime itself came from nixPkgs. Treat that as a hard rule, not a first thing to try — it can silently break the interpreter for a completely unrelated reason. Preload via absolute path instead.
Thanks to @travismcelfresh who pinpointed the actual LD_LIBRARY_PATH-timing mechanism — that was the piece that made the rest of this tractable.
5 days ago
Nice work!

