2 months ago
I am aware of the two open incidents (GitHub auto-deploys, and deployments queued across all regions). I am reporting this because the symptom is specifically a build/runtime image mismatch, not a delay: the build logs confirm langchain-core 0.3.83 is installed and AsyncCallbackManagerForLLMRun imports successfully at build time, yet the running container fails on that exact import. Building the identical Dockerfile locally with docker build --no-cache produces a working image (celery worker starts, symbol present). Happy to re-test once the incidents are resolved — I would just like to know whether the runtime container can be serving a different image digest than the one just built.
6 Replies
2 months ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 2 months ago
2 months ago
You can add NO_CACHE=1 to the Service Variables if you want to try rebuild without cache
2 months ago
worth separating "different image" from "same image, different import resolution", because the second explains this better.
if your Dockerfile does COPY . . after the pip install, everything in your repo lands in /app after the packages are installed. at runtime cwd is /app and python searches there first, so a stale local directory, a vendored copy, or a committed .venv can shadow the installed package. correct version installed, import of a newer symbol still fails. it would also explain why a clean local docker build --no-cache works, your local checkout likely doesn't carry whatever's in the build context.
run this as the start command temporarily and it settles it:
python -c "import langchain_core, sys; print(langchain_core.version); print(langchain_core.file); print(sys.path[:5])"
if file is anywhere other than site-packages, that's the cause. if it shows an older version from site-packages, then the installed layer really is stale and your digest theory is right, worth pairing that with pip show langchain-core at runtime to compare against the build log.
either way that output tells you which it is rather than guessing. adding a .dockerignore with .venv, pycache and any local package dirs removes the shadowing possibility entirely.
2 months ago
Thanks both.
@mayori — already tried NO_CACHE=1. The build then re-ran fully (I could see every package being installed, langchain-core 0.3.83, and my build-time assert passing) but the runtime container still failed on the same import.
@manuproject — good hypothesis, and I had the same one early on, but I think it's ruled out: the Dockerfile already has a .dockerignore excluding .venv/, pycache/ and *.py[cod], plus an explicit RUN rm -rf /app/.venv and a .pyc purge after COPY . .. Inspecting the built image, there is no langchain_core, no .venv and no site-packages anywhere under /app. There's also no committed venv in the repo (git ls-files returns nothing for it).
That said, your diagnostic is the right move — I'm running it as the start command now and will post the raw output, since it distinguishes "different image" from "different import resolution" definitively.
One extra data point: building the same Dockerfile locally with docker build --no-cache and running the worker's exact start command (celery -A app.celery_app worker --loglevel=info --concurrency=2 --pool=solo) starts the worker cleanly — tasks registered, no ImportError. Same Dockerfile, same poetry.lock, same commit.
2 months ago
Diagnostic output, run as the start command on the worker:
PATH: ['', '/usr/local/lib/python311.zip', '/usr/local/lib/python3.11', '/usr/local/lib/python3.11/lib-dynload', '/usr/local/lib/python3.11/site-packages']
CORE_FILE: /usr/local/lib/python3.11/site-packages/langchain_core/init.py
CORE_META: 0.3.83
ANTH_META: 0.3.22
CB_FILE: /usr/local/lib/python3.11/site-packages/langchain_core/callbacks/init.py
HAS_SYMBOL: False
So it is neither of the two theories. sys.path is clean, nothing shadows the package, and the module resolves from site-packages — no vendored copy, no .venv, no stale local dir (thanks @manuproject, that ruled it out cleanly).
But look at the contradiction: the metadata says 0.3.83, while the actual callbacks/init.py at that exact path does not contain AsyncCallbackManagerForLLMRun. In a correct 0.3.83 install that file is 139 lines and mentions the symbol three times (verified in a locally built image). So the dist-info poetry wrote is 0.3.83, but the .py files in langchain_core/ are from an older release — a partially overwritten package directory.
That is why every version-based check passed while the runtime import failed: importlib.metadata.version() reads the dist-info, not the files.
My working theory is a stale layer where an older langchain_core was present, and poetry install --sync updated the metadata without replacing every file — the overlay/union FS kept old files alive. Notably this survived NO_CACHE=1, which is the part I still don't understand: with caching disabled the whole layer was rebuilt from python:3.11-slim, so where would the old files come from?
Fix that works (local docker build --no-cache verified, worker starts clean): physically delete the package trees, then reinstall, and assert on the symbol rather than the version:
RUN poetry install --no-interaction --no-ansi --no-root --sync \
&& rm -rf /usr/local/lib/python3.11/site-packages/langchain_core \
/usr/local/lib/python3.11/site-packages/langchain_core-*.dist-info \
/usr/local/lib/python3.11/site-packages/langchain_anthropic \
/usr/local/lib/python3.11/site-packages/langchain_anthropic-*.dist-info \&& pip install --no-cache-dir --force-reinstall --no-deps \
"langchain-core==0.3.83" "langchain-anthropic==0.3.22" \&& python3 -c "import langchain_core.callbacks as c; assert hasattr(c,'AsyncCallbackManagerForLLMRun'), c.file"
Deploying that now. Leaving the thread open because the underlying question stands: can a Railway build layer serve partially-overwritten package files, even with NO_CACHE=1? If so it's worth knowing, since any version-based check will happily pass over it.
2 months ago
Resolved the ambiguity — this is not a dependency problem at all.
I added a gate to the Dockerfile that runs from /app, i.e. the worker's exact runtime cwd, asserting on the symbol rather than the version, and importing the very module chain that crashes:
RUN cd /app \
&& python3 -c "import langchain_core.callbacks as c; assert hasattr(c,'AsyncCallbackManagerForLLMRun'), c.file; import app.tasks.match_tasks; print('runtime import chain OK from', c.file)"
Build output (step 8, with NO_CACHE=1):
runtime import chain OK from /usr/local/lib/python3.11/site-packages/langchain_core/callbacks/init.py
Image digest produced: sha256:1038ba092749749a9059c1f5b196b33a5cea3304d1b8c158aa9b88c3563af704
Runtime output of that same deployment, same file, same path:
CB_FILE: /usr/local/lib/python3.11/site-packages/langchain_core/callbacks/init.py
HAS_SYMBOL: False
CORE_META: 0.3.83
PATH: ['', '/usr/local/lib/python311.zip', '/usr/local/lib/python3.11', '/usr/local/lib/python3.11/lib-dynload', '/usr/local/lib/python3.11/site-packages']
The build proves the symbol is present in that file and that import app.tasks.match_tasks succeeds from /app. The container then reports the same file lacks the symbol. sys.path is clean, no shadowing, no .venv, no vendored copy — and --force-reinstall after an explicit rm -rf of the package trees did not change the runtime behaviour.
A build cannot assert a file's contents and the container from that image observe different contents. Unless the running container is not from sha256:1038ba09….
Could someone from Railway confirm which image digest deployment on the playmakerly-worker service is actually running? Project a338dd04-62d2-4392-b4f6-47ff526a51ac, service 122611cf-6542-44ad-80a1-794853cb3232, build at 2026-08-18 09:18 UTC.
Things already ruled out: NO_CACHE=1 (build fully re-ran, every package reinstalled), layer caching (Caching Disabled in the logs), builder misconfig (railway.json pins DOCKERFILE, logs show load build definition from backend/Dockerfile), local reproduction (docker build --no-cache + the worker's exact start command works — tasks registered, no ImportError).
2 months ago
Closing this from my side — the app is fixed, but the underlying behaviour is
worth flagging because it is reproducible and silently misleads any
version-based check.
WHAT I OBSERVED (3 independent times, same signature)
The build asserts something about the filesystem, and the container started
from that build reports the opposite.
Most clear-cut case. Dockerfile step:
RUN pip uninstall -y charset-normalizer \
&& python3 -c "import requests; print('requests OK without charset-normalizer')"
Build log (with NO_CACHE=1, Caching Disabled confirmed in the log):
requests OK without charset-normalizer
entrypoints + agent import OK
Image digest: sha256:7179aa3ae794083a048c3f5fb8c57718ccf7cda93859fa9eea829ae8d473f030
created: 2026-08-18T16:51:23Z
Container started 2026-08-18T17:00:47Z (9 min after that image was created).
Diagnostic run as the start command:
CN_INSTALLED 3.4.4
/usr/local/lib/python3.11/site-packages/charset_normalizer
The build removed the package and verified its absence. The container has it.
Earlier instances of the same pattern:
-
Build installed charset-normalizer with
--no-binary :all:(pure Python,md.py, no .so). Runtime traceback showed
src/charset_normalizer/cd.pyx,i.e. the compiled extension.
-
Build ran `python -c "from langchain_core.callbacks import
AsyncCallbackManagerForLLMRun; import app.tasks.match_tasks"` successfully
from /app. The container failed on that exact import, same file path.
RULED OUT
-
Layer cache:
NO_CACHE=1,Caching Disabledin the log, all steps re-ran. -
Old service state: deleted the service and recreated it from scratch.
-
Builder misconfig: railway.json pins DOCKERFILE; logs show
load build definition from backend/Dockerfile. -
Import shadowing: sys.path clean, module resolved from site-packages,
.dockerignore excludes .venv, no vendored copy in the repo.
-
Dependency versions: poetry.lock is hash-pinned; at runtime the metadata
reported the pinned version — only the files disagreed.
-
My code: same commit,
docker build --no-cachelocally → api, celeryworker and beat all start, no ImportError.
WHAT I CANNOT EXPLAIN
Whether the runtime pulls a different image than the one just built, or
something re-materialises packages at container start. I have no visibility
into either. I am not claiming to know the cause — only that the build and the
runtime disagreed about the same file paths, repeatedly, across a fresh
service and with caching disabled.
WHY IT MATTERS BEYOND MY CASE
The failure surfaced as:
ImportError: cannot import name 'AsyncCallbackManagerForLLMRun'
from 'langchain_core.callbacks'
because langchain-core resolves that symbol through a lazy getattr that
imports requests -> langsmith. So a broken charset-normalizer got reported as
a langchain problem, and every version/symbol assertion in the build passed
while production crash-looped. I spent a full day on this. Anyone hitting a
compiled-extension failure inside a lazy import chain will get an error
message pointing at the wrong package entirely.
HOW I RESOLVED IT
Stopped depending on the chain: dropped langchain/langgraph/langsmith
entirely and rewrote the agent on the Anthropic SDK directly (it was only
doing a tool-use loop the SDK does natively). All three services have been
stable since.
If anyone from Railway wants to look: project a338dd04-62d2-4392-b4f6-47ff526a51ac,
image digest sha256:7179aa3a…, container started 17:00:47Z on 2026-08-18.
Happy to re-run any diagnostic.
FOR ANYONE LANDING HERE WITH THE SAME ImportError
Don't trust the message. Run the failing import chain directly at runtime and
print the traceback:
python -c "import langsmith.run_helpers"
That is what finally showed the real AttributeError.