Build installs correct dependency but runtime container fails on the same import — image mismatch?
bonpiedlaroute
HOBBYOP

2 months ago

I am aware of the two open incidents (GitHub auto-deploys, and deployments queued across all regions). I am reporting this because the symptom is specifically a build/runtime image mismatch, not a delay: the build logs confirm langchain-core 0.3.83 is installed and AsyncCallbackManagerForLLMRun imports successfully at build time, yet the running container fails on that exact import. Building the identical Dockerfile locally with docker build --no-cache produces a working image (celery worker starts, symbol present). Happy to re-test once the incidents are resolved — I would just like to know whether the runtime container can be serving a different image digest than the one just built.

$10 Bounty

6 Replies

Railway
BOT

2 months ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 2 months ago


2 months ago

You can add NO_CACHE=1 to the Service Variables if you want to try rebuild without cache


manuproject
FREE

2 months ago

worth separating "different image" from "same image, different import resolution", because the second explains this better.

if your Dockerfile does COPY . . after the pip install, everything in your repo lands in /app after the packages are installed. at runtime cwd is /app and python searches there first, so a stale local directory, a vendored copy, or a committed .venv can shadow the installed package. correct version installed, import of a newer symbol still fails. it would also explain why a clean local docker build --no-cache works, your local checkout likely doesn't carry whatever's in the build context.

run this as the start command temporarily and it settles it:

python -c "import langchain_core, sys; print(langchain_core.version); print(langchain_core.file); print(sys.path[:5])"

if file is anywhere other than site-packages, that's the cause. if it shows an older version from site-packages, then the installed layer really is stale and your digest theory is right, worth pairing that with pip show langchain-core at runtime to compare against the build log.

either way that output tells you which it is rather than guessing. adding a .dockerignore with .venv, pycache and any local package dirs removes the shadowing possibility entirely.


bonpiedlaroute
HOBBYOP

2 months ago

Thanks both.

@mayori — already tried NO_CACHE=1. The build then re-ran fully (I could see every package being installed, langchain-core 0.3.83, and my build-time assert passing) but the runtime container still failed on the same import.

@manuproject — good hypothesis, and I had the same one early on, but I think it's ruled out: the Dockerfile already has a .dockerignore excluding .venv/, pycache/ and *.py[cod], plus an explicit RUN rm -rf /app/.venv and a .pyc purge after COPY . .. Inspecting the built image, there is no langchain_core, no .venv and no site-packages anywhere under /app. There's also no committed venv in the repo (git ls-files returns nothing for it).

That said, your diagnostic is the right move — I'm running it as the start command now and will post the raw output, since it distinguishes "different image" from "different import resolution" definitively.

One extra data point: building the same Dockerfile locally with docker build --no-cache and running the worker's exact start command (celery -A app.celery_app worker --loglevel=info --concurrency=2 --pool=solo) starts the worker cleanly — tasks registered, no ImportError. Same Dockerfile, same poetry.lock, same commit.


bonpiedlaroute
HOBBYOP

2 months ago

Diagnostic output, run as the start command on the worker:

PATH: ['', '/usr/local/lib/python311.zip', '/usr/local/lib/python3.11', '/usr/local/lib/python3.11/lib-dynload', '/usr/local/lib/python3.11/site-packages']

CORE_FILE: /usr/local/lib/python3.11/site-packages/langchain_core/init.py

CORE_META: 0.3.83

ANTH_META: 0.3.22

CB_FILE: /usr/local/lib/python3.11/site-packages/langchain_core/callbacks/init.py

HAS_SYMBOL: False

So it is neither of the two theories. sys.path is clean, nothing shadows the package, and the module resolves from site-packages — no vendored copy, no .venv, no stale local dir (thanks @manuproject, that ruled it out cleanly).

But look at the contradiction: the metadata says 0.3.83, while the actual callbacks/init.py at that exact path does not contain AsyncCallbackManagerForLLMRun. In a correct 0.3.83 install that file is 139 lines and mentions the symbol three times (verified in a locally built image). So the dist-info poetry wrote is 0.3.83, but the .py files in langchain_core/ are from an older release — a partially overwritten package directory.

That is why every version-based check passed while the runtime import failed: importlib.metadata.version() reads the dist-info, not the files.

My working theory is a stale layer where an older langchain_core was present, and poetry install --sync updated the metadata without replacing every file — the overlay/union FS kept old files alive. Notably this survived NO_CACHE=1, which is the part I still don't understand: with caching disabled the whole layer was rebuilt from python:3.11-slim, so where would the old files come from?

Fix that works (local docker build --no-cache verified, worker starts clean): physically delete the package trees, then reinstall, and assert on the symbol rather than the version:

RUN poetry install --no-interaction --no-ansi --no-root --sync \

&& rm -rf /usr/local/lib/python3.11/site-packages/langchain_core \

       /usr/local/lib/python3.11/site-packages/langchain_core-*.dist-info \

       /usr/local/lib/python3.11/site-packages/langchain_anthropic \

       /usr/local/lib/python3.11/site-packages/langchain_anthropic-*.dist-info \

&& pip install --no-cache-dir --force-reinstall --no-deps \

  "langchain-core==0.3.83" "langchain-anthropic==0.3.22" \

&& python3 -c "import langchain_core.callbacks as c; assert hasattr(c,'AsyncCallbackManagerForLLMRun'), c.file"

Deploying that now. Leaving the thread open because the underlying question stands: can a Railway build layer serve partially-overwritten package files, even with NO_CACHE=1? If so it's worth knowing, since any version-based check will happily pass over it.


bonpiedlaroute
HOBBYOP

2 months ago

Resolved the ambiguity — this is not a dependency problem at all.

I added a gate to the Dockerfile that runs from /app, i.e. the worker's exact runtime cwd, asserting on the symbol rather than the version, and importing the very module chain that crashes:

RUN cd /app \

&& python3 -c "import langchain_core.callbacks as c; assert hasattr(c,'AsyncCallbackManagerForLLMRun'), c.file; import app.tasks.match_tasks; print('runtime import chain OK from', c.file)"

Build output (step 8, with NO_CACHE=1):

runtime import chain OK from /usr/local/lib/python3.11/site-packages/langchain_core/callbacks/init.py

Image digest produced: sha256:1038ba092749749a9059c1f5b196b33a5cea3304d1b8c158aa9b88c3563af704

Runtime output of that same deployment, same file, same path:

CB_FILE: /usr/local/lib/python3.11/site-packages/langchain_core/callbacks/init.py

HAS_SYMBOL: False

CORE_META: 0.3.83

PATH: ['', '/usr/local/lib/python311.zip', '/usr/local/lib/python3.11', '/usr/local/lib/python3.11/lib-dynload', '/usr/local/lib/python3.11/site-packages']

The build proves the symbol is present in that file and that import app.tasks.match_tasks succeeds from /app. The container then reports the same file lacks the symbol. sys.path is clean, no shadowing, no .venv, no vendored copy — and --force-reinstall after an explicit rm -rf of the package trees did not change the runtime behaviour.

A build cannot assert a file's contents and the container from that image observe different contents. Unless the running container is not from sha256:1038ba09….

Could someone from Railway confirm which image digest deployment on the playmakerly-worker service is actually running? Project a338dd04-62d2-4392-b4f6-47ff526a51ac, service 122611cf-6542-44ad-80a1-794853cb3232, build at 2026-08-18 09:18 UTC.

Things already ruled out: NO_CACHE=1 (build fully re-ran, every package reinstalled), layer caching (Caching Disabled in the logs), builder misconfig (railway.json pins DOCKERFILE, logs show load build definition from backend/Dockerfile), local reproduction (docker build --no-cache + the worker's exact start command works — tasks registered, no ImportError).


bonpiedlaroute
HOBBYOP

2 months ago

Closing this from my side — the app is fixed, but the underlying behaviour is

worth flagging because it is reproducible and silently misleads any

version-based check.

WHAT I OBSERVED (3 independent times, same signature)

The build asserts something about the filesystem, and the container started

from that build reports the opposite.

Most clear-cut case. Dockerfile step:

RUN pip uninstall -y charset-normalizer \

&& python3 -c "import requests; print('requests OK without charset-normalizer')"

Build log (with NO_CACHE=1, Caching Disabled confirmed in the log):

requests OK without charset-normalizer

entrypoints + agent import OK

Image digest: sha256:7179aa3ae794083a048c3f5fb8c57718ccf7cda93859fa9eea829ae8d473f030

created: 2026-08-18T16:51:23Z

Container started 2026-08-18T17:00:47Z (9 min after that image was created).

Diagnostic run as the start command:

CN_INSTALLED 3.4.4

/usr/local/lib/python3.11/site-packages/charset_normalizer

The build removed the package and verified its absence. The container has it.

Earlier instances of the same pattern:

  • Build installed charset-normalizer with --no-binary :all: (pure Python,

    md.py, no .so). Runtime traceback showed src/charset_normalizer/cd.pyx,

    i.e. the compiled extension.

  • Build ran `python -c "from langchain_core.callbacks import

    AsyncCallbackManagerForLLMRun; import app.tasks.match_tasks"` successfully

    from /app. The container failed on that exact import, same file path.

RULED OUT

  • Layer cache: NO_CACHE=1, Caching Disabled in the log, all steps re-ran.

  • Old service state: deleted the service and recreated it from scratch.

  • Builder misconfig: railway.json pins DOCKERFILE; logs show

    load build definition from backend/Dockerfile.

  • Import shadowing: sys.path clean, module resolved from site-packages,

    .dockerignore excludes .venv, no vendored copy in the repo.

  • Dependency versions: poetry.lock is hash-pinned; at runtime the metadata

    reported the pinned version — only the files disagreed.

  • My code: same commit, docker build --no-cache locally → api, celery

    worker and beat all start, no ImportError.

WHAT I CANNOT EXPLAIN

Whether the runtime pulls a different image than the one just built, or

something re-materialises packages at container start. I have no visibility

into either. I am not claiming to know the cause — only that the build and the

runtime disagreed about the same file paths, repeatedly, across a fresh

service and with caching disabled.

WHY IT MATTERS BEYOND MY CASE

The failure surfaced as:

ImportError: cannot import name 'AsyncCallbackManagerForLLMRun'

from 'langchain_core.callbacks'

because langchain-core resolves that symbol through a lazy getattr that

imports requests -> langsmith. So a broken charset-normalizer got reported as

a langchain problem, and every version/symbol assertion in the build passed

while production crash-looped. I spent a full day on this. Anyone hitting a

compiled-extension failure inside a lazy import chain will get an error

message pointing at the wrong package entirely.

HOW I RESOLVED IT

Stopped depending on the chain: dropped langchain/langgraph/langsmith

entirely and rewrote the agent on the Anthropic SDK directly (it was only

doing a tool-use loop the SDK does natively). All three services have been

stable since.

If anyone from Railway wants to look: project a338dd04-62d2-4392-b4f6-47ff526a51ac,

image digest sha256:7179aa3a…, container started 17:00:47Z on 2026-08-18.

Happy to re-run any diagnostic.

FOR ANYONE LANDING HERE WITH THE SAME ImportError

Don't trust the message. Run the failing import chain directly at runtime and

print the traceback:

python -c "import langsmith.run_helpers"

That is what finally showed the real AttributeError.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...