Production frontend intermittently using staging API URL
rafaelcoias
PROOP

24 days ago

Hi everyone,

I’ve found a bug with environment variables/API URLs that I wanted to report.

I have two servers in my production environment: one for the backend and one for the frontend. I also have two equivalent servers under my staging environment. Both environments use the same environment variable key names, but with different values for their respective API URLs.

The important thing is that I have not changed or modified the environment variables between deployments. The production frontend server consistently has the production API URL configured, and the staging frontend server consistently has the staging API URL configured.

There are also no hardcoded URLs anywhere in my code. The frontend gets the API URL exclusively from the environment variable.

Despite this, sometimes when I update/deploy the production frontend, it ends up using the staging API URL.

I can verify this clearly because:

The production Railway server shows the correct production API URL in its environment variables.

The code does not contain any hardcoded staging or production URLs.

The frontend's Network tab shows requests actually being sent to the staging API URL.

I have not changed the environment variable keys or values between these deployments.

This appears to happen specifically because production and staging use the same environment variable key name, with different values depending on the environment.

This makes me strongly suspect that there is some kind of environment variable caching, propagation, or cross-environment contamination where the staging value is somehow being used when building/deploying the production frontend.

I encountered this once before and was able to fix it by forcing a new deployment. A normal redeploy did not fix it. Unfortunately, it happened again today and my users were unable to use the application because the production frontend was calling the staging API.

This is particularly concerning because everything is configured correctly on my side, and there is nothing in the code that could cause the staging URL to be selected. It seems like the production and staging environments are somehow sharing or mixing up values for the same environment variable key.

Thanks in advance!

Solved$20 Bounty

2 Replies

Railway
BOT

24 days ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • 24 days ago


24 days ago

Hey, have you tried Cmd + K -> Deploy Latest Commit? Have you got "Skipped Builds" feature flag enabled in the service settings? If so, try disabling it.


dicecodes
PRO

23 days ago

To add the "why" behind the Skipped Builds suggestion above, since it explains both the bug and how to stop it recurring:

Frontend frameworks that use build-time env vars (Next.js NEXT_PUBLIC_, Vite VITE_, CRA REACT_APP_*) bake those values directly into the compiled JS bundle at build time — they are not read at runtime. Once a bundle is built, whatever values were present at that specific build are frozen into it permanently, no matter what the dashboard shows afterward.

If Skipped Builds is enabled and a production deploy gets skipped, Railway reuses a previously built artifact instead of building fresh. If that reused artifact happens to be one built while variable resolution was in a bad state, you'd see exactly this symptom: the dashboard correctly shows the production URL, but the actual served bundle has whatever was baked in during that earlier build. This also explains why forcing a new deployment fixed it before — a fresh build re-bakes the currently correct values, while a normal redeploy that gets skipped just re-serves the stale bundle.

Two things worth doing:

  1. Disable Skipped Builds on this service, as suggested above, to stop it from ever reusing a stale artifact.
  2. If you want a structural fix rather than relying on remembering to force-rebuild: switch to a runtime env injection pattern instead of build-time baking (e.g. a small script that writes a window.ENV config object at container start, read by the frontend at runtime) — this makes the served config always reflect the actual running container's environment variables, immune to build caching entirely.

If disabling Skipped Builds alone doesn't fully resolve it, worth double-checking that your production and staging services aren't accidentally sharing the same build cache key (e.g. via a shared Docker layer cache setting) rather than each having environment-scoped caches.


Status changed to Solved medim • about 11 hours ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...