Production outage: PostgreSQL 28P01 after managed password regeneration
jonatamelodev
HOBBYOP

4 hours ago

Environment: production

A single password regeneration was performed using the managed PostgreSQL “Regenerate password” action.

The PostgreSQL service remains Online, but every fresh backend deployment fails during TypeORM startup with PostgreSQL error code 28P01 and routine auth_failed.

The backend DATABASE_URL is a managed reference to the PostgreSQL service. In the PostgreSQL service:

  • PGPASSWORD already references POSTGRES_PASSWORD.
  • DATABASE_URL and DATABASE_PUBLIC_URL are already reference expressions derived from POSTGRES_PASSWORD.
  • These references existed before the password regeneration.

A completely fresh backend deployment was created after PostgreSQL was healthy, but authentication still fails with 28P01.

No manual password edits, ALTER USER commands, migrations, restores, rollbacks, volume changes, or database recreation were performed.

The Railway dashboard also displayed an active incident affecting services and databases in US West.

Please verify whether the password stored for the PostgreSQL role matches the managed POSTGRES_PASSWORD and safely resynchronize them without recreating the service, deleting the volume, or modifying application data.

Please confirm whether redeploying the backend will be sufficient after resynchronization.

Similar reported case:

https://station.railway.com/questions/production-incident-postgre-sql-authenti-e1184b4f

Solved

6 Replies

Railway
BOT

4 hours ago

Your Postgres logs confirm the mismatch: since the regeneration, every connection as the postgres user is rejected with a password authentication failure. That means the password stored in the database's data directory no longer matches POSTGRES_PASSWORD. The database only reads that variable when it initializes an empty data directory, so redeploying Postgres will not resync it on its own.

To bring the two back in line without touching the volume or your data, set POSTGRES_PASSWORD on the Postgres service back to the password the database actually holds. That is the value from before the regeneration, which you may still have in an earlier copy of DATABASE_URL or a local env file. Because PGPASSWORD, DATABASE_URL and DATABASE_PUBLIC_URL are references to it, they follow automatically. Then redeploy the backend. Once both sides match, Regenerate works normally again, as described in our database docs.

To answer your last question: yes, once the variable and the database agree, redeploying the backend is enough. It picks up the referenced DATABASE_URL on startup, and nothing else needs to change.


Status changed to Awaiting User Response Railway • about 4 hours ago


Railway

Your Postgres logs confirm the mismatch: since the regeneration, every connection as the `postgres` user is rejected with a password authentication failure. That means the password stored in the database's data directory no longer matches `POSTGRES_PASSWORD`. The database only reads that variable when it initializes an empty data directory, so redeploying Postgres will not resync it on its own. To bring the two back in line without touching the volume or your data, set `POSTGRES_PASSWORD` on the Postgres service back to the password the database actually holds. That is the value from before the regeneration, which you may still have in an earlier copy of `DATABASE_URL` or a local env file. Because `PGPASSWORD`, `DATABASE_URL` and `DATABASE_PUBLIC_URL` are references to it, they follow automatically. Then redeploy the backend. Once both sides match, Regenerate works normally again, as described in [our database docs](https://docs.railway.com/databases/database-view). To answer your last question: yes, once the variable and the database agree, redeploying the backend is enough. It picks up the referenced `DATABASE_URL` on startup, and nothing else needs to change.

jonatamelodev
HOBBYOP

4 hours ago

Thank you. I still need help because restoring the previous password is not an acceptable recovery path.

The previous credential must be treated as compromised and cannot be placed back into POSTGRES_PASSWORD, DATABASE_URL, a deployment, or any other active configuration.

I need a safe recovery procedure that changes the password stored by the existing PostgreSQL role to match the newly generated POSTGRES_PASSWORD, without:

  • restoring the previous credential;
  • recreating the PostgreSQL service;
  • deleting or replacing the volume;
  • restoring an old backup;
  • modifying application data.

Please escalate this to a Railway employee and provide the supported procedure for safely resynchronizing the existing PostgreSQL role with the current managed password. The production backend remains unavailable with PostgreSQL error 28P01.

If verification requires sensitive values, please do not request that they be posted in this public thread.


Status changed to Awaiting Railway Response Railway • about 3 hours ago


Railway
BOT

3 hours ago

The only supported way to change a database password is the Regenerate action in the database's Config tab, and it only works when the variable and the database already agree. We don't have a supported procedure that rewrites the role's stored password to match the new POSTGRES_PASSWORD from the variable side.

The password the database accepts right now is still the previous one. Regenerating didn't change it inside the database, so that credential has been the live one since the mismatch started.

The documented recovery doesn't leave that credential in place. Set POSTGRES_PASSWORD back to the value the database actually holds, so both sides match, then use Regenerate straight away. That changes the password inside the database and in the variable together, to a new value. The old credential stops working at that point. PGPASSWORD, DATABASE_URL and DATABASE_PUBLIC_URL follow automatically because they reference POSTGRES_PASSWORD. Redeploy the backend after the Regenerate step and it will pick up the new DATABASE_URL. None of this recreates the service or touches the volume or your data.

You don't need to share any password on this thread for this.


Status changed to Awaiting User Response Railway • about 3 hours ago


Railway

The only supported way to change a database password is the Regenerate action in the database's Config tab, and it only works when the variable and the database already agree. We don't have a supported procedure that rewrites the role's stored password to match the new POSTGRES_PASSWORD from the variable side. The password the database accepts right now is still the previous one. Regenerating didn't change it inside the database, so that credential has been the live one since the mismatch started. The documented recovery doesn't leave that credential in place. Set POSTGRES_PASSWORD back to the value the database actually holds, so both sides match, then use Regenerate straight away. That changes the password inside the database and in the variable together, to a new value. The old credential stops working at that point. PGPASSWORD, DATABASE_URL and DATABASE_PUBLIC_URL follow automatically because they reference POSTGRES_PASSWORD. Redeploy the backend after the Regenerate step and it will pick up the new DATABASE_URL. None of this recreates the service or touches the volume or your data. You don't need to share any password on this thread for this.

jonatamelodev
HOBBYOP

3 hours ago

The issue is still unresolved, and new diagnostics show a split state that the automated response does not account for.

Inside the currently running PostgreSQL container:

  • Local Unix-socket authentication succeeds.
  • The postgres role was synchronized to the current POSTGRES_PASSWORD through the local socket.
  • Password-authenticated TCP access to 127.0.0.1 succeeds with the current POSTGRES_PASSWORD.
  • PGPASSWORD matches POSTGRES_PASSWORD.
  • DATABASE_URL contains the current POSTGRES_PASSWORD.

However:

  • DATABASE_URL fails with 28P01.
  • Direct private-network authentication through PGHOST / RAILWAY_PRIVATE_DOMAIN fails with 28P01.
  • PGHOST and RAILWAY_PRIVATE_DOMAIN match.
  • Direct authentication through the public Railway TCP proxy also fails with 28P01.
  • Redeploying PostgreSQL did not change this result.

Therefore, the PostgreSQL process reached through the local socket accepts the current credential, while both Railway network routes reach an endpoint that rejects it.

The previously exposed credential cannot be restored as a normal operating credential. Please escalate this to a Railway employee to inspect which deployment instance and mounted volume are targeted by SSH, private networking, and the public TCP proxy.

Please do not mark this solved or repeat the generic regeneration instructions without addressing the verified split between local and network authentication.


Status changed to Awaiting Railway Response Railway • about 3 hours ago


Railway
BOT

3 hours ago

We checked which instance the network routes reach. The Postgres service has a single running instance, from the redeploy at 19:26 UTC, and it is mounted on the existing volume. The 28P01 rejections from private networking and the TCP proxy are logged by that same Postgres process, at 19:28, 19:38, 19:44 and 19:48 UTC. So the network routes are not reaching a different instance or volume. They reach the same database you are testing inside the container, and that database is rejecting the password.

Each of those rejected connections matched the catch-all host all all all scram-sha-256 rule, line 128 of the pg_hba.conf in your data directory. Unix-socket and 127.0.0.1 connections can match an earlier, more specific line in that same file. If that line does not require a password, a successful local login does not show that the role's stored password equals the current POSTGRES_PASSWORD. Check the lines above 128 to see which rule your local tests are matching.


Status changed to Awaiting User Response Railway • about 3 hours ago


jonatamelodev
HOBBYOP

2 hours ago

Resolved. Production is operational again and all application tests passed.

The effective recovery was:

  1. Access the existing PostgreSQL container through Railway SSH.
  2. Connect through the local Unix socket.
  3. Load the current POSTGRES_PASSWORD directly into a psql variable using \getenv.
  4. Execute ALTER ROLE postgres WITH PASSWORD :'new_password';.
  5. Validate password authentication through Railway private networking.
  6. Validate the PostgreSQL service's DATABASE_URL.
  7. Redeploy the backend.

After PostgreSQL authentication was repaired, the backend's ${{Postgres.DATABASE_URL}} reference still resolved to stale credentials. Replacing the backend DATABASE_URL with direct Railway references to the current Postgres components restored connectivity:

postgresql://${{Postgres.PGUSER}}:${{Postgres.POSTGRES_PASSWORD}}@${{Postgres.RAILWAY_PRIVATE_DOMAIN}}:5432/${{Postgres.PGDATABASE}}

Validation completed successfully:

  • Backend remained online.
  • Login succeeded.
  • Profile data loaded.
  • Friends loaded.
  • Photos loaded.
  • Places loaded.
  • Refresh-token session worked.
  • Session persisted after reopening the application.

No application data, schema, migration, backup, service, or volume was recreated or modified. Only the postgres role password and the backend connection reference were corrected.

Please investigate why the managed Regenerate action left the role password inconsistent and why the backend reference to Postgres.DATABASE_URL remained stale after fresh deployments.


Status changed to Awaiting Railway Response Railway • about 2 hours ago


Status changed to Solved Railway • about 2 hours ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...