17 days ago
Hello Railway Support,
We are investigating an issue where Railway deployments report SUCCESS, but the expected service does not remain available afterward as a persistent runtime/HTTP instance.
Observed behavior:
the build/deployment completes successfully;
Railway reports the deployment as SUCCESS;
however, no persistent running HTTP/runtime instance remains available afterward;
the behavior has been reproduced independently in two service/deployment paths.
We are currently avoiding further configuration changes while investigating so that we do not alter the evidence.
Could you please check the deployment/runtime lifecycle for this project and help us determine why a deployment can complete successfully without leaving an active persistent runtime/HTTP instance?
We can provide the relevant deployment IDs, service names, timestamps, and sanitized build/deploy logs if needed.
We will not include credentials, secrets, API keys, tokens, or database URLs in this thread.
Thank you.
5 Replies
17 days ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 17 days ago
17 days ago
what do you mean by "no persistent running http runtime instance" - do you mean that you can't access your application? is the deployments HTTP logs seeing any traffic when you try to visit the page?
milo
what do you mean by "no persistent running http runtime instance" - do you mean that you can't access your application? is the deployments HTTP logs seeing any traffic when you try to visit the page?
17 days ago
Thanks for the follow-up.
By “no persistent running HTTP/runtime instance” we mean the following observable behavior:
The relevant deployments reached SUCCESS.
We were not able to obtain an observable response from the service endpoints we exposed, specifically /health and /probe.
For the final deployment of p03-validator-m1, Railway deploy/runtime logs were empty and HTTP logs were empty.
For p03-observer-temp, the deploy log contained only Starting Container; after that, we observed no further runtime output and no HTTP log entries.
During the observed window, CPU and memory metrics for both services remained at 0.
So, from our side, yes: operationally we could not successfully access the application endpoints after deployment.
However, we do not yet know whether that means:
the process exited,
the instance was drained or stopped,
routing never reached the instance,
the service was not actually kept alive,
or another Railway platform/runtime lifecycle condition occurred.
Regarding HTTP logs: they were empty for the relevant deployments. We did not observe HTTP traffic corresponding to our access attempts.
Because the HTTP logs are empty, we also cannot prove from our side that the requests reached the Railway proxy/service layer. This is one of the reasons we are asking for platform-side clarification.
Existing evidence:
p03-validator-m1
Service ID: 3443b5e0-d1c3-48b8-8e83-de553d9fd23a
Deployment: 144cfe5e-20d7-4ab5-80e2-c0ffd749c168
Status: SUCCESS
deploy logs: empty
HTTP logs: empty
observed CPU: 0
observed memory: 0
p03-observer-temp
Service ID: 745fd5b9-c2b1-40b6-a7d7-90622b15a098
Deployment: da8a35d8-2390-4733-8918-619008d95b97
Status: SUCCESS
deploy log: Starting Container
HTTP logs: empty
observed CPU: 0
observed memory: 0
We have intentionally not redeployed or run additional probes since preserving the incident evidence.
If possible, could you check the platform-side lifecycle of these two deployments and tell us whether the containers exited, were stopped/drained, failed to remain scheduled, or never became reachable behind the HTTP proxy?
Please do not modify or restart the services while investigating.
17 days ago
Thanks — we checked the actual Railway service configuration read-only, and this rules out some of the possibilities you mentioned.
Both services are configured as long-lived HTTP servers rather than one-shot scripts.
For p03-validator-m1:
the start command runs code using Bun.serve(...)
it binds to 0.0.0.0
it listens on process.env.PORT (fallback 3000)
Railway has healthcheckPath=/health
Runtime V2, one replica
a Railway service domain exists
For p03-observer-temp:
the start command ends with exec node /tmp/observer.js
that script creates an http.createServer(...)
it binds to 0.0.0.0
it listens on process.env.PORT || 3000
Railway has healthcheckPath=/health
healthcheck timeout is 30 seconds
Runtime V2, one replica
a Railway service domain exists
So these are not intended to run and exit normally after doing a task.
This also means the “no configured healthcheck” explanation does not match the current configuration: /health is configured on both services.
Yet the relevant deployments showed:
SUCCESS
no persistent memory usage
no observable HTTP response
empty HTTP logs
and, for the conventional Node container, only Starting Container in deploy logs.
We have made no changes or redeploys since preserving the incident.
Before we run another diagnostic deployment, could you please check whether Railway has platform-side lifecycle/exit information for these exact deployments, particularly the process exit code or healthcheck lifecycle?
Relevant deployments:
p03-validator-m1
144cfe5e-20d7-4ab5-80e2-c0ffd749c168
p03-observer-temp
da8a35d8-2390-4733-8918-619008d95b97
One additional detail: both Railway-generated domains currently report no explicit target port in the service-domain metadata. Could you confirm whether that is expected when the service uses the injected PORT, or whether a specific target port needs to be set?
We would prefer to understand the recorded lifecycle of these preserved deployments before changing the start command or redeploying.
Please do not modify or restart the services while investigating.
17 days ago
Hi Railway team,
Our workspace has now been upgraded to Pro.
We attempted to create a Private Thread for this same P03 investigation, but Central Station prevents it because this existing Community Thread is detected as an already open duplicate.
We want to preserve the full history and technical evidence already contained here, so we do not want to delete this thread.
Could a Railway team member either:
move/convert this case to a Private Thread, if that is supported, or
tell us the correct way to continue this same case privately without deleting its existing history?
If conversion is not possible, we are also happy to continue the platform-side lifecycle investigation in this existing thread under our Pro support entitlement.
Please do not modify, restart or redeploy any services while investigating.
Thank you.
Status changed to Open 0x5b62656e5d • 14 days ago
costabrava010818
Hi Railway team, Our workspace has now been upgraded to Pro. We attempted to create a Private Thread for this same P03 investigation, but Central Station prevents it because this existing Community Thread is detected as an already open duplicate. We want to preserve the full history and technical evidence already contained here, so we do not want to delete this thread. Could a Railway team member either: move/convert this case to a Private Thread, if that is supported, or tell us the correct way to continue this same case privately without deleting its existing history? If conversion is not possible, we are also happy to continue the platform-side lifecycle investigation in this existing thread under our Pro support entitlement. Please do not modify, restart or redeploy any services while investigating. Thank you.
14 days ago
The thread will be kept public, as there is no sensitive data.
Also, can you make sure you’re running the correct image? Please share a screenshot of the deployments tab of your service.