24 days ago
Hello Railway Support,
I need assistance determining whether output files from a completed research batch job can be recovered from its original stopped container.
The job completed successfully, but no persistent volume was attached and no external artifact backup was verified. The container is no longer running, so the model and detailed results are currently inaccessible through the container file tools.
Project and deployment details:
Account plan: Pro
Workspace: My Projects
Project: Quant Trading Lab Research
Project ID: c9f905fb-cb85-48d7-857d-06830054c804
Service: three-year-trainer-github
Service ID: 6f46a589-7ff2-410a-95f0-d33c4e7edcfc
Environment: production
Environment ID: 8a1b1114-996f-41d8-96d3-0a64b2adce38
Deployment ID: e88d31e7-3b59-4d68-926c-81cd2ed8883a
Runtime: V2
Region: sfo
What happened:
The Python job trained and evaluated a research model using three years of historical market data. It printed its final RESEARCH_FIT_ONLY summary on 12 September 2026 at 10:32:27 UTC / 13:32:27 Saudi time.
At the latest check, the deployment still showed SUCCESS, but the replica status was 0 running, 0 crashed, and 1 total. No persistent volume was attached.
The container file-access tool returned:
“No running instance found. The service may be offline, sleeping, or still deploying.”
The job was configured to save its outputs under:
/data/models/three_year_research/
The expected files are:
research_model.joblib
out_of_sample_predictions.csv.gz
walk_forward_report.json
model_card.json
net_backtest.json
fit_summary.json
experiment_plan.json
events.csv.gz
The final aggregate summary remains available in the deployment logs, but it cannot replace the trained model or detailed predictions and reports.
Assistance requested
- Please determine whether the original stopped container’s writable filesystem still exists and whether these runtime-generated files are recoverable. They were created after startup, rather than included in the original build image.
- If recovery is possible, please preserve the filesystem and advise on a non-destructive export of the output directory. Recovering the trained model and detailed evaluation results is the priority.
- If recovery is not possible, please explicitly confirm this so I can arrange a controlled rerun with persistent storage and verified artifact backups.
Please do not restart, redeploy, remove, replace, or attach storage to the existing deployment without my explicit approval. I do not want to risk overwriting recoverable data or inadvertently launching the training job again.
Thank you for investigating.
1 Replies
24 days ago
The service has no persistent volume, and no detached volume exists in the project, so the files written during the job lived entirely on the container's ephemeral filesystem. Once the container stopped, that filesystem was discarded and cannot be recovered. For the rerun, attach a volume to the service so the output directory persists independently of the container lifecycle.
Status changed to Awaiting User Response Railway • 24 days ago
17 days ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • 17 days ago