Recovery request — trained model artifacts inaccessible after batch job completed without a persistent volume
taroutson2002
PROOP

24 days ago

Hello Railway Support,

I need assistance determining whether output files from a completed research batch job can be recovered from its original stopped container.

The job completed successfully, but no persistent volume was attached and no external artifact backup was verified. The container is no longer running, so the model and detailed results are currently inaccessible through the container file tools.

Project and deployment details:

Account plan: Pro

Workspace: My Projects

Project: Quant Trading Lab Research

Project ID: c9f905fb-cb85-48d7-857d-06830054c804

Service: three-year-trainer-github

Service ID: 6f46a589-7ff2-410a-95f0-d33c4e7edcfc

Environment: production

Environment ID: 8a1b1114-996f-41d8-96d3-0a64b2adce38

Deployment ID: e88d31e7-3b59-4d68-926c-81cd2ed8883a

Runtime: V2

Region: sfo

What happened:

The Python job trained and evaluated a research model using three years of historical market data. It printed its final RESEARCH_FIT_ONLY summary on 12 September 2026 at 10:32:27 UTC / 13:32:27 Saudi time.

At the latest check, the deployment still showed SUCCESS, but the replica status was 0 running, 0 crashed, and 1 total. No persistent volume was attached.

The container file-access tool returned:

“No running instance found. The service may be offline, sleeping, or still deploying.”

The job was configured to save its outputs under:

/data/models/three_year_research/

The expected files are:

research_model.joblib

out_of_sample_predictions.csv.gz

walk_forward_report.json

model_card.json

net_backtest.json

fit_summary.json

experiment_plan.json

events.csv.gz

The final aggregate summary remains available in the deployment logs, but it cannot replace the trained model or detailed predictions and reports.

Assistance requested

  1. Please determine whether the original stopped container’s writable filesystem still exists and whether these runtime-generated files are recoverable. They were created after startup, rather than included in the original build image.
  2. If recovery is possible, please preserve the filesystem and advise on a non-destructive export of the output directory. Recovering the trained model and detailed evaluation results is the priority.
  3. If recovery is not possible, please explicitly confirm this so I can arrange a controlled rerun with persistent storage and verified artifact backups.

Please do not restart, redeploy, remove, replace, or attach storage to the existing deployment without my explicit approval. I do not want to risk overwriting recoverable data or inadvertently launching the training job again.

Thank you for investigating.

Solved

1 Replies

Railway
BOT

24 days ago

The service has no persistent volume, and no detached volume exists in the project, so the files written during the job lived entirely on the container's ephemeral filesystem. Once the container stopped, that filesystem was discarded and cannot be recovered. For the rerun, attach a volume to the service so the output directory persists independently of the container lifecycle.


Status changed to Awaiting User Response Railway • 24 days ago


Railway
BOT

17 days ago

This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!

Status changed to Solved Railway • 17 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...