a month ago
Service 0a02ddfc-5913-4e84-961f-dc8888af5ca8, project
49e739bd-587c-4d35-8be7-2c985e1018c8. Single replica, 8 GB, Singapore.
My service streams large file downloads. A single job runs for roughly
five minutes and holds a few hundred MB while it does, then returns to
baseline. Several concurrent jobs would multiply that.
I want to confirm production memory against measurements taken in a
local harness, so I need to know what the Metrics view can actually
show me.
-
What sampling interval does the memory graph use, and is it the same
at every time range or does it aggregate as the window widens?
-
If the graph aggregates, is it showing peak or average within each
bucket? A two-minute spike that gets averaged into a wider bucket
would be invisible, which is exactly the event I care about.
-
Is there any way to see the peak memory a replica reached, rather
than a sampled series? Even a single high-water figure per deployment
would be more useful to me than a smoothed graph.
-
If a container is killed for exceeding its memory limit, what appears
in the deploy log, and is the memory reading at the moment of the
kill retained anywhere?
I am asking because I would rather rely on your metrics than build my
own instrumentation, but only if they can resolve an event of this
duration.
1 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 1 month ago
a month ago
Thanks, railway metrics - raw is what I needed for the sampling question.
Question 4 is still open though: if a container is killed for exceeding its memory limit, what appears in the deploy log, and is the memory reading at the moment of the kill retained anywhere? That is the one I most need, because it decides whether an out-of-memory event leaves any evidence at all or whether I only find out from the absence of something.
Status changed to Solved imchikachirag • about 1 month ago