Note

Batch size and memory limit are the same setting

An inference service that occasionally batches large requests will evict itself, and the eviction looks like a mystery.

Memory limits describe the peak, not the average. An inference container sitting comfortably at 400 MB most of the time and allocating 2 GB when a large batch arrives is not a 400 MB container.

The eviction is confusing because the dashboard shows healthy average usage. Average usage is not what the limit is compared against.

Either cap the batch size in the service so the peak is bounded, or size the limit for the real peak. Doing neither means the container works until traffic shape changes. Doing both is fine and is what I would now default to.