Engineering / machine-learning

Machine learning operations

A model endpoint is a production system with a second correctness problem: it can be healthy, fast, and wrong.

The lifecycle is the product

Training, evaluation, packaging, serving, monitoring, and retraining form one system. A notebook that produces a model is only the beginning; the useful question is whether the deployed artifact can be identified and its behaviour compared over time.

  • Track data and model versions beside code revisions.
  • Separate infrastructure health from prediction quality and data drift.
  • Make a failed or rolled-back model version as observable as a failed container.

Serving on Kubernetes

K3s and Kubernetes make scheduling and isolation available at different scales, but they do not decide what should be measured. Request latency, saturation, error rate, and prediction distribution should meet the same operational standard as any other service.

A 200 response proves transport. It does not prove inference quality.

Responsible iteration

Retraining should be triggered by a clear signal and reviewed like a software change. If a model cannot be explained to the person operating the service, automation has only moved the uncertainty somewhere harder to see.

Machine learning operations — Gokul Upadhyay Guragain