ProductionMLOps2026

Model lifecycle on AWS SageMaker

Training, registry and deployment workflows on SageMaker for a healthcare-compliant environment, where a model version has to be traceable to its data.

Why it existsA model in a regulated setting needs the same provenance as a release: which data, which code, which parameters, approved by whom.

Overview

MLOps work at acAIberry, building out model lifecycle workflows on AWS SageMaker inside an environment with healthcare data-handling requirements.

The distinguishing constraint was provenance. Deploying a model that performs well is straightforward. Deploying a model and being able to state, months later, exactly which dataset version, which training code and which hyperparameters produced it is a different engineering problem.

Problem

A model that reaches an endpoint through a notebook has no reconstructible history. The dataset it saw may have moved, the notebook has been edited, and the parameters live in a cell that has been overwritten.

In a healthcare-compliant environment that is not merely untidy. If a model informs anything consequential, the inability to reproduce it is a control failure. The requirement was to make training a pipeline execution rather than a human session, and to make every registered model version carry its own lineage.

Approach

Move training out of notebooks and into SageMaker Pipelines, and make the model registry the only route to an endpoint.

Training becomes a pipeline: a processing step that prepares data from a versioned source, a training step with parameters supplied as pipeline inputs rather than embedded in code, an evaluation step, and a conditional registration step that only registers a model version if evaluation metrics clear a threshold.

Deployment reads from the registry. An endpoint is updated from an approved registered model version, never from a training job output directly. Approval is a deliberate state transition, which gives the audit trail a human decision point.

Architecture

Training pipeline

SageMaker Pipelines with discrete steps: data processing, training, evaluation, conditional registration. Parameters are pipeline inputs. Dataset location is a versioned object store prefix, pinned per execution.

Registry

SageMaker Model Registry as the single gate. Each version records the pipeline execution that produced it, the dataset prefix, the training image and the evaluation metrics. Approval status controls deployability.

Serving

Endpoints updated from approved registry versions. Deployment is a pipeline action, not a console action.

Observability

Endpoint invocation and latency metrics into CloudWatch, with the same alerting path the platform used for its ordinary services. Training job failures surface as pipeline failures rather than as an absence of output.

Technology

AWS SageMaker for training, pipelines, registry and hosted endpoints. Amazon S3 for versioned datasets and artefacts. IAM for role separation between training, registration and deployment. CloudWatch for metrics and alerting. Docker for the training and inference container images. Python for pipeline definitions and step code.

Implementation

Data preparation was pulled out of the training script into its own processing step. That separation is what makes a training run reproducible, because the input to training becomes a concrete artefact in object storage rather than the result of code that ran a moment earlier.

Evaluation gates registration. A training job that produces a model failing the metric threshold completes successfully and registers nothing. This sounds obvious and is easy to omit, and omitting it means a bad model version sits in the registry waiting for someone to approve it by accident.

Role separation was enforced with distinct IAM roles for training, registration and deployment, rather than one broad execution role. The training role can read the dataset prefix and write artefacts. It cannot approve a model version or update an endpoint.

Challenges

Reproducibility is mostly about inputs. Pinning code was the easy half. Pinning data required a convention for versioned dataset prefixes and the discipline to never train against a mutable location.

Notebooks are hard to give up. Exploratory work genuinely belongs in a notebook. The line that held was that notebooks may explore and may not produce anything deployable. Everything on the path to an endpoint goes through a pipeline.

Container image size. Training images accreted dependencies quickly, and image pull time became a visible share of short training jobs. A stripped base image and a separate inference image with only serving dependencies addressed it.

Decisions

Registry as the only path to an endpoint. It closes the shortcut from training job straight to production, which is the shortcut everyone takes under time pressure.

Conditional registration on evaluation metrics. The gate belongs in the pipeline, where it is mechanical, not in a review checklist, where it is optional.

Separate IAM roles per lifecycle stage. More configuration, and it means a compromised training job cannot deploy a model.

Results

Training runs as a reproducible pipeline execution with pinned data, parameterised inputs and metric-gated registration. Deployments read approved registry versions, and each version carries the lineage needed to reconstruct it. Endpoint metrics flow into the same CloudWatch alerting the platform already used.

Scope note: this was a one-month traineeship contribution to an existing platform, focused on the lifecycle workflow rather than on model development.

Stack

  1. Training

    • SageMaker Pipelines
    • Separate processing step
    • Parameterised inputs
    • Versioned S3 dataset prefixes
  2. Governance

    • Model Registry as sole gate
    • Metric-gated registration
    • Explicit approval transition
    • Per-stage IAM roles
  3. Serving

    • Endpoints from approved versions
    • Dedicated inference image
    • Pipeline-driven deployment
  4. Operations

    • CloudWatch metrics
    • Shared alerting path
    • Pipeline failure surfacing

Measurements

Training entrypoint
PipelineNotebooks cannot produce deployables
Registration
Metric-gatedConditional step in the pipeline
IAM roles
SeparatedTrain, register and deploy are distinct