A model that reached production through a notebook has no history you can reconstruct. The dataset may have moved. The notebook has been edited since. The hyperparameters were in a cell that got overwritten. The model works, and you cannot say why or reproduce it.
In an ordinary setting that is untidy. In a healthcare-compliant environment it is a control failure, because a model informing anything consequential has to be reconstructible.
Banning notebooks does not work
The first instinct is to prohibit them. This fails, because exploratory data work genuinely belongs in a notebook. Interactive, iterative, visual, disposable: those are the right properties for figuring out whether an idea has legs. Take the notebook away and you have not improved provenance, you have made exploration painful and pushed it somewhere less visible.
The line that held was narrower. Notebooks may explore. Notebooks may not produce anything deployable. Everything on the path to an endpoint goes through a pipeline.
What the pipeline has to do
On SageMaker this became four steps, and the ordering matters more than the tooling.
A processing step that prepares data from a versioned source and writes the prepared dataset as a concrete artefact. This is the step people skip, and skipping it is what makes runs irreproducible. If training reads data that some code produced a moment ago in the same process, the input to training is not a thing you can point at. Once preparation writes to a pinned object-store prefix, it is.
A training step whose parameters are pipeline inputs. Not constants in the script, not values in a config file that gets edited in place. Inputs, recorded per execution.
An evaluation step producing metrics as structured output.
A conditional registration step that registers a model version only if those metrics clear a threshold. This one sounds too obvious to state and is easy to omit. Omit it and a training run that produced a bad model registers it anyway, where it sits looking exactly like a good one, waiting for somebody to approve it by mistake.
The registry is the only door
Deployment reads from the model registry. An endpoint is updated from an approved registered version, never from a training job output directly.
That closes the shortcut everyone takes under time pressure, which is to grab the artefact from the training job and push it. It also gives the audit trail a human decision point, because approval is a deliberate state transition with a name and a timestamp attached.
Each registered version carries the pipeline execution that produced it, the dataset prefix, the training image and the evaluation metrics. Reconstructing a model months later is reading a record, not an investigation.
Separate the roles
Three IAM roles rather than one broad execution role: training, registration, deployment.
The training role can read its dataset prefix and write artefacts. It cannot approve a model version and it cannot update an endpoint. That is more configuration, and it means a compromised training job cannot put a model into production.
This is the same reasoning as any other least-privilege argument, and it is easier to win here because the lifecycle stages are already distinct in everyone's mental model. You are encoding a separation people already believe in.
Pinning data is the hard half
Pinning code is solved. Git exists.
Pinning data requires a convention and the discipline to hold it. Versioned dataset prefixes in object storage, and never, under any circumstance, training against a mutable location. The moment one pipeline reads from a path that something else writes to, reproducibility is gone and nothing will tell you.
Most of the effort in this work was not the pipeline definitions. It was establishing that convention and then noticing the three places that quietly violated it.