Delivery

The pipeline is the control

In a regulated environment a CI pipeline is not automation that happens to be convenient. It is the evidence that software reached production legitimately, and evidence has to outlive a log retention window.

The first thing I got wrong on a HITRUST-regulated platform was assuming the pipeline already did its job because deployments worked.

They did work. Code merged, images built, scans ran, the application landed in production. By any ordinary engineering measure the delivery path was healthy. Then someone asked a question that the pipeline could not answer: on the fourteenth of last month, which image was promoted to production, which scans passed before it went, what did those scans find, and who approved the promotion?

The pipeline had done all of those things. It had recorded none of them anywhere that still existed.

Logs are not evidence

CI job logs are written for a person debugging a failure in the next twenty minutes. They are verbose, unstructured, and subject to a retention policy that exists to control storage cost. Every property that makes a log useful for debugging makes it useless as a record.

Evidence has the opposite requirements. It should be structured, minimal, addressed to a specific commit, and retained on the same schedule as the thing it describes. If a release lives for two years, the record of how it got to production lives for two years.

The reframing that fixed this: a stage that corresponds to a control must emit an artefact, not a log line. Scanning is not "we ran a scanner and you can see it in the output". Scanning is "here is a document stating which scanner, which version, against which image digest, and these were the findings by severity".

What goes in the bundle

We collected one structured record per pipeline run.

  • The commit SHA.
  • The image digest, never the tag. A tag can be moved after the fact; a digest cannot. This single substitution removed an entire category of "which build is actually running" ambiguity.
  • Tool identity and version for every scanner. A finding means nothing without knowing what looked for it.
  • Findings by severity, in full.
  • Findings that were accepted rather than fixed, with the acceptance recorded alongside.
  • Test results.
  • The identity that approved the promotion.

That last pair matters more than the rest. Every organisation has findings it cannot fix inside a release window. If there is no recorded path for accepting a risk, engineers route around the scan, because the scan is now an obstacle to shipping rather than a source of information. Give acceptance a home and it becomes reviewable. Leave it homeless and your scanning is theatre.

The gate has to be able to say no

Collecting evidence is only half of it. The production deployment stage was changed to refuse to start unless the evidence bundle for that commit exists and is complete.

This is the part that converts a convention into a control. Before, a pipeline that skipped a scan still deployed, and the skip was discoverable only if someone went looking. After, a pipeline that skipped a scan cannot reach production, and it does not matter whether the skip was a misconfiguration or a deliberate shortcut under deadline pressure.

Evidence costs latency, so buy width instead of length

Every gate you add lengthens the path to production, and a delivery path that becomes painful gets circumvented. That tension is real and it does not resolve itself through good intentions.

The answer was structural rather than cultural. Evidence-producing stages have no dependency on each other: a dependency scan does not need the container scan to finish, and neither needs the infrastructure configuration check. Running them in parallel meant that adding a control made the pipeline wider rather than slower. Wall-clock time to production stayed roughly where it was while the number of controls went up.

That is the sentence I would keep if I could only keep one. When compliance work makes delivery slower, look for stages you have accidentally serialised before you start negotiating the controls away.

Vocabulary is half the work

A genuinely large share of this was translation rather than engineering. A control document says change management. The team says open a merge request and get a review. Those are the same activity, and until somebody writes the mapping down, the same argument recurs every week with both sides convinced the other is being obstructive.

We wrote the mapping down: this control, this pipeline stage, this artefact. It stopped the weekly argument, and it made gaps visible, because a control with no corresponding stage had nowhere to point.

What I would tell someone starting this

Do not begin by adding gates. Begin by asking your existing pipeline the audit question: for a specific deployment last month, prove how it got there. Whatever you cannot answer is your actual backlog, and it is usually shorter than it feels.

Continue reading