Infrastructure

K3s is small, not simple

Choosing K3s over managed Kubernetes cut the operational surface substantially and cut the conceptual surface not at all. That distinction is worth understanding before you make the same choice.

K3s was the right decision for the inference workloads at Fusemachines. It was right for reasons I did not fully understand when I made it, and it left behind a category of work I had expected it to remove.

What it genuinely removes

The footprint is real. A single binary, a lightweight datastore option, no separate etcd cluster to operate, and a control plane that runs comfortably on hardware that would be embarrassing for a full Kubernetes deployment. For a small team serving a handful of models, that is a large reduction in the surface you are responsible for.

It also removes a decision. Managed Kubernetes gives you control-plane features you may never use, priced accordingly. If your workload is a set of stateless inference deployments behind a service, most of that is inventory rather than capability.

What it does not remove

Kubernetes concepts. All of them.

Pods still get evicted. Resource requests and limits still need to be right, and getting them wrong still produces the same confusing failures. Readiness and liveness probes still need to distinguish between starting up and being broken, and an inference container that loads a large model at startup will still be killed by a liveness probe with an impatient initial delay. Scheduling still has to be reasoned about.

The single sharpest thing I learned: an inference container that occasionally allocates a large batch will get itself evicted, and the eviction will look like a mystery until you understand that the limit describes the peak rather than the average. Batch size and memory limit are the same setting expressed twice, and they have to agree.

None of that is a K3s problem. It is a Kubernetes problem, and K3s is Kubernetes.

The trap in the framing

"Lightweight Kubernetes" gets read as "Kubernetes without the hard parts". It is not. It is Kubernetes with a smaller operational bill.

If a team picks K3s to avoid learning Kubernetes, they have chosen to learn Kubernetes anyway, in production, without the safety rails a managed control plane would have provided. That is a worse position than either alternative.

The correct reason to pick it is the one that applied to us: the team already understood the concepts, the workload did not need managed control-plane features, and the operational bill was worth reducing.

What made it work

Two things, neither of them K3s-specific.

Helm for the manifests, so an environment difference is a values file rather than a divergent copy of a YAML tree. The moment you have two environments and no packaging, they start drifting, and the drift is discovered during an incident.

Image digests rather than tags in the deployments. A tag can be moved, and a deployment referencing a moved tag is a deployment whose contents you cannot state with confidence. Digests make "what is running" answerable.

The recommendation

K3s is a good choice when the cluster is small, the team already has Kubernetes fluency, and reduced operational surface has real value.

It is a poor choice when it is being selected as a way around learning Kubernetes, or when the workload actually needs the managed features. Small is not simple, and the two get conflated in exactly the situations where the distinction matters most.

Continue reading