Infrastructure

Multi-AZ RDS is not automatically correct

It roughly doubles your database cost to buy a synchronous standby. Sometimes that is exactly the right purchase. Deciding it by default is how architecture reviews stop being reviews.

The reflex answer to "make the database highly available" is multi-AZ. It is often correct. Applying it without a stated reason is a habit, not a design.

What you are actually buying

A synchronous standby in a second availability zone, and automated failover to it. The write path acknowledges after the standby has the data, so a zone failure does not lose committed transactions. Recovery is minutes, without human involvement.

The cost is roughly double, because you are paying for a second instance that serves no read traffic. That is not a criticism, it is the deal: you are buying availability, not capacity.

The question that should come first

Not "should this be highly available", which nobody ever answers no to. Instead: what is the recovery time objective, what is the recovery point objective, and what does an hour of unavailability actually cost this workload?

For a payment path, an internal tool that the whole company blocks on, or anything with a contractual availability commitment, multi-AZ is straightforwardly correct and the cost conversation is short.

For a good number of workloads I have seen it applied to, the honest answers are that a thirty-minute recovery is tolerable, losing five minutes of data is tolerable, and an hour of downtime costs some inconvenience. Those answers point at single-AZ with a tested restore, and the difference is a meaningful share of a small project budget.

"Tested restore" is the load-bearing phrase

Choosing single-AZ obliges you to something, and this is where the decision usually goes wrong. Not choosing multi-AZ is fine. Not choosing multi-AZ and also never testing a restore is how you find out during an incident that your recovery time objective was fiction.

Testing means performing the restore. Restore the most recent snapshot into a fresh instance, time it, point a non-production application at it, confirm the data is intact and current enough. Write down the number you measured.

An untested backup is a hypothesis. Snapshot age, restore duration and the state of the application after a restore are all things that surprise people, and they surprise them at the worst possible time.

Write the decision down

The specific practice that changed how these conversations went at Adex: every availability decision was recorded as a short entry naming the failure domain, the chosen response, and the cost.

Zone failure. Response: single-AZ with snapshot restore, measured at roughly twenty minutes. Cost avoided: the standby instance. Accepted by: named person, on this date.

Three lines. What they buy is that the decision stops being re-litigated every time someone new looks at the architecture, and it becomes possible to disagree with it specifically. "Why is this single-AZ" has an answer. If the workload becomes more important later, there is a written thing to revisit rather than an assumption to excavate.

Undocumented defaults cannot be reviewed, which means an architecture review of undocumented defaults is not a review. It is a reading.

Where I would spend the money instead

On a workload where the recovery objectives genuinely tolerate a restore window, the standby cost buys more elsewhere.

Automated snapshot verification, so a broken backup is discovered on a schedule instead of during an incident. Right-sizing the primary against observed utilisation, which at Adex was consistently the largest available saving and required no architectural change at all. Scheduling non-production environments to stop outside working hours, which is the least clever and most reliable saving available.

The general shape: high availability is one lever among several, it is expensive, and pulling it by default means the other levers never get considered.

Continue reading