ProductionCloud architecture2024

Highly available AWS reference architectures

Terraform-described, cost-conscious AWS architectures with least-privilege IAM and segmented networks, built during the Adex apprenticeship.

Why it existsAvailability, cost and blast radius are the same conversation. Designing one without the others produces an architecture that only works on a slide.

Overview

Apprenticeship work at Adex International designing highly available AWS architectures and describing them in Terraform. This is the work that led directly to ARCH, because the gap between the diagram and the Terraform was visible on every engagement.

The architectures were conventional by design. The value was in the trade-off reasoning and in the fact that the result was reproducible.

Problem

Three recurring failures in cloud architecture work.

  • Availability treated as a feature bought by duplicating everything, without regard to what actually fails or what the duplication costs.
  • Permissions granted broadly at build time because narrowing them later is nobody's task. The wildcard policy that gets things working never gets revisited.
  • Networks flat inside the perimeter. A firewall at the edge and unrestricted movement behind it.

None of these are exotic. They are what happens when infrastructure is assembled under time pressure without a written rationale.

Approach

Design against explicit failure domains, provision with Terraform, and grant permissions from the workload backwards.

Availability decisions start from what fails: an availability zone, an instance, a database primary. Each gets a stated response. Where the response is expensive, that is a recorded decision with a cost attached rather than an unexamined default.

Every architecture is Terraform from the start. Not documented and later codified, described in code as the primary artefact.

IAM policies are written from the workload's actual API calls rather than from a service-level wildcard. Slower to author, and it means the policy describes the workload.

Architecture

Network

VPC with public and private subnets spread across at least two availability zones. Public subnets carry only the load balancer and NAT. Application and data tiers sit in private subnets with no route to an internet gateway. Security groups reference other security groups rather than CIDR ranges wherever possible, so the rules describe relationships instead of addresses.

Compute

EC2 in autoscaling groups behind an Application Load Balancer, spread across availability zones. Health checks at the target group, so a failed instance is replaced rather than kept in rotation.

Data

Amazon RDS with multi-AZ where the recovery objective justified the cost, single-AZ with tested snapshot restore where it did not. That choice was made per workload and written down.

Storage and identity

S3 with public access blocked at the account level, encryption enforced, and bucket policies limited to the roles that need them. IAM roles per workload, no long-lived access keys on instances.

Observability

CloudWatch metrics and alarms on the signals that indicate the failure modes identified in the design, rather than on a default dashboard.

Technology

Terraform for provisioning. AWS: EC2, VPC, IAM, RDS, S3, Application Load Balancer, CloudWatch. Route 53 for DNS. CloudTrail for API audit and GuardDuty for threat detection on the accounts that carried anything sensitive.

Implementation

Terraform was structured as composable modules, network, compute, data, rather than one root configuration per environment. Environments differ by variable file. This is the structure that later became the module layout ARCH generates.

Cost work was mostly right-sizing against observed utilisation instead of requested capacity, plus scheduling non-production environments to stop outside working hours. Neither is clever. Both were consistently the largest available saving.

Security group rules reference source security groups rather than CIDR blocks. The rule then reads as "the application tier may reach the database tier", which survives a network change that a hardcoded CIDR does not.

IAM policies were narrowed by starting from a deny-by-default position and adding actions as the workload demonstrated it needed them, with CloudTrail as the record of what was actually called.

Challenges

Least privilege is a scheduling problem. Everyone agrees with it and it loses to deadlines, because a broad policy works immediately and a narrow one takes an afternoon. Making the narrow policy part of the initial build, not a follow-up ticket, was the only approach that held.

Multi-AZ RDS is not automatically correct. It roughly doubles the database cost for a synchronous standby. For workloads that could tolerate a restore window, single-AZ with a tested restore procedure was the better trade. Making that a written decision stopped it being re-argued.

Terraform state discipline. Shared state without locking produced exactly the conflict everyone warns about. Remote state with locking, separated per environment, from the beginning.

Decisions

Terraform first, not Terraform eventually. Infrastructure created in the console and codified later is never faithfully codified.

Security groups referencing security groups. Rules that describe relationships instead of addresses stay correct through network changes.

Availability per failure domain with a stated cost. Turns "make it highly available" into a decision someone can disagree with, which is the point.

No long-lived keys on compute. Instance roles only. A leaked key with no expiry is a different class of incident.

Results

Architectures delivered as Terraform modules with documented availability, cost and access-control decisions. Right-sizing and non-production scheduling produced the bulk of the cost reduction. IAM was narrowed from service-level grants to workload-derived policies, and network segmentation replaced flat internal connectivity.

This apprenticeship is also where the AWS Solutions Architect Associate preparation stopped being abstract, and where ARCH started as a note about the diagram-to-Terraform gap.

Stack

  1. Network

    • Multi-AZ VPC
    • Private application and data tiers
    • Security groups referencing security groups
    • Route 53
  2. Compute and data

    • EC2 autoscaling behind an ALB
    • Target group health checks
    • RDS multi-AZ where justified
    • Tested snapshot restore where not
  3. Identity

    • Workload-derived IAM policies
    • Instance roles, no long-lived keys
    • Account-level S3 public access block
  4. Operations

    • Terraform modules per concern
    • Remote state with locking
    • CloudWatch alarms on identified failure modes
    • CloudTrail and GuardDuty

Measurements

Availability floor
2 AZsPer failure-domain decision, with cost stated
Provisioning
TerraformPrimary artefact, not documentation
Largest cost lever
Right-sizingPlus non-production scheduling