All infographics

VIRTUAL COMPUTE / 04

Amazon EC2

Running virtual servers in the cloud

Download sheet SVG

THE BIG PICTURE

A scalable application fleet

  1. ClientsApplication traffic
  2. Load balancerDistribute requests
  3. EC2 fleetAuto Scaling group
EBS volumesAttached to each instance

An example architecture: a load balancer distributes traffic to an Auto Scaling fleet, with EBS storage attached per instance.

01Instance types and sizing

  • Letter sets the purpose: M general, C compute, R memory, I storage, P and G GPU
  • Suffixes add traits: g is Graviton (Arm), a is AMD, d adds local NVMe instance store
  • Each size step roughly doubles vCPU and memory; metal sizes give the whole host
  • T types are burstable: they earn CPU credits below baseline and spend them above it
  • Standard mode throttles once credits run out; watch CPUCreditBalance

02AMIs and launch templates

  • An AMI holds volume snapshots, launch permissions and the block device mapping
  • AMIs are regional; copy one to each region you launch in
  • Deregistering an AMI keeps snapshots by default; optionally delete them in the same request
  • Launch templates version AMI, type, security groups, storage and user data together
  • On Linux, cloud-init runs user data once at first boot; never put secrets in it

03EBS volume types

  • gp3 is the usual SSD, with IOPS and throughput set apart from size; prefer it to gp2
  • io2 and io2 Block Express give provisioned IOPS for latency-sensitive databases
  • st1 and sc1 are HDD types for large sequential reads; neither can be a boot volume
  • A volume lives in one AZ; move data across AZs or regions with snapshots
  • Type, size and IOPS can change online; the filesystem must be extended separately

04Instance store

  • Instance store is disk physically on the host; it is not network attached
  • Data survives a reboot but is lost on stop, terminate, or underlying hardware failure
  • Resizing requires a stop, which erases instance store contents; copy data off first
  • Use it for caches, scratch space and data replicated elsewhere, never as the only copy
  • Only some instance types include it; check the storage details before you choose

05Pricing and purchasing

  • On-Demand has no term; most OS options bill per second with a 60-second minimum; SLES hourly
  • A stopped instance skips compute charges; its EBS volumes and any Elastic IP still bill
  • Savings Plans and Reserved Instances trade a one or three year term for a lower rate
  • Spot uses spare capacity at a discount and can be reclaimed on a two-minute notice
  • Capacity Reservations hold one-AZ capacity and bill unused slots; future-dated ones require a term

06Placement groups

  • Cluster groups pack instances together in one AZ for low latency and high throughput
  • Spread groups use separate hardware; seven running instances per AZ maximum
  • Partition groups put instances in separate rack partitions, for large systems like Kafka
  • Launch cluster members in one request to reduce insufficient-capacity errors
  • An instance joins a placement group at launch; moving it later requires it stopped

07Networking and public IPs

  • Each instance keeps a primary ENI; how many extra ENIs it takes depends on type
  • Security groups are stateful; new custom groups start with no inbound rules and allow outbound
  • A security group can name another security group as a source instead of an IP range
  • Elastic IPs keep their address across stop and start; auto-assigned public IPs do not
  • Public IPv4 addresses bill per hour, attached or not, so release unused Elastic IPs

08Instance metadata (IMDSv2)

  • Metadata lives at 169.254.169.254 and serves instance identity and role credentials
  • IMDSv2 needs a session token: PUT to /latest/api/token, then send it on each GET
  • Set HttpTokens to required to refuse IMDSv1 requests on an instance
  • Containers need an extra hop; raise the metadata hop limit if they cannot reach it
  • Use SDK credential chains rather than reading raw metadata URLs in application code

09Access: keys, roles, SSM

  • AWS keeps only the public half of a key pair; a lost private key cannot be recovered
  • An instance profile attaches an IAM role; SDKs get rotating temporary credentials
  • Do not store long-lived access keys on instances; the attached role replaces them
  • Session Manager opens a shell through the SSM agent, so port 22 need not be open
  • Session Manager needs role permissions; session logs can go to S3 or CloudWatch

10Scaling and load balancing

  • Auto Scaling groups hold capacity between min and max and replace failed instances
  • Target tracking policies hold a metric, such as average CPU, near a chosen target
  • Spread a group across AZs; a mixed-instances policy can blend Spot and On-Demand
  • ALB routes HTTP and HTTPS at layer 7; NLB routes TCP, UDP and TLS at layer 4
  • Enable ELB health checks on the group so failed LB checks also replace instances

11CloudWatch monitoring

  • Basic metrics come every 5 minutes; detailed monitoring is 1 minute and costs extra
  • Memory, disk and process metrics are not built in; install the CloudWatch agent
  • StatusCheckFailed metrics split host faults from instance faults; alarm on both
  • A system check alarm can trigger the EC2 recover action on supported instance types

12Common pitfalls

  • Unattached EBS volumes and old snapshots keep billing after the instance is gone
  • Security groups open to 0.0.0.0/0 on ports 22 or 3389 invite brute-force logins
  • Cross-AZ and internet transfer bill per GB; check transfer lines in cost reports
  • Root volumes are deleted on termination by default; data volumes are not
  • Termination protection is off by default; while on, it blocks terminate calls

Go to the source

Use AWS documentation for current limits, availability, and pricing.

EC2 On-Demand pricing EC2 Capacity Reservations Deregister an EC2 AMI VPC security groups