VIRTUAL COMPUTE / 04
Amazon EC2
Running virtual servers in the cloud
THE BIG PICTURE
A scalable application fleet
- ClientsApplication traffic
- Load balancerDistribute requests
- EC2 fleetAuto Scaling group
EBS volumesAttached to each instance
An example architecture: a load balancer distributes traffic to an Auto Scaling fleet, with EBS storage attached per instance.
01Instance types and sizing
- Letter sets the purpose: M general, C compute, R memory, I storage, P and G GPU
- Suffixes add traits: g is Graviton (Arm), a is AMD, d adds local NVMe instance store
- Each size step roughly doubles vCPU and memory; metal sizes give the whole host
- T types are burstable: they earn CPU credits below baseline and spend them above it
- Standard mode throttles once credits run out; watch CPUCreditBalance
02AMIs and launch templates
- An AMI holds volume snapshots, launch permissions and the block device mapping
- AMIs are regional; copy one to each region you launch in
- Deregistering an AMI keeps snapshots by default; optionally delete them in the same request
- Launch templates version AMI, type, security groups, storage and user data together
- On Linux, cloud-init runs user data once at first boot; never put secrets in it
03EBS volume types
- gp3 is the usual SSD, with IOPS and throughput set apart from size; prefer it to gp2
- io2 and io2 Block Express give provisioned IOPS for latency-sensitive databases
- st1 and sc1 are HDD types for large sequential reads; neither can be a boot volume
- A volume lives in one AZ; move data across AZs or regions with snapshots
- Type, size and IOPS can change online; the filesystem must be extended separately
04Instance store
- Instance store is disk physically on the host; it is not network attached
- Data survives a reboot but is lost on stop, terminate, or underlying hardware failure
- Resizing requires a stop, which erases instance store contents; copy data off first
- Use it for caches, scratch space and data replicated elsewhere, never as the only copy
- Only some instance types include it; check the storage details before you choose
05Pricing and purchasing
- On-Demand has no term; most OS options bill per second with a 60-second minimum; SLES hourly
- A stopped instance skips compute charges; its EBS volumes and any Elastic IP still bill
- Savings Plans and Reserved Instances trade a one or three year term for a lower rate
- Spot uses spare capacity at a discount and can be reclaimed on a two-minute notice
- Capacity Reservations hold one-AZ capacity and bill unused slots; future-dated ones require a term
06Placement groups
- Cluster groups pack instances together in one AZ for low latency and high throughput
- Spread groups use separate hardware; seven running instances per AZ maximum
- Partition groups put instances in separate rack partitions, for large systems like Kafka
- Launch cluster members in one request to reduce insufficient-capacity errors
- An instance joins a placement group at launch; moving it later requires it stopped
07Networking and public IPs
- Each instance keeps a primary ENI; how many extra ENIs it takes depends on type
- Security groups are stateful; new custom groups start with no inbound rules and allow outbound
- A security group can name another security group as a source instead of an IP range
- Elastic IPs keep their address across stop and start; auto-assigned public IPs do not
- Public IPv4 addresses bill per hour, attached or not, so release unused Elastic IPs
08Instance metadata (IMDSv2)
- Metadata lives at 169.254.169.254 and serves instance identity and role credentials
- IMDSv2 needs a session token: PUT to /latest/api/token, then send it on each GET
- Set HttpTokens to required to refuse IMDSv1 requests on an instance
- Containers need an extra hop; raise the metadata hop limit if they cannot reach it
- Use SDK credential chains rather than reading raw metadata URLs in application code
09Access: keys, roles, SSM
- AWS keeps only the public half of a key pair; a lost private key cannot be recovered
- An instance profile attaches an IAM role; SDKs get rotating temporary credentials
- Do not store long-lived access keys on instances; the attached role replaces them
- Session Manager opens a shell through the SSM agent, so port 22 need not be open
- Session Manager needs role permissions; session logs can go to S3 or CloudWatch
10Scaling and load balancing
- Auto Scaling groups hold capacity between min and max and replace failed instances
- Target tracking policies hold a metric, such as average CPU, near a chosen target
- Spread a group across AZs; a mixed-instances policy can blend Spot and On-Demand
- ALB routes HTTP and HTTPS at layer 7; NLB routes TCP, UDP and TLS at layer 4
- Enable ELB health checks on the group so failed LB checks also replace instances
11CloudWatch monitoring
- Basic metrics come every 5 minutes; detailed monitoring is 1 minute and costs extra
- Memory, disk and process metrics are not built in; install the CloudWatch agent
- StatusCheckFailed metrics split host faults from instance faults; alarm on both
- A system check alarm can trigger the EC2 recover action on supported instance types
12Common pitfalls
- Unattached EBS volumes and old snapshots keep billing after the instance is gone
- Security groups open to 0.0.0.0/0 on ports 22 or 3389 invite brute-force logins
- Cross-AZ and internet transfer bill per GB; check transfer lines in cost reports
- Root volumes are deleted on termination by default; data volumes are not
- Termination protection is off by default; while on, it blocks terminate calls
Go to the source
Use AWS documentation for current limits, availability, and pricing.