CONTAINER ORCHESTRATION / 08
Amazon ECS
Schedule, deploy and scale your containers
THE BIG PICTURE
From task definition to running service
- Task definitionImage + resources + roles
- ECS serviceMaintain desired task count
- Running tasksApplication containers
Compute capacityFargate / EC2 / Managed Instances
Example service: ECS maintains the desired task count. Capacity providers select compute; standalone tasks can run without a service.
01Clusters, tasks and services
- ECS is a managed container orchestrator; the selected compute option runs your containers
- A cluster is a regional grouping of tasks, services and their compute capacity
- A task definition describes an application; a task is one running copy of that definition
- A service maintains a desired task count and replaces tasks that fail or stop
- Standalone tasks suit jobs that finish; EventBridge Scheduler can launch them on a schedule
02Task definitions and images
- Task definitions version images, CPU, memory, networking, roles, logs and volume settings together
- A task can contain an app and sidecars that share its lifecycle; scale independent apps separately
- Use a unique build tag or image digest so each release can be identified and reproduced
- Register a new task definition revision, then update the service to deploy that revision
- Put container health checks in the task definition; image-only Docker checks are not monitored
03Choosing compute capacity
- Fargate runs each task in an isolated environment while AWS manages its underlying servers
- Self-managed EC2 gives host and instance control; you maintain the OS, agent and instance capacity
- ECS Managed Instances handles EC2 provisioning, scaling and patching and supports specialized hardware
- ECS Anywhere registers external servers to the ECS control plane for supported container workloads
- Validate task compatibility and feature support before moving between compute options
04Capacity providers
- A capacity provider strategy chooses compute; a cluster default can be overridden per service or task
- Base places an initial task count on one provider; weights distribute tasks after that base
- FARGATE and FARGATE_SPOT can share a strategy to combine regular and interruptible capacity
- A cluster can mix provider types, but one strategy cannot mix Fargate, EC2 ASG and Managed Instances
- EC2 ASG providers with managed scaling adjust host capacity for tasks assigned to that provider
05Networking and load balancing
- The awsvpc mode gives each task an ENI and private IP; Fargate requires this network mode
- Choose subnets and security groups for awsvpc tasks; allow application ports from intended callers
- An ALB routes HTTP and HTTPS; an NLB handles transport traffic such as TCP, UDP and TLS
- Use IP target groups for awsvpc tasks; instance targets are for applicable EC2 network modes
- Image pulls, logs and AWS APIs need reachable endpoints through NAT, public routing or VPC endpoints
06IAM roles and secrets
- The task role grants AWS permissions to application code; SDKs obtain temporary credentials
- The execution role lets the ECS or Fargate agent pull images, publish logs and fetch startup secrets
- The EC2 instance role serves the host agent; keep application permissions in separate task roles
- Reference Secrets Manager or Parameter Store instead of placing secret values in images or definitions
- Secrets injected into environment variables do not refresh after rotation; launch new tasks
07Deployments and health
- Rolling deployments replace existing tasks with new ones while the service remains running
- Minimum healthy and maximum percent control how many tasks a rolling deployment keeps or adds
- Blue/green, canary and linear strategies support traffic shifting with compatible routing integrations
- Enable the rolling deployment circuit breaker and rollback to recover from failed deployments
- Set health-check grace time for startup and let apps finish requests during graceful shutdown
08Service scaling and availability
- Service Auto Scaling changes desired task count within the minimum and maximum you configure
- Target tracking can follow service CPU or memory; compatible ALB services can track requests per target
- Scheduled scaling prepares for known peaks; cooldowns limit rapid repeated scaling actions
- Task scaling and EC2 host scaling solve separate problems; ensure enough host capacity for new tasks
- Use multiple AZ subnets and replicas; check service placement and Availability Zone rebalancing
09Service-to-service communication
- Service Connect gives ECS services stable short names, managed proxies and communication metrics
- Name container ports in the task definition, then select those ports in Service Connect settings
- A Cloud Map namespace groups Service Connect endpoints across services in the same AWS Region
- Allow the required proxy ports through security groups; discovery does not create network access
- Clients outside Service Connect need another discovery or routing method, such as a load balancer
10Storage and data lifetime
- Ephemeral task storage suits temporary files; keep durable application state outside the task
- Supported Linux tasks can mount EFS for shared files that survive task replacement
- Task-managed EBS creates a new volume per task, optionally initialized from an existing snapshot
- Service-managed EBS volumes are deleted when tasks stop; standalone tasks can preserve their volumes
- EC2 host bind mounts depend on that host; a replacement task on another host will not inherit the files
11Logs, metrics and debugging
- Configure the awslogs driver to send container stdout and stderr to CloudWatch Logs
- CloudWatch service metrics show CPU and memory; Container Insights adds detail at extra cost
- Use EventBridge task and deployment events to react to stopped tasks or failed rollouts
- Inspect service events, stoppedReason and container exit codes when tasks cannot start or remain healthy
- ECS Exec uses Systems Manager for container commands; enable it on new tasks and scope IAM access
12Costs and common pitfalls
- ECS orchestration has no extra fee for EC2 or Fargate; Managed Instances adds a management fee
- Fargate bills requested resources from image pull; EC2 bills whole instances, including idle capacity
- Include load balancers, NAT, public IPv4, storage, transfer and log ingestion in the estimate
- A service replaces stopped tasks until you reduce desired count or delete it; stopping one is temporary
- Pending tasks can signal insufficient compute or subnet IPs; image pull failures also need IAM and routing checks
Go to the source
Use AWS documentation for current limits, availability, and pricing.