All infographics

SERVERLESS COMPUTE / 01

AWS Lambda

Run code without thinking about servers

Download sheet SVG

THE BIG PICTURE

From event to execution

  1. Event sourcesAPI Gateway · S3 · SQS
  2. InvocationSync · async · polling
  3. Lambda functionYour code + execution role
  4. AWS servicesStorage · data · monitoring

An event invokes your function. Its execution role controls access to other AWS services.

01Invocation models

  • Synchronous: the caller waits for the result and receives errors directly
  • Asynchronous: Lambda queues the event, returns at once and retries on error
  • Event source mappings: Lambda polls queues and streams and invokes in batches
  • Synchronous requests and buffered responses are capped at 6 MB; streamed responses at 200 MB
  • Asynchronous event payloads are capped at 1 MB; pass larger data by reference

02Environment and cold starts

  • Default Lambda runs one invocation per environment; Managed Instances can run several
  • Warm environments are reused, so module-level globals and /tmp persist between calls
  • Cold start time covers code download, runtime start and your module imports
  • Reduce cold starts with smaller packages, lazy imports and fewer dependencies
  • Provisioned concurrency initializes ahead; SnapStart restores snapshots on supported runtimes

03Concurrency and scaling

  • Concurrency is the number of invocations executing at the same time
  • Default Lambda scales each function by up to 1,000 environments every 10 seconds
  • Reserved concurrency caps a function and carves its share from the regional pool
  • Provisioned concurrency keeps initialized environments ready for a version or alias
  • Throttled sync calls fail with 429; async and polled invocations are retried

04Event sources

  • S3 and EventBridge invoke asynchronously; deliveries can repeat, so be idempotent
  • API Gateway and function URLs invoke synchronously; the caller waits for the reply
  • SQS is polled: set the queue visibility timeout to at least six times the function timeout
  • DynamoDB and Kinesis streams are polled per shard; a failing record blocks that shard
  • EventBridge rules and Scheduler can invoke functions on a cron or rate schedule

05Versions and rollouts

  • $LATEST is mutable; a published version is a fixed snapshot of code and config
  • An alias points at one version; a weighted alias splits traffic between two versions
  • Roll out by shifting a small weight to the new version, watching errors, then moving on
  • Point triggers and provisioned concurrency at an alias, not at $LATEST
  • SAM and CodeDeploy can shift traffic in steps and roll back on failed alarms

06Packages and layers

  • Direct zip upload is capped at 50 MB; code plus layers may total 250 MB unzipped
  • A function can attach up to five layers to share libraries across functions
  • Container images can be up to 10 GB and bring their own OS and runtime
  • Bigger packages mean longer download and unpack time on cold starts
  • Build dependencies on the target architecture (arm64 or x86_64) and pin versions

07Permissions and VPC access

  • The execution role sets what the function may call in AWS; keep it least-privilege
  • Resource-based policies control who may invoke the function, such as S3 or another account
  • VPC attachment places the function's network interfaces in your subnets, with no public IP
  • VPC IPv4 internet egress uses NAT; supported AWS APIs can use VPC endpoints
  • Keep secrets in Secrets Manager or Parameter Store, not in plain environment variables

08Retries and failures

  • Async invocations retry failed runs twice by default; the retry count can be changed
  • Max event age bounds how long an async event may wait before it is discarded
  • On-failure and on-success destinations receive outcome records for async invocations
  • Partial batch responses let a stream or queue retry only the items that failed
  • For SQS sources, use the queue's redrive policy to a dead-letter queue

09Logs, metrics and tracing

  • Stdout and stderr go to CloudWatch Logs, in a log group named /aws/lambda/FUNCTION
  • Key metrics: Invocations, Errors, Throttles, Duration and ConcurrentExecutions
  • IteratorAge shows how far a stream consumer lags behind the stream
  • Enable active X-Ray tracing to see calls to downstream services in a request
  • Log JSON with a request ID so one invocation's lines can be grouped and searched

10Pricing model

  • Default Lambda bills per request and per GB-second: memory times billed duration
  • Memory runs from 128 MB to 10,240 MB; CPU scales with memory, so more can cut time
  • Billed duration rounds up to the nearest millisecond
  • Provisioned concurrency is billed for as long as it is held, used or not
  • A free tier covers some monthly requests and compute; check current terms

11When to use it

  • Good fit: event-driven work, spiky traffic, glue code and short tasks
  • Default timeout limit is 15 minutes; Managed Instances allow 90 for some async/batch sources
  • For steady load, compare default Lambda costs with Managed Instances or EC2
  • Weak fit: GPU work and workloads that must hold long-lived client connections
  • Weak fit: latency-critical paths that cannot absorb cold starts without warm capacity

12Common pitfalls

  • One database connection per call exhausts the database; reuse clients or use RDS Proxy
  • Timeouts shorter than a downstream call leave half-finished work that gets retried
  • Background threads may freeze once the handler returns; await work before returning
  • Unbounded concurrency can overwhelm downstream systems; cap it with reserved concurrency
  • A function that writes to the source it triggers on can loop and run up cost

Go to the source

Use AWS documentation for current limits, availability, and pricing.

Lambda quotas Lambda execution environment Lambda Managed Instances execution environment Lambda SnapStart