SERVERLESS COMPUTE / 01
AWS Lambda
Run code without thinking about servers
THE BIG PICTURE
From event to execution
- Event sourcesAPI Gateway · S3 · SQS
- InvocationSync · async · polling
- Lambda functionYour code + execution role
- AWS servicesStorage · data · monitoring
An event invokes your function. Its execution role controls access to other AWS services.
01Invocation models
- Synchronous: the caller waits for the result and receives errors directly
- Asynchronous: Lambda queues the event, returns at once and retries on error
- Event source mappings: Lambda polls queues and streams and invokes in batches
- Synchronous requests and buffered responses are capped at 6 MB; streamed responses at 200 MB
- Asynchronous event payloads are capped at 1 MB; pass larger data by reference
02Environment and cold starts
- Default Lambda runs one invocation per environment; Managed Instances can run several
- Warm environments are reused, so module-level globals and /tmp persist between calls
- Cold start time covers code download, runtime start and your module imports
- Reduce cold starts with smaller packages, lazy imports and fewer dependencies
- Provisioned concurrency initializes ahead; SnapStart restores snapshots on supported runtimes
03Concurrency and scaling
- Concurrency is the number of invocations executing at the same time
- Default Lambda scales each function by up to 1,000 environments every 10 seconds
- Reserved concurrency caps a function and carves its share from the regional pool
- Provisioned concurrency keeps initialized environments ready for a version or alias
- Throttled sync calls fail with 429; async and polled invocations are retried
04Event sources
- S3 and EventBridge invoke asynchronously; deliveries can repeat, so be idempotent
- API Gateway and function URLs invoke synchronously; the caller waits for the reply
- SQS is polled: set the queue visibility timeout to at least six times the function timeout
- DynamoDB and Kinesis streams are polled per shard; a failing record blocks that shard
- EventBridge rules and Scheduler can invoke functions on a cron or rate schedule
05Versions and rollouts
- $LATEST is mutable; a published version is a fixed snapshot of code and config
- An alias points at one version; a weighted alias splits traffic between two versions
- Roll out by shifting a small weight to the new version, watching errors, then moving on
- Point triggers and provisioned concurrency at an alias, not at $LATEST
- SAM and CodeDeploy can shift traffic in steps and roll back on failed alarms
06Packages and layers
- Direct zip upload is capped at 50 MB; code plus layers may total 250 MB unzipped
- A function can attach up to five layers to share libraries across functions
- Container images can be up to 10 GB and bring their own OS and runtime
- Bigger packages mean longer download and unpack time on cold starts
- Build dependencies on the target architecture (arm64 or x86_64) and pin versions
07Permissions and VPC access
- The execution role sets what the function may call in AWS; keep it least-privilege
- Resource-based policies control who may invoke the function, such as S3 or another account
- VPC attachment places the function's network interfaces in your subnets, with no public IP
- VPC IPv4 internet egress uses NAT; supported AWS APIs can use VPC endpoints
- Keep secrets in Secrets Manager or Parameter Store, not in plain environment variables
08Retries and failures
- Async invocations retry failed runs twice by default; the retry count can be changed
- Max event age bounds how long an async event may wait before it is discarded
- On-failure and on-success destinations receive outcome records for async invocations
- Partial batch responses let a stream or queue retry only the items that failed
- For SQS sources, use the queue's redrive policy to a dead-letter queue
09Logs, metrics and tracing
- Stdout and stderr go to CloudWatch Logs, in a log group named /aws/lambda/FUNCTION
- Key metrics: Invocations, Errors, Throttles, Duration and ConcurrentExecutions
- IteratorAge shows how far a stream consumer lags behind the stream
- Enable active X-Ray tracing to see calls to downstream services in a request
- Log JSON with a request ID so one invocation's lines can be grouped and searched
10Pricing model
- Default Lambda bills per request and per GB-second: memory times billed duration
- Memory runs from 128 MB to 10,240 MB; CPU scales with memory, so more can cut time
- Billed duration rounds up to the nearest millisecond
- Provisioned concurrency is billed for as long as it is held, used or not
- A free tier covers some monthly requests and compute; check current terms
11When to use it
- Good fit: event-driven work, spiky traffic, glue code and short tasks
- Default timeout limit is 15 minutes; Managed Instances allow 90 for some async/batch sources
- For steady load, compare default Lambda costs with Managed Instances or EC2
- Weak fit: GPU work and workloads that must hold long-lived client connections
- Weak fit: latency-critical paths that cannot absorb cold starts without warm capacity
12Common pitfalls
- One database connection per call exhausts the database; reuse clients or use RDS Proxy
- Timeouts shorter than a downstream call leave half-finished work that gets retried
- Background threads may freeze once the handler returns; await work before returning
- Unbounded concurrency can overwhelm downstream systems; cap it with reserved concurrency
- A function that writes to the source it triggers on can loop and run up cost
Go to the source
Use AWS documentation for current limits, availability, and pricing.