NOSQL DATABASE / 05
Amazon DynamoDB
A fully managed database at any scale
THE BIG PICTURE
Keys shape your data access
- ApplicationGet · put · query
- Table + keysPartition key / sort key
- PartitionsData distributed by key
DynamoDB StreamsOptional stream of item changes
Partition keys distribute data; an optional sort key orders items within a partition key. Streams can publish item changes.
01Tables, items and keys
- A table holds items; an item is a set of named attributes
- Only key attributes are declared up front; other attributes are schemaless
- Simple key: partition key alone. Composite: partition key plus sort key
- Key attributes must be string, number or binary, never map or list
- Partition key is hashed to pick a partition; sort key orders items sharing it
02Partition key design
- Pick a high-cardinality partition key so traffic spreads over partitions
- Low-cardinality keys, like a status flag, concentrate load on few partitions
- Hot partitions can throttle even when total table capacity looks sufficient
- Write sharding appends a suffix to split one hot key into N keys
- Reads for a sharded key must query every suffix and merge the results
03Secondary indexes
- GSI: own partition key and optional sort key; can be added or dropped later
- GSI reads are eventually consistent only; projection picks copied attributes
- Sparse GSI: items lacking the index key attributes are not indexed
- LSI: same partition key, different sort key; must be set at table creation
- LSI supports strongly consistent reads and uses the table's capacity
04Query vs scan
- Query reads items for one partition key value, optionally by sort key
- Scan reads every item in a table or index; avoid it on request paths
- Filter expressions run after the read, so filtered-out items still cost
- Each Query or Scan call reads at most 1 MB, so large results span several calls
- Parallel scan splits a scan into segments; the total read is unchanged
05Read consistency and DAX
- Reads are eventually consistent by default and cost half as much
- ConsistentRead=true requests a strong read from a table or an LSI
- Strong reads cost double, so use them only where a stale value would cause harm
- DAX caches eventually consistent reads; it suits hot, frequently read items
- DAX passes strong reads through; apps need the DAX client library
06Transactions and batches
- TransactWriteItems applies up to 100 actions atomically, all or nothing
- Transactions span tables in one account and Region; they cannot target indexes
- Conflicts cancel the whole call; retry it with a ClientRequestToken
- Transactional calls cost double the normal read or write units
- Batch calls are not atomic; resend UnprocessedItems and UnprocessedKeys
07TTL for expiring items
- TTL reads a number attribute holding a Unix epoch time in seconds
- Expired items are deleted in the background, usually well after their expiry time
- Expired items still appear in reads and scans until actually deleted
- TTL deletes use no write capacity in the source Region; replicas do bill
- TTL deletes have a service identity in the source Region stream; replicated deletes do not
08DynamoDB Streams
- Streams log item-level changes; changes to one item stay in order
- View types: keys only, new image, old image, or both images
- Stream records are kept for 24 hours, then they expire
- Each record appears once, but a retried batch means handlers must be idempotent
- Design for at most two readers per shard; AWS recommends one for global tables
09Global tables
- Global tables replicate one table across AWS Regions, one replica per Region
- Default MREC mode replicates asynchronously, usually with only a short delay
- MREC conflicts resolve per item by last writer wins; any replica accepts writes
- MREC strong reads are current only for items last written in that Region
- MRSC replicates synchronously across supported Regions; TTL and transactions are unsupported
10Backups and recovery
- On-demand backups are full copies, kept until you delete them
- PITR keeps continuous backups for up to 35 days, configurable from 1
- PITR restores to any second in that window, into a new table
- PITR covers only times after it was enabled; turn it on early
- Export to S3 reads a snapshot without using table read capacity
11Capacity modes and pricing
- On-demand: pay per read and write request, no capacity planning
- Provisioned: set read and write units per second, optionally auto scaled
- 1 read unit = one strongly consistent read up to 4 KB; 1 write unit = 1 KB
- Storage is billed per GB-month; Standard-IA suits tables where storage dominates
- Backups, PITR, DAX nodes and replicated writes have separate charges
12Modelling and pitfalls
- Model from access patterns first; keys follow the queries you need
- Overloaded PK and SK values let one Query fetch related item types
- Item size limit is 400 KB; put large blobs in S3 and keep a pointer
- Pitfall: stop paging only when LastEvaluatedKey is absent; pages can be empty
- Pitfall: retrying throttled calls without backoff amplifies the throttling
Go to the source
Use AWS documentation for current limits, availability, and pricing.