All infographics

NOSQL DATABASE / 05

Amazon DynamoDB

A fully managed database at any scale

Download sheet SVG

THE BIG PICTURE

Keys shape your data access

  1. ApplicationGet · put · query
  2. Table + keysPartition key / sort key
  3. PartitionsData distributed by key
DynamoDB StreamsOptional stream of item changes

Partition keys distribute data; an optional sort key orders items within a partition key. Streams can publish item changes.

01Tables, items and keys

  • A table holds items; an item is a set of named attributes
  • Only key attributes are declared up front; other attributes are schemaless
  • Simple key: partition key alone. Composite: partition key plus sort key
  • Key attributes must be string, number or binary, never map or list
  • Partition key is hashed to pick a partition; sort key orders items sharing it

02Partition key design

  • Pick a high-cardinality partition key so traffic spreads over partitions
  • Low-cardinality keys, like a status flag, concentrate load on few partitions
  • Hot partitions can throttle even when total table capacity looks sufficient
  • Write sharding appends a suffix to split one hot key into N keys
  • Reads for a sharded key must query every suffix and merge the results

03Secondary indexes

  • GSI: own partition key and optional sort key; can be added or dropped later
  • GSI reads are eventually consistent only; projection picks copied attributes
  • Sparse GSI: items lacking the index key attributes are not indexed
  • LSI: same partition key, different sort key; must be set at table creation
  • LSI supports strongly consistent reads and uses the table's capacity

04Query vs scan

  • Query reads items for one partition key value, optionally by sort key
  • Scan reads every item in a table or index; avoid it on request paths
  • Filter expressions run after the read, so filtered-out items still cost
  • Each Query or Scan call reads at most 1 MB, so large results span several calls
  • Parallel scan splits a scan into segments; the total read is unchanged

05Read consistency and DAX

  • Reads are eventually consistent by default and cost half as much
  • ConsistentRead=true requests a strong read from a table or an LSI
  • Strong reads cost double, so use them only where a stale value would cause harm
  • DAX caches eventually consistent reads; it suits hot, frequently read items
  • DAX passes strong reads through; apps need the DAX client library

06Transactions and batches

  • TransactWriteItems applies up to 100 actions atomically, all or nothing
  • Transactions span tables in one account and Region; they cannot target indexes
  • Conflicts cancel the whole call; retry it with a ClientRequestToken
  • Transactional calls cost double the normal read or write units
  • Batch calls are not atomic; resend UnprocessedItems and UnprocessedKeys

07TTL for expiring items

  • TTL reads a number attribute holding a Unix epoch time in seconds
  • Expired items are deleted in the background, usually well after their expiry time
  • Expired items still appear in reads and scans until actually deleted
  • TTL deletes use no write capacity in the source Region; replicas do bill
  • TTL deletes have a service identity in the source Region stream; replicated deletes do not

08DynamoDB Streams

  • Streams log item-level changes; changes to one item stay in order
  • View types: keys only, new image, old image, or both images
  • Stream records are kept for 24 hours, then they expire
  • Each record appears once, but a retried batch means handlers must be idempotent
  • Design for at most two readers per shard; AWS recommends one for global tables

09Global tables

  • Global tables replicate one table across AWS Regions, one replica per Region
  • Default MREC mode replicates asynchronously, usually with only a short delay
  • MREC conflicts resolve per item by last writer wins; any replica accepts writes
  • MREC strong reads are current only for items last written in that Region
  • MRSC replicates synchronously across supported Regions; TTL and transactions are unsupported

10Backups and recovery

  • On-demand backups are full copies, kept until you delete them
  • PITR keeps continuous backups for up to 35 days, configurable from 1
  • PITR restores to any second in that window, into a new table
  • PITR covers only times after it was enabled; turn it on early
  • Export to S3 reads a snapshot without using table read capacity

11Capacity modes and pricing

  • On-demand: pay per read and write request, no capacity planning
  • Provisioned: set read and write units per second, optionally auto scaled
  • 1 read unit = one strongly consistent read up to 4 KB; 1 write unit = 1 KB
  • Storage is billed per GB-month; Standard-IA suits tables where storage dominates
  • Backups, PITR, DAX nodes and replicated writes have separate charges

12Modelling and pitfalls

  • Model from access patterns first; keys follow the queries you need
  • Overloaded PK and SK values let one Query fetch related item types
  • Item size limit is 400 KB; put large blobs in S3 and keep a pointer
  • Pitfall: stop paging only when LastEvaluatedKey is absent; pages can be empty
  • Pitfall: retrying throttled calls without backoff amplifies the throttling

Go to the source

Use AWS documentation for current limits, availability, and pricing.

DynamoDB constraints DynamoDB Streams and TTL How global tables work DynamoDB transactions