After reading a concept, test whether you can explain it through a concrete request or failure. This map follows the 12 chapters of DDIA’s first edition, matching this site’s notes; other editions may use different chapter numbers. Questions are editorial exercises, not questions attributed to the book’s authors or a company’s interview bank.
Chapter map: from concepts to traceable questions
| Chapter | What to learn | Practice question | Apply it |
|---|---|---|---|
| 1. Reliability, Scalability, Maintainability | Turn broad goals into workload assumptions and measurable service behavior. | What fails first when traffic grows 10×, and what would you measure? | Capacity calculator |
| 2. Data Models and Query Languages | Choose a model from access patterns, relationships, and update boundaries. | How would you model posts, follows, and feed retrieval? | News feed |
| 3. Storage and Retrieval | Connect indexes, write amplification, and read cost to a workload. | Which operations dominate a key-value store, and what does an index cost? | Key-value store |
| 4. Encoding and Evolution | Plan schema compatibility across stored data and mixed software versions. | How can a producer add a field without a synchronized consumer release? | Kafka and event consumers |
| 5. Replication | State the consistency guarantee and trace lag or failover. | After an update, where does a user read to avoid seeing the old value? | Key-value store replicas |
| 6. Partitioning | Choose keys, route requests, and account for skew and rebalancing. | How do you move a hot shard while reads and writes continue? | Sharding guide |
| 7. Transactions | Tie isolation levels to an invariant and a concrete concurrent schedule. | Two users reserve the last room: which operation decides the winner? | Booking race demonstration |
| 8. The Trouble with Distributed Systems | Treat timeout as uncertainty; account for pauses, clocks, and partial failure. | A charge times out: how do you find out whether money moved? | Payment system |
| 9. Consistency and Consensus | Separate linearizability, ordering, and agreement instead of conflating them. | What must be agreed upon during failover, and what may remain stale? | Key-value store failover |
| 10. Batch Processing | Make large computations reproducible with explicit inputs and outputs. | How would you rebuild a derived view without disrupting live traffic? | Crawler and indexing pipeline |
| 11. Stream Processing | Explain event time, replay, state, and the boundary of processing guarantees. | A consumer crashes after an external effect but before saving its offset: what happens? | Kafka and replay |
| 12. The Future of Data Systems | Connect derived data, end-to-end correctness, auditability, and responsibility. | Which records prove a checkout completed, and how do you repair disagreement? | Ecommerce checkout |
Choose a path for your weakest area
Correctness first
Chapter 7 → Chapter 8 → Chapter 9
Trace competing reservations, unknown payment outcomes, and a failover. State the invariant before choosing a mechanism. Open the matching guide →
Storage and scale
Chapter 2 → Chapter 3 → Chapter 5 → Chapter 6
Choose a data model, explain an index, then add replication and partitioning. Identify a hot key and a stale read. Open the matching guide →
Evolution and derived data
Chapter 4 → Chapter 10 → Chapter 11 → Chapter 12
Change an event schema and rebuild a projection. Explain replay, late events, external effects, and an audit trail. Open the matching guide →
Worked example: payment arrives after the inventory hold expires
Start with the rule: a hold expires at 10 seconds, and only confirmation before the deadline may move the booking to BOOKED. Payment callbacks may be delayed or repeated; the expiry worker may also run late.
- Chapter 7: the atomic boundary. Check booking state and deadline in the same transaction as the state transition and inventory changes. Repeating the operation must not release inventory again.
- Chapter 8: unknown outcomes. A charge timeout does not prove failure. Reconcile using a stable payment identity or retry under the provider’s idempotency contract instead of creating a second payment.
- Chapter 11: duplication and recovery. For a late success, persist one refund intent. Consumers may retry, but the refund call uses a stable key. Queuing the job does not mean the refund has completed.
This answer connects the product rule, concurrency, and recovery. A policy based on the provider’s charge timestamp would need different clock, evidence, and compensation rules; changing one comparison is insufficient.
Run the booking and payment timeline →
Three checks after each chapter
- Which invariant must hold, and which stale results or failures are acceptable?
- Can I trace two requests and one failure through the changing state?
- What latency, storage, operational, or recovery cost does the chosen mechanism add?
Browse all DDIA chapter notes · Check capacity assumptions
Book and edition information: Designing Data-Intensive Applications. This independent study aid does not replace the book.