Of the 52 questions in the system design bank, one comes up in cohort feedback more than any other. Members write more questions about it, get more wrong, and learn more from getting it wrong.
It’s the Google Docs question. Design a real-time collaborative editor.
Here’s why it’s the hardest — and why that makes it the most valuable one to understand deeply even if you never get asked it directly.
Why it’s different from every other question
Most system design questions have a “store data, serve data” structure underneath. Messages are stored and delivered. Files are stored and synced. Videos are stored and streamed. The interesting problems are about doing these operations reliably at scale.
Collaborative editing is different. The problem isn’t storage or delivery. It’s merging — two users type simultaneously into the same document, and both changes must survive in a consistent order on both clients, in real time, without either user’s edits being lost.
This is a fundamentally different problem. Storage and delivery can be solved with good distributed systems primitives. Merging concurrent edits requires a different class of solution: Operational Transformation or CRDTs.
The naive approaches fail interestingly
Most candidates start with one of two intuitions.
The first: “the server is the source of truth — last write wins.” This silently destroys data. User A types “Hello” and user B types “World” simultaneously. Whichever write arrives at the server last overwrites the first. One user loses their edit invisibly. Not acceptable.
The second: “the server merges conflicting writes.” This sounds right until you try to implement it. What does “merge” mean for concurrent text edits? If A inserts a character at position 5 and B deletes a character at position 3 simultaneously, what’s the correct merged result? The positions have shifted. The server cannot naively apply both operations in sequence.
This is where OT comes in.
What Operational Transformation actually does
Each edit is represented as an operation: insert(position, character) or delete(position). When two operations are generated concurrently — before either user has seen the other’s edit — they must be transformed against each other before both can be applied.
The classic example: document is “Hello.” User A inserts “ World” at position 5. User B simultaneously inserts “!” at position 5.
After A’s operation, position 5 has shifted — “!” should now be inserted at position 11. The transform function adjusts B’s operation before applying it. Result: “Hello World!”
The server maintains a history of all operations. When it receives an operation, it transforms that operation against all operations the client hasn’t yet seen, then broadcasts the transformed version to all clients. The document converges to the same state on all clients.
Why this matters for questions you will get asked
Understanding OT deeply unlocks several adjacent concepts:
The operation log as a data model. An OT-based system stores operations, not document state. The document is the result of replaying all operations from the beginning. This is event sourcing — the same pattern used in payment ledgers, audit trails, and any system where history matters more than current state.
The difference between mutable and immutable data models. Most system design candidates default to mutable state (update the row when something changes). The operation log is immutable — you only ever append. Understanding why immutable append-only logs are often better than mutable state is one of the higher-order concepts in distributed systems.
CRDTs as an alternative. Conflict-free Replicated Data Types solve the same merging problem differently — by designing data structures where all operations commute (can be applied in any order and produce the same result). Notion uses CRDTs. Google Docs uses OT. Both work.
If you understand one of these deeply, you understand a design principle that appears in message delivery systems, distributed databases, version control systems, and collaborative tools.
The Google Docs question is hard because it requires understanding a class of problem — concurrent state merging — that doesn’t appear in most backend engineering work. But once you understand it, you see it everywhere.
The full walkthrough is in the Vault under Real-Time Messaging. The drill card is the most complex card in the set. Both are there when you’re ready to go deep.
CTA: The cohort curriculum is built around closing exactly these three gaps — state merging, immutable operation logs, and real-time OT mechanics
Subscription link
https://systemdr.systemdrd.com/subscribe
The deeper concepts start where this lesson ends.
Explore the premium content for advanced techniques and system-level thinking.

