Design a Chat System (WhatsApp/Messenger) — System Design Walkthrough
"Design a chat system" (WhatsApp, Messenger, Slack DMs) is a favorite because it forces you to reason about real-time delivery, connection management, and delivery guarantees — topics that don't come up in a simple CRUD service. This walkthrough runs the full problem end to end.
Step 1 — Clarify requirements
Functional
- 1:1 messaging, and (usually) group chats.
- Real-time delivery when the recipient is online.
- Offline delivery — messages wait and arrive when the user reconnects.
- Delivery/read receipts (sent, delivered, read) and typing/presence indicators.
- Message history.
Non-functional
- Low latency for online delivery (feels instant).
- Reliability — messages must not be lost, and shouldn't be duplicated or reordered within a conversation.
- High availability and horizontal scale to hundreds of millions of connections.
Confirm scope early: are we doing groups? media? end-to-end encryption? Each meaningfully changes the design.
Step 2 — The connection model
The defining decision: how does the server push a message to an online user in real time? Options: short polling (wasteful), long polling (better), or a persistent WebSocket connection (the standard answer). Clients hold an open WebSocket to a connection/gateway service; the server pushes messages down that socket instantly.
This creates the central scaling challenge: a user is connected to one specific gateway server, and you must be able to find which one when someone sends them a message. That's the routing problem below.
Step 3 — High-level design and message flow
Components:
- Gateway/connection service — holds the WebSocket connections. Stateful and horizontally scaled.
- Session registry — maps
user_id → which gateway server they're connected to(kept in a fast store like Redis). - Message service — persists messages and orchestrates delivery.
- Message store — durable history.
- Push service — sends a mobile push (APNs/FCM) when the recipient is offline.
The flow for A → B:
- A sends a message over its WebSocket to its gateway.
- The message service persists it first (durability), assigns a sequence number, and acks A ("sent").
- It looks up B in the session registry. If B is online, it forwards the message to B's gateway, which pushes it down B's socket → "delivered."
- If B is offline, the message stays in B's store; a mobile push notification is triggered. B pulls it on reconnect.
Persisting before acking is what makes the system reliable — the message survives even if delivery fails.
Step 4 — Storage: inbox vs. conversation log
Two common models:
- Per-conversation log: one append-only message list per conversation, shared by participants. Efficient storage, natural ordering, great for history. Retrieving "all my conversations" requires an index.
- Per-user inbox: each user has their own message queue. Simple delivery and per-user read state, but duplicates message data across recipients (bad for large groups).
A common approach: a per-conversation log as the source of truth (partitioned by conversation ID), plus lightweight per-user metadata (last-read pointer, unread counts). Ordering within a conversation is guaranteed by a per-conversation sequence number, not wall-clock time.
Step 5 — Delivery guarantees, receipts, and idempotency
- At-least-once + idempotency. Networks fail mid-send, so the client retries. Give each message a client-generated message ID; the server dedupes on it, turning at-least-once into effectively exactly-once storage. Without this, retries create duplicates.
- Ordering comes from the per-conversation sequence number; clients render by it.
- Receipts are just small status messages flowing the other way: the recipient's client sends "delivered"/"read" events that propagate back to the sender.
Step 6 — Groups and presence
- Group fan-out: a message to a group of N members must reach all N. For small/medium groups, fan out to each member's delivery path. For very large groups, this is the same write-amplification problem as a social feed — you may cap group size or handle huge "broadcast" groups differently.
- Presence (online/typing): cheap to get wrong at scale. Don't broadcast every keystroke to everyone; debounce typing events, and update presence via heartbeats with a short TTL rather than a constant stream. Presence is best-effort — it's fine if it's a second stale.
Step 7 — Follow-ups interviewers love
- "How do you find which server a user is connected to?" The session registry (
user_id → gateway), updated on connect/disconnect. - "What happens when a gateway server crashes?" Its connections drop; clients reconnect (to a new server) and re-register; undelivered messages are still in the store.
- "How do you avoid losing messages?" Persist before ack; deliver from the store; dedupe on client message ID.
- "How do you handle media (photos/video)?" Upload to blob storage first, then send a message containing the media URL — don't push large blobs through the message path.
- "End-to-end encryption?" Keys live on devices; the server routes ciphertext it can't read — which also means server-side search and some features change.
Common mistakes
- Acking before persisting — a crash then loses the message. Persist first.
- No idempotency key — retries silently duplicate messages.
- Broadcasting presence/typing to everyone — a self-inflicted scaling problem; debounce and use TTLs.
- Pushing media through the message channel — upload to blob storage and send a link.
- Forgetting the offline path — real-time is only half of it; the store + push notifications are the other half.
Practice this out loud
The chat system rewards clear reasoning about the connection model and delivery guarantees — and interviewers will probe "what if the server crashes mid-delivery?" hard. On Whitepad, a real-time chat design is a preset problem: a senior AI interviewer runs it by voice, watches your whiteboard, and presses on exactly these failure modes, then scores you on the rubric. Practice the follow-ups until they're routine.
Practice this out loud
Reading is the easy part. Sit across from a senior AI interviewer that talks, watches your whiteboard, and scores you like the real thing — your first mock is free.
Start a free mock →