← All posts
Problem WalkthroughsAugust 7, 202612 min read

Design a Chat System (WhatsApp/Messenger) — System Design Walkthrough

"Design a chat system" (WhatsApp, Messenger, Slack DMs) is a favorite because it forces you to reason about real-time delivery, connection management, and delivery guarantees — topics that don't come up in a simple CRUD service. This walkthrough runs the full problem end to end.

Step 1 — Clarify requirements

Functional

  • 1:1 messaging, and (usually) group chats.
  • Real-time delivery when the recipient is online.
  • Offline delivery — messages wait and arrive when the user reconnects.
  • Delivery/read receipts (sent, delivered, read) and typing/presence indicators.
  • Message history.

Non-functional

  • Low latency for online delivery (feels instant).
  • Reliability — messages must not be lost, and shouldn't be duplicated or reordered within a conversation.
  • High availability and horizontal scale to hundreds of millions of connections.

Confirm scope early: are we doing groups? media? end-to-end encryption? Each meaningfully changes the design.

Step 2 — The connection model

The defining decision: how does the server push a message to an online user in real time? Options: short polling (wasteful), long polling (better), or a persistent WebSocket connection (the standard answer). Clients hold an open WebSocket to a connection/gateway service; the server pushes messages down that socket instantly.

This creates the central scaling challenge: a user is connected to one specific gateway server, and you must be able to find which one when someone sends them a message. That's the routing problem below.

Step 3 — High-level design and message flow

Components:

  • Gateway/connection service — holds the WebSocket connections. Stateful and horizontally scaled.
  • Session registry — maps user_id → which gateway server they're connected to (kept in a fast store like Redis).
  • Message service — persists messages and orchestrates delivery.
  • Message store — durable history.
  • Push service — sends a mobile push (APNs/FCM) when the recipient is offline.

The flow for A → B:

  1. A sends a message over its WebSocket to its gateway.
  2. The message service persists it first (durability), assigns a sequence number, and acks A ("sent").
  3. It looks up B in the session registry. If B is online, it forwards the message to B's gateway, which pushes it down B's socket → "delivered."
  4. If B is offline, the message stays in B's store; a mobile push notification is triggered. B pulls it on reconnect.

Persisting before acking is what makes the system reliable — the message survives even if delivery fails.

Step 4 — Storage: inbox vs. conversation log

Two common models:

  • Per-conversation log: one append-only message list per conversation, shared by participants. Efficient storage, natural ordering, great for history. Retrieving "all my conversations" requires an index.
  • Per-user inbox: each user has their own message queue. Simple delivery and per-user read state, but duplicates message data across recipients (bad for large groups).

A common approach: a per-conversation log as the source of truth (partitioned by conversation ID), plus lightweight per-user metadata (last-read pointer, unread counts). Ordering within a conversation is guaranteed by a per-conversation sequence number, not wall-clock time.

Step 5 — Delivery guarantees, receipts, and idempotency

  • At-least-once + idempotency. Networks fail mid-send, so the client retries. Give each message a client-generated message ID; the server dedupes on it, turning at-least-once into effectively exactly-once storage. Without this, retries create duplicates.
  • Ordering comes from the per-conversation sequence number; clients render by it.
  • Receipts are just small status messages flowing the other way: the recipient's client sends "delivered"/"read" events that propagate back to the sender.

Step 6 — Groups and presence

  • Group fan-out: a message to a group of N members must reach all N. For small/medium groups, fan out to each member's delivery path. For very large groups, this is the same write-amplification problem as a social feed — you may cap group size or handle huge "broadcast" groups differently.
  • Presence (online/typing): cheap to get wrong at scale. Don't broadcast every keystroke to everyone; debounce typing events, and update presence via heartbeats with a short TTL rather than a constant stream. Presence is best-effort — it's fine if it's a second stale.

Step 7 — Follow-ups interviewers love

  • "How do you find which server a user is connected to?" The session registry (user_id → gateway), updated on connect/disconnect.
  • "What happens when a gateway server crashes?" Its connections drop; clients reconnect (to a new server) and re-register; undelivered messages are still in the store.
  • "How do you avoid losing messages?" Persist before ack; deliver from the store; dedupe on client message ID.
  • "How do you handle media (photos/video)?" Upload to blob storage first, then send a message containing the media URL — don't push large blobs through the message path.
  • "End-to-end encryption?" Keys live on devices; the server routes ciphertext it can't read — which also means server-side search and some features change.

Common mistakes

  • Acking before persisting — a crash then loses the message. Persist first.
  • No idempotency key — retries silently duplicate messages.
  • Broadcasting presence/typing to everyone — a self-inflicted scaling problem; debounce and use TTLs.
  • Pushing media through the message channel — upload to blob storage and send a link.
  • Forgetting the offline path — real-time is only half of it; the store + push notifications are the other half.

Practice this out loud

The chat system rewards clear reasoning about the connection model and delivery guarantees — and interviewers will probe "what if the server crashes mid-delivery?" hard. On Whitepad, a real-time chat design is a preset problem: a senior AI interviewer runs it by voice, watches your whiteboard, and presses on exactly these failure modes, then scores you on the rubric. Practice the follow-ups until they're routine.

Practice this out loud

Reading is the easy part. Sit across from a senior AI interviewer that talks, watches your whiteboard, and scores you like the real thing — your first mock is free.

Start a free mock →