Design YouTube (Video Streaming) — System Design Interview Walkthrough
"Design YouTube" (or Netflix) is a large, satisfying system design problem: it spans uploads, heavy processing, enormous storage, and global low-latency delivery. The trick is to recognize it's really two very different systems glued together — a write-side ingestion/processing pipeline and a massively read-heavy delivery path — and to spend your time on the parts that actually carry the load. Here's the walkthrough.
Step 1 — Clarify requirements
Functional
- Creators upload videos.
- Viewers stream videos smoothly, at a quality that adapts to their connection.
- Browse/search by metadata; view counts; (optional) comments, recommendations.
Non-functional
- Extremely read-heavy — views vastly outnumber uploads.
- Low startup latency and no buffering during playback, worldwide.
- Huge, durable storage — petabytes and growing.
- High availability; uploads can be processed asynchronously (a few minutes' delay is fine).
Confirm scope: are recommendations and comments in? Usually you focus on upload → process → store → stream, and treat the rest as extensions.
Step 2 — Back-of-the-envelope
The numbers justify the architecture. Even modest assumptions — say 500 hours of video uploaded per minute at scale, each transcoded into several resolutions — produce petabytes of storage growth and make transcoding a massive compute job. On the read side, billions of views per day means the delivery path, not the database, is where the system lives or dies. Two takeaways: storage must be cheap blob storage, and delivery must be pushed to a CDN.
Step 3 — High-level design
Two pipelines:
Write path (ingestion):
Upload service → raw video in blob storage → transcoding pipeline → multiple encoded renditions back to blob storage → metadata DB updated → CDN warmed.
Read path (playback):
Client → metadata/API service (gets the manifest + stream URLs) → CDN serves the video segments.
Plus a metadata service (titles, descriptions, video-to-file mappings, view counts) backed by a database, and a search index.
The interesting deep-dives are transcoding and delivery.
Step 4 — Upload and transcoding
Raw uploads are large, so accept them directly into blob/object storage (often via a resumable, chunked upload so a dropped connection doesn't restart the whole file). Uploading is decoupled from processing: once the raw file lands, drop a job on a queue.
Transcoding is the heavy lifting. A fleet of workers picks up jobs and encodes each video into:
- Multiple resolutions/bitrates (240p → 4K) so playback can adapt.
- Segmented formats for adaptive streaming — HLS or DASH — which split each rendition into small chunks (a few seconds each) plus a manifest listing them.
Transcoding is embarrassingly parallel (split the video, encode segments concurrently, reassemble), so it scales horizontally with the worker fleet. When done, the renditions go to blob storage and the metadata DB is updated to mark the video ready.
Step 5 — Delivery: CDN + adaptive bitrate
This is the heart of the read path. You never stream petabytes from origin — you push content to a CDN with edge locations near users. The player fetches the manifest, then requests segments from the nearest edge.
Adaptive bitrate streaming (ABR) is what makes playback smooth: the client monitors its bandwidth and, segment by segment, requests a higher- or lower-quality chunk. A user on a fast connection gets 1080p; when their network dips, the next segment drops to 480p — no buffering, no manual quality switch. Explaining ABR + HLS/DASH segments is the key technical signal for this problem.
Step 6 — Metadata and view counting at scale
- Metadata store: video records (id → title, uploader, rendition URLs, status) in a database, plus a search index for discovery. This is modest compared to the blobs.
- View counting: a hot, high-write path. Do not do a synchronous
UPDATE views = views + 1on every play — that hammers the database and creates contention on popular videos. Instead, emit a view event to a queue/stream and aggregate asynchronously (batch increments, approximate counts), reconciling periodically. Real-time exact counts aren't necessary. - Thumbnails/previews are generated during transcoding and served from the CDN too.
Step 7 — Follow-ups interviewers love
- "How do you handle a huge upload that fails midway?" Chunked/resumable upload to blob storage; resume from the last received chunk.
- "How do you keep startup latency low?" Serve the manifest fast, prefetch the first segments, and rely on CDN edge proximity.
- "How do popular videos not overwhelm origin?" The CDN absorbs the reads; origin only serves cache fills.
- "How do you count views without slowing playback?" Async event pipeline, not synchronous DB writes.
- "Live streaming instead of VOD?" Different beast — low-latency segmenting, no full pre-transcode; acknowledge the distinction.
Common mistakes
- Trying to stream from a database or origin. Video lives in blob storage and is delivered by a CDN — full stop.
- Ignoring transcoding. "Just store the upload" misses that you need multiple renditions and segments for adaptive playback.
- Synchronous view counting. A per-play DB increment doesn't survive a viral video.
- Forgetting ABR. Without adaptive bitrate, playback quality can't respond to changing networks — the defining feature of good streaming.
Practice this out loud
This problem has a lot of surface area, so interviewers steer hard — "okay, go deep on the transcoding pipeline" or "how does the player pick a quality?". On Whitepad, a senior AI interviewer runs a video-streaming design by voice, watches your whiteboard, and pushes on the upload pipeline, ABR delivery, and view-counting trade-offs in real time — then scores you on the rubric. Practice the deep-dives until you can go anywhere the interviewer points.
Practice this out loud
Reading is the easy part. Sit across from a senior AI interviewer that talks, watches your whiteboard, and scores you like the real thing — your first mock is free.
Start a free mock →