Design a URL Shortener — System Design Interview Walkthrough
"Design a URL shortener" (think Bitly or TinyURL) is one of the most common system design interview questions, and for good reason: it's simple enough to finish in 45 minutes but rich enough to expose how you think about IDs, storage, caching, and read-heavy scale. This walkthrough runs the full problem the way you'd run it live.
Step 1 — Clarify requirements
Before drawing anything, pin down scope. Reasonable answers for this problem:
Functional
- Given a long URL, return a short URL.
- Visiting the short URL redirects to the original.
- Optional: custom aliases, link expiry, basic analytics (click counts).
Non-functional
- Read-heavy: redirects vastly outnumber creations — often 100:1 or more.
- Low latency on redirects (this is the hot path — tens of milliseconds).
- High availability: a dead redirect service breaks every link ever created.
- Short codes should be hard to guess for unlisted links.
Stating the 100:1 read/write ratio early is a strong signal — it tells the interviewer you already know caching and the read path will dominate the design.
Step 2 — Back-of-the-envelope estimation
Say we expect 100 million new URLs per month. That's roughly 40 writes/second. At a 100:1 read ratio, that's ~4,000 redirects/second, with peaks several times higher.
Storage: 100M/month × 12 × 5 years ≈ 6 billion URLs. At ~500 bytes per record (short code, long URL, metadata), that's about 3 TB over five years — comfortably within a single well-provisioned database with room to shard later.
These numbers matter because they justify decisions: 3 TB says "don't over-engineer the storage"; 4,000 reads/second says "cache aggressively."
Step 3 — API design
Keep it minimal:
POST /urls { "long_url": "https://…", "alias": "optional" } -> { "short_url": "https://wp.dev/aX9k2" }
GET /{short_code} -> 301/302 redirect to the long URL
A subtle but important choice: 301 (permanent) vs. 302 (temporary) redirect. A 301 lets browsers and proxies cache the redirect, which slashes load on your service — but it also means you lose per-click analytics and can't easily change the destination. A 302 keeps every click flowing through you (good for analytics, changeable links) at the cost of more traffic. Naming this trade-off out loud scores points.
Step 4 — The core question: how do you generate the short code?
This is the heart of the problem. Three common approaches:
Option A — Hash the URL (e.g., MD5/SHA) and take the first N characters. Simple, but hashes collide, so you need collision detection and retry. Two different long URLs can also map to different codes for the same content, which is usually fine.
Option B — Auto-increment ID, then Base62-encode it. Use a monotonic counter, encode the integer in Base62 ([a-zA-Z0-9]). A 7-character Base62 code covers 62⁷ ≈ 3.5 trillion URLs — plenty. This is clean and collision-free by construction. The catch: a single counter is a bottleneck and a single point of failure, and sequential codes are guessable/enumerable.
Option C — A distributed ID generator. Hand out ID ranges to each app server (e.g., a key-generation service that leases blocks of the counter), or use a scheme like Snowflake. This removes the single-counter bottleneck while keeping codes collision-free. This is usually the strongest answer for scale.
A good deep-dive: start with Base62 of a counter, then, when the interviewer asks "what happens when one counter can't keep up?", evolve to a key-generation service that pre-allocates ranges. Showing that evolution is exactly the reasoning they want.
Step 5 — Storage
The access pattern is a simple key-value lookup: short code → long URL. That points at a key-value or wide-column store (or even a well-indexed relational table given the modest 3 TB). The schema is trivial:
short_code (PK) | long_url | created_at | expires_at | owner_id
Because reads dominate and are point lookups by primary key, this shards cleanly on short_code when you outgrow one node.
Step 6 — The read path and caching
Redirects are 100x the writes, so the read path deserves the most attention. Put a cache (e.g., Redis) in front of the database keyed by short code. With a strong hit rate, the vast majority of redirects never touch the database.
Two refinements worth mentioning:
- Cache the hot set. A small fraction of links get most of the traffic (a viral link, a marketing campaign). An LRU cache naturally keeps those resident.
- Push reads to the edge. For 301s especially, a CDN can serve redirects close to users, cutting latency and origin load dramatically.
Step 7 — Scale, failure, and the follow-ups interviewers love
Once the happy path is solid, expect probing questions. Have answers ready:
- "What happens when the cache goes down?" Reads fall through to the database; make sure it can absorb the miss traffic, and consider request coalescing to avoid a thundering herd on a hot key.
- "How do you handle expiry?" Store
expires_at, check it on read, and clean up lazily (on access) or with a background job. - "How do you prevent abuse / malicious links?" Rate-limit creation per user/IP, and run destinations against a safe-browsing check.
- "How do you count clicks without slowing redirects?" Don't write synchronously on the hot path — emit an event to a queue and aggregate asynchronously.
- "How do you keep custom aliases unique?" A uniqueness constraint on
short_codeplus a clear conflict error.
Common mistakes
- Over-engineering the storage. 3 TB does not need an exotic setup; say so.
- Ignoring the read/write asymmetry. If you spend all your time on writes, you missed the point — redirects are the system.
- Hand-waving key generation. This is the crux; have a real, collision-safe scheme and know its bottleneck.
- Forgetting the 301-vs-302 trade-off. It's a small detail that signals real product thinking.
Practice this one out loud
Reading a walkthrough is the easy part. The interview is a conversation — the interviewer will interrupt with "what if the counter is the bottleneck?" and watch how you adapt. You can practice exactly that: the URL shortener is one of the preset problems on Whitepad, where a senior AI interviewer runs this phased script by voice, watches your whiteboard, and pushes on your key-generation and caching choices in real time — then scores you against the same rubric a real interviewer uses.
Design it once on paper. Then design it out loud, under questioning, until the follow-ups feel routine.
Practice this out loud
Reading is the easy part. Sit across from a senior AI interviewer that talks, watches your whiteboard, and scores you like the real thing — your first mock is free.
Start a free mock →