Load Balancer System Design: A Walkthrough With Diagram

By Aniruddha · September 28, 2026 · 8 min read

Every system design answer that says "add more servers" needs a load balancer in front of them. It is one of the first boxes you draw, and interviewers like to dig into it: how does it choose a server, what happens when one dies, and what breaks when requests move between servers?

Below is a 50 second animation of the whole thing, drawn in CanvasKeep: three servers, two routing algorithms, a server failing its health check, and a user getting logged out. The rest of this post explains each step and the parts a short video skips.

1. Why one server is not enough

With one server, every request lands on the same box. You can scale it up with more CPU and memory, but that gets expensive fast, it has a hard ceiling, and it is a single point of failure: if it dies, everything is down.

Scaling out means running many smaller servers and putting a load balancer in front. Clients see one address; the load balancer spreads requests across the servers behind it. Need more capacity? Add a server.

2. The diagram

All traffic enters through the load balancer, which picks a server for each request. The servers keep no session state of their own; they read it from a shared store.

Load balancer system design diagram: a client sends all traffic to a load balancer, which routes to servers A, B and C, and every server reads sessions from a shared Redis session store
Click to open full size.

Get the editable diagram

An .excalidraw file. Open it on excalidraw.com, or in a CanvasKeep diagram with Open from the menu.

Download

3. Layer 4 vs Layer 7

  • Layer 4 (transport). Routes by IP and port and never looks inside the request. Very fast, works for any TCP or UDP protocol.
  • Layer 7 (application). Reads the HTTP request, so it can send /api to one pool and /img to another, terminate TLS, add headers, and retry on another server. Most web apps use this.

Common choices: NGINX and HAProxy you run yourself, or managed ones like AWS ALB (Layer 7) and NLB (Layer 4).

4. Picking a server

Round robin

Servers take turns: A, B, C, A, B, C. Simple and even when requests are similar. But it is blind: a server stuck on slow requests still gets its turn.

Least connections

Send the request to the server with the fewest open connections. In the video, B is busy with three slow uploads, so new requests go to A and C. This handles uneven request times much better.

Others worth knowing

  • Weighted round robin. Bigger servers get proportionally more turns.
  • IP hash or consistent hashing. The same client or key always maps to the same server. Useful for caches, and it gives stickiness without cookies.
  • Power of two choices. Pick two servers at random and use the less busy one. Close to least connections, with far less coordination.

5. Health checks

The load balancer pings each server every few seconds, for example GET /health. A server that fails a few checks in a row is taken out of rotation, and traffic flows to the rest. When it passes again, it comes back.

GET /health
Server A  200
Server B  200
Server C  timeout   -> out of rotation

These are active checks. Many load balancers also do passive checks: if real requests to a server start failing, it is marked down without waiting for the next ping. Before removing a server on purpose (a deploy, say), drain it: stop sending new requests and let in-flight ones finish.

6. The catch: sessions

A user logs in on server A, and A keeps the session in memory. The next request goes to B, which has never heard of them: 401, logged out. Two ways to fix it:

  • Sticky sessions. The load balancer pins each user to one server, usually with a cookie. Easy, but load gets uneven, and if that server dies its users lose their sessions.
  • A shared session store (what the diagram uses). Sessions live in Redis, and any server can read them. Servers become stateless, so any of them can handle any request and you can add or remove them freely.

Signed tokens such as JWTs are a third option: the session travels with each request, so there is nothing to store, at the cost of making logout and revocation harder.

7. What the 60 second version leaves out

  • The load balancer as a single point of failure. Run an active and a standby pair with a floating IP, or use a managed service.
  • TLS termination. Layer 7 load balancers usually decrypt HTTPS so the servers behind them do not have to.
  • Global load balancing. Across regions, DNS-based routing or anycast sends users to the nearest healthy region before any regional load balancer sees them.
  • Slow start. A server that just came back should get a little traffic at first, not its full share, while its caches warm up.
  • Autoscaling. Pair the load balancer with an autoscaler that adds servers under load and registers them automatically.

Interview checklist

  • Explain scale up vs scale out and the single point of failure.
  • Choose Layer 4 or Layer 7 and say why.
  • Pick an algorithm: round robin for similar requests, least connections for uneven ones.
  • Add health checks and connection draining.
  • Make servers stateless with a shared session store.
  • Make the load balancer itself redundant.

Load balancers sit at the top of most designs, including the URL shortener and the rate limiter.

FAQ

What is the difference between a Layer 4 and a Layer 7 load balancer?

A Layer 4 load balancer routes by IP address and port without reading the request, which makes it fast and protocol-agnostic. A Layer 7 load balancer reads the HTTP request, so it can route by URL path, host or headers, terminate TLS, and retry failed requests, at some extra cost per request.

Round robin or least connections?

Round robin is fine when requests are short and similar, since every server gets an equal share. Least connections is better when request times vary a lot, like uploads or slow queries, because it sends new work to whichever server is least busy right now.

What are sticky sessions?

The load balancer sends all of a user's requests to the same server, usually by setting a cookie. It keeps in-memory sessions working, but load becomes uneven and users on a server that dies lose their session. Storing sessions in a shared store like Redis avoids both problems.

Isn't the load balancer a single point of failure?

It would be, so it is run in pairs or clusters. A common setup is an active and a standby load balancer sharing a floating IP that moves to the standby if the active one fails. Managed load balancers from cloud providers handle this for you.

Sources and further reading

More system design in 60 seconds

The diagram was drawn in CanvasKeep, which keeps Excalidraw diagrams in workspaces and folders with version history.

Keep your system design diagrams in one place

The Excalidraw canvas, with workspaces, nested folders, and named versions. Free to start.

Start drawing