
Top 8 Load Balancer Algorithms You Should Know About
8 ways a balancer picks the next server. Each one has its own picture, under the section that explains it.

8 ways a balancer picks the next server. Each one has its own picture, under the section that explains it.
A health check was green on all three boxes. CPU on the first one was pinned. The office NAT had hashed onto that one server, and every laptop in the building followed it there.
The algorithm was doing what it was told. I had just never watched the requests land.
There are 8 below. Each one gets its own figure. It starts sending when the figure is on screen, and it stops when you scroll away. Speed is under the picture. Reset clears the lap and starts it again.
Request 1 goes to A, 2 to B, 3 to C, 4 back to A. nginx uses this unless you say otherwise. HAProxy calls it balance roundrobin.
Round robin does not look at the servers. It keeps a list, and it walks that list. Request 1 is A, request 2 is B, request 3 is C. Request 4 is A again, because the list ran out.
| Request | Server | Why that server |
|---|---|---|
| 1 | A | First name in the list |
| 2 | B | Next name |
| 3 | C | Next name |
| 4 | A | The list ended, so it starts over |
| 5 | B | Same seat as request 2 |
| 6 | C | Same seat as request 3 |
Six requests is two full laps, so A, B, and C each hold two. Request 7 starts a third lap on A. Nothing in that walk asks whether A is already busy.
It is fair when every request costs about the same. It is a bad fit when one request uploads a video and the next one hits /health. The video stays on A while B and C keep taking cheap checks.
| Question | Answer |
|---|---|
| What it counts | The next name in a fixed list |
| What it ignores | CPU, open connections, response time, and request size |
| nginx | This is the default. You write a directive only when you want a different algorithm |
| HAProxy | balance roundrobin |
| Same client, next request | Can land on a different server. This is not sticky |
| Equal machines, short requests | This is enough |
| One huge upload, then health checks | The upload stays where it landed. The cheap checks keep walking the list |
| A server that is slow | It still gets its turn |
| Against least connections | Least connections skips a busy box. Round robin does not |
Use this when the machines are different sizes. A c7g.xlarge and a t4g.small should not take the same count. nginx writes it as weight=5 on the upstream. HAProxy writes it as weight on the server line.
The lap is fixed. A takes five, then B takes one, then C takes one, then A starts again. A server that finishes early does not get to jump the queue.
The next request goes to the server with the fewest open requests. Long requests pile up, so the balancer stops sending them more work.
nginx calls this least_conn. HAProxy calls it leastconn. An Application Load Balancer does not give you a round-robin dropdown. It uses least outstanding requests, which is this idea: in-flight requests, not a completed count.
Least connections on its own still treats a large box and a small box as equals. Weights fix that. One open request on a weight-5 server is less busy than one open request on a weight-1 server.
HAProxy writes this as balance leastconn plus a weight on each server. The picture still leans on A, and a server that drains its queue can take the next request before the lap says so.
Connections are not the whole story. A server with two open requests that answers in 12 ms can be a better pick than an idle server that answers in 40 ms.
HAProxy calls this least response time. It multiplies open connections by a recent response time and picks the smaller product. I use it when the boxes are the same size on paper and one of them is actually slow.
The same client IP keeps landing on the same server. That is sticky routing without a cookie.
It is also how one office NAT melts a single box. Every laptop shares the public address, so they all hash together. nginx calls this ip_hash. A Network Load Balancer does not hash the IP for the life of a laptop. It picks a target for a flow, and that choice sticks for the life of the connection.
IP hash sticks a client to a server. Consistent hashing sticks a key to a server, and it tries not to move every key when you add or remove a node.
Envoy calls a form of this ring hash. You meet it in caches and in meshes, where the point is that key A stays on the same box for the next request. Maglev, which Google uses in front of a lot of traffic, is the same family.
Pick two servers at random. Send the request to the one with less load. You skip a scan of the whole pool, and you still dodge the server that is already buried.
Plain random is the version that does not take the second look. Power of two is the one worth remembering. Envoy and a lot of service meshes use a form of it. You will not find a radio button with this name on an Application Load Balancer.
If the requests are short and the machines match, round robin is enough. If the machines differ and the requests are short, add weights. If some requests stay open, use least connections, and add weights if the machines differ too. If one box is slow even when it looks idle, look at response time.
If you need a client to stay on one box, know that IP hash will also glue a whole NAT to that box. If you need a cache key to stay put, use consistent hashing. Power of two is what I want inside a mesh, where no human is going to retune weights at 2am.
Hope you enjoyed this one. If you want to argue about sticky sessions versus a shared store, find me on X at https://x.com/harundotdev.
I email when a new post goes up. One send a week, and only if there's something new.
Want this applied on your account? Start with DevOps.
Related reading
Kept below the post instead of in a sidebar, with a slow continuous motion for a cleaner editorial feel.
On this post
Comments
A reply stays under the note it answers.
No comments yet.
If you have a note on Top 8 Load Balancer Algorithms You Should Know About, sign in and leave it.