Skip to content

Throttling

Deliberately limit how much traffic a client (or the system as a whole) can send, before an overload happens — the proactive counterpart to circuit breakers, which react after things have already started failing.

flowchart LR Junior["Junior: rate limiting vs. reacting to overload after the fact"] --> Middle["Middle: token bucket and leaky bucket algorithms"] Middle --> Senior["Senior: per-client vs. global limits, and fairness"] Senior --> Professional["Professional: distributed rate limiting at scale"]
flowchart LR Requests[Incoming requests] --> Bucket{"Token bucket:\ntokens available?"} Bucket -->|yes| Allow[Allow, consume a token] Bucket -->|no| Reject["Reject (429),\nor queue"]

Choose a level

Level Guide You are done when
Junior Proactive limiting vs. reactive failure You can explain why throttling is applied before overload, not after.
Middle Token bucket and leaky bucket You can trace both algorithms and explain when each fits better.
Senior Per-client limits and fairness You can design a rate-limiting scheme that prevents one client from starving others.
Professional Distributed rate limiting You can design a rate limiter that's consistent across many stateless service instances.

Practice rule

For any public or shared API, ask: "if one client sent 1,000x their normal traffic right now, what would happen to every other client?" If the answer is "they'd all be affected," you need throttling, not just capacity.