Rate Limiting Explained: How APIs Protect Themselves from Too Many Requests

What rate limiting is, why APIs use it, how common algorithms work, and what the HTTP 429 status code means for developers.

Share on Linkedin Share on WhatsApp

Estimated reading time: 7 minutes

Article image Rate Limiting Explained: How APIs Protect Themselves from Too Many Requests

Every time an app loads your feed, checks the weather or processes a payment, it is likely sending requests to an API. Behind the scenes, the server receiving those requests needs a way to make sure no single user, bug or attacker can overwhelm it. That mechanism is called rate limiting. If you are learning backend development or just want to understand why you sometimes see a “Too Many Requests” message, this guide explains the idea in plain language.

What is rate limiting?

Rate limiting is a technique that controls how many requests a client can make to a service within a given period of time. For example, an API might allow 100 requests per minute for each user. Once a client goes over that limit, the server stops processing the extra requests for a while and returns an error instead.

Think of a restaurant that can only cook so many meals per hour. If every customer ordered at once, the kitchen would collapse and everyone would wait. Limiting the pace of orders keeps the service working for all customers.

Why APIs need it

  • Protecting availability: a sudden burst of traffic, whether malicious or accidental, can slow down or crash a service.
  • Fair usage: it prevents one heavy user from consuming resources that others need.
  • Security: it makes brute-force attacks, such as trying many passwords in a row, much slower and easier to detect.
  • Cost control: servers, databases and third-party services cost money per use, so limits keep the bill predictable.
  • Business plans: many providers offer different limits for free and paid tiers.

What the client sees: HTTP 429

When a client exceeds the limit, a well-designed API responds with the HTTP status code 429 Too Many Requests. Many APIs also add helpful headers, such as Retry-After, which tells the client how many seconds to wait before trying again. Some services also expose headers that show the limit, how many requests remain and when the counter resets, although the exact names vary between providers, so always check the documentation of the API you use.

Common rate limiting algorithms

There are several ways to implement a limit. Each balances simplicity, accuracy and flexibility differently.

AlgorithmHow it worksStrengthWeakness
Fixed windowCounts requests in fixed time blocks, such as each minuteVery simpleBursts at the edges of two windows can briefly double the traffic
Sliding windowCounts requests over a moving period of timeSmoother and more accurateNeeds more memory and computation
Token bucketA bucket fills with tokens at a steady rate; each request uses one tokenAllows short bursts while keeping a long-term averageSlightly more complex to tune
Leaky bucketRequests enter a queue and are processed at a constant rateProduces a steady output flowRequests can wait in the queue or be dropped when it is full

The token bucket is especially popular. Imagine a bucket that receives one token every second, up to a maximum capacity. Each request consumes one token. If the bucket is full, a user who has been idle can send a quick burst of requests. When the bucket runs out, requests must wait for new tokens.

Who is being limited?

Before choosing a limit, a team has to decide what to count it against. This is called the key of the limit. Common choices are:

  • The IP address of the client, which is simple but can penalize many users behind the same network.
  • The user account or API key, which is fairer for authenticated services.
  • A specific endpoint, since some operations (like login or search) are more sensitive or expensive than others.
  • A combination of these, for example a limit per user plus a global limit for the whole service.

Where it is implemented

Rate limiting can live in different layers. Many teams apply it at an API gateway or reverse proxy in front of the application, so excess traffic never reaches the code. Others implement it inside the application, often using a fast shared store such as an in-memory database so that several server instances can share the same counters. In practice, it is common to combine both: a coarse limit at the edge and finer rules inside the application.

How to be a good API client

If you consume APIs, you will eventually hit a limit. Good habits make this painless:

  • Read the documentation to know the limits before you build.
  • Respect Retry-After when it is provided, instead of retrying immediately.
  • Use exponential backoff: wait a little after the first failure, longer after the second, and so on, ideally with some random variation to avoid many clients retrying at the same moment.
  • Cache responses that do not change often to reduce unnecessary requests.
  • Batch requests when the API supports it.

Tips for designing limits

  • Start from real usage: look at how legitimate users behave and set limits comfortably above that.
  • Apply stricter limits to sensitive actions, such as login attempts and password resets.
  • Return clear error messages and a proper 429 status so clients can react correctly.
  • Monitor how often limits are triggered, since frequent hits may signal a bug, an attack or limits that are too tight.
  • Document the rules so developers know what to expect.

Conclusion

Rate limiting is a small idea with a big impact: it keeps services stable, fair and safer, and it shapes how clients should behave. Understanding it helps you build more reliable APIs and write client code that handles limits gracefully. If you want to go deeper into backend development, from APIs to databases and security, the programming courses on Cursa are a good next step.

NTFS, exFAT, FAT32 and APFS: Choosing the Right File System for a Drive

Understand what a file system does and how NTFS, exFAT, FAT32, APFS and ext4 differ, so you can format drives without losing compatibility.

Text Encoding Explained: ASCII, Unicode and Why You Sometimes See Strange Symbols

Learn how computers store text, what ASCII and Unicode actually are, why UTF-8 became the standard, and how to fix files that display garbled characters.

Idempotency in APIs: Why Retrying a Request Should Be Safe

Learn what idempotency means in backend development, which HTTP methods provide it, and how idempotency keys prevent duplicate operations.

What Is a CDN? How Content Delivery Networks Make Websites Fast

Learn what a CDN is, how edge caching and cache headers work, what a cache hit means, and when a CDN helps — or does not.

Semantic Versioning Explained: What a Number Like 2.4.1 Actually Tells You

MAJOR.MINOR.PATCH is a promise, not decoration. Learn to read version numbers and understand dependency range symbols.

What Is a Virtual Machine? Virtualization Explained for Beginners

Learn what a virtual machine is, how hypervisors work, how VMs differ from containers, and when to use each one.

How HTTPS Works: Certificates, the TLS Handshake and What the Padlock Really Means

A beginner-friendly walkthrough of HTTPS: what TLS certificates prove, how the handshake works, and what the browser padlock does not guarantee.

Big O Notation Explained: How to Talk About Code Efficiency

A beginner-friendly guide to Big O notation: what it measures, the most common complexity classes, and how to reason about the cost of your code.