Design a scalable distributed rate limiting middleware that restricts incoming requests per user/IP (e.g., max 100 requests per minute) with sub-2ms overhead across global datacenters.
Makes REST/GraphQL API calls with Authorization header or API Key.
Intercepts incoming requests and executes Rate Limiter plugin before forwarding to backend services.
Stores tier configuration rules (e.g. Free Tier = 100 req/min, Paid = 10,000 req/min).
Atomic Redis Lua scripts managing token refill and counter decrements in memory.
Processes business logic after request passes rate limiter check.
Request arrives with Header Authorization: Bearer token123
Pass client_id, requested_tokens, bucket_capacity, refill_rate
Atomic evaluation completed in < 1ms.
Allowed! Forward request to internal service.
Blocked! Return Retry-After header immediately without overloading backend.