policybook / rate-limiter

rate-limiter

Policies for deciding whether a request may proceed. They differ far less in *how much* they let through than most comparisons suggest, and far more in what they cost to run and how they behave at the edges.

Run these side by side → Same trace, same seed, stepped in lockstep.

Choosing one

If you needUseBecause

A sensible default

token bucket

Bursts allowed and bounded, no seam, exact retry hint.

The same thing with the least state

GCRA

Identical decisions from one integer per key instead of three: 34 bytes against 42.

Smoothing, not budgeting

leaky bucket

Its default capacity of 1 forces even spacing. At equal parameters it is the token bucket.

An exact “N in any window” guarantee

sliding log

The only policy that makes that sentence literally true, at 834 bytes per key.

To limit across processes without coordination

sliding counter

Epoch-aligned windows shard with two counters and no messages.

The simplest thing a stored procedure can do

fixed window

One INCR with a TTL, and a boundary where twice the limit slips through.

RPM and TPM, the LLM-API shape

dual bucket

Two ceilings checked together, charged atomically.

Benchmarks

Accept rate on each trace, measured by this repository's benchmarks.

Policy steady bursty many-keys overload
Dual bucket 1.00001.00001.00000.6747
GCRA 1.00001.00001.00000.3429
Leaky bucket 1.00001.00001.00000.3429
Token bucket 1.00001.00001.00000.3429
Fixed window 0.99270.96921.00000.3374
Sliding log 0.97720.96921.00000.3374
Sliding counter 0.98470.96921.00000.3373

Every policy