How do you prevent a cache stampede or thundering herd?
A stampede happens when a popular key expires and many clients simultaneously hit the database to rebuild it. Several mitigations exist.
- Locking or mutex: only the first request rebuilds while others wait briefly or serve stale data, using SET NX with a short TTL as a rebuild lock.
- Probabilistic early expiration: refresh before expiry with random jitter so rebuilds spread out.
- Stale-while-revalidate: serve the expired value while one worker refreshes in the background.
- TTL jitter: add a random offset so many keys do not expire at once.
- Request coalescing: deduplicate concurrent misses for the same key in-process.
- Never expire hot keys, refreshing them with a scheduled job.
Combine TTL jitter with a rebuild lock for most read-heavy systems, and monitor hit rate and database load during incidents.