TT Lab
Get started
Learn Learning paths Courses

Redis and Caching

What Keys Without a TTL Create

Continue in TT Lab

Summary

Expiration does not happen immediately when the time comes. Keys are deleted when they are accessed (lazy deletion), or a background process samples keys and deletes them (active expiration).

Why this was needed

There are often situations where a Redis instance's memory keeps growing and you cannot find the cause. Usually the answer is one thing — keys without a TTL are mixed in.

The problem is that this is not noticeable. Cache code mostly uses SETEX, but somewhere one place uses only SET, or updates with SET key val and wipes out the TTL that was there. The second case is especially tricky — SET removes the existing TTL. To keep the TTL, you need the KEEPTTL option.

How it works

You first need to know the expiration mechanism precisely. Redis uses two methods together. Lazy deletion checks expiration and deletes at the moment a key is accessed. Active expiration samples keys that have a TTL at random in the background and deletes the expired ones. So a key whose expiration time has passed but which nobody accesses remains in memory for a while. This is why DBSIZE is larger than expected.

When maxmemory is reached, the eviction policy kicks in. There are eight choices, but three matter in practice.

noeviction is the default, and when the limit is reached it rejects writes. If you use it as a cache with this setting, all writes suddenly fail one day.

allkeys-lru or allkeys-lfu treats all keys as eviction candidates. For a pure cache use, this is the one. The author's recommendation is allkeys-lfu — because frequency matches cache hit rate better than recency does.

volatile-lru evicts only keys that have a TTL set. Use it when cache and permanent data are mixed in one instance. However, with this setting, if you omit the TTL on a cache key, that key consumes memory while never being evicted.

Three ways a cache collapses

TTL design exists to prevent these three. The names are similar and confusing, but the causes and countermeasures differ.

Name What happens How to prevent it
Stampede The moment a popular key expires, thousands of requests hit the DB at once Jitter, locks, early recomputation
Penetration Continually looking up nonexistent keys — nothing remains in the cache, so every time goes to the DB Cache "absent" too, briefly; Bloom filter
Avalanche Many keys expire at once, or the cache server goes down Jitter, spreading TTLs, multi-layer cache

Stampede is prevented with a lock. Only the first request goes to the DB and the rest wait.

def get_or_load(key, loader, ttl=300):
    v = r.get(key)
    if v is not None:
        return v
    # SET NX 로 잠금 하나만 얻는다. 얻지 못한 요청은 잠깐 기다렸다 다시 읽는다.
    if r.set(f"lock:{key}", "1", nx=True, ex=10):
        try:
            v = loader()
            r.set(key, v, ex=ttl + random.randint(-30, 30))   # 지터
            return v
        finally:
            r.delete(f"lock:{key}")
    time.sleep(0.05)
    return r.get(key) or loader()      # 그래도 없으면 어쩔 수 없이 직접

A better method is early recomputation. If you refresh probabilistically before the TTL ends, the moment of expiration itself disappears. The approach that raises the refresh probability as the remaining TTL gets shorter (probabilistic early expiration) is widely used.

Penetration is handled by caching "it does not exist" too. However, you should set it briefly (30–60 seconds) so that it is reflected quickly when real data appears.

The moment the cache and the original drift apart

If there are writes, you need invalidation, and invalidation has an ordering problem.

❌ 캐시를 먼저 지우고 DB 를 쓴다
   1. DEL cache          2. (다른 요청이 옛 값을 읽어 캐시에 다시 넣는다)
   3. UPDATE db          → 캐시에 옛 값이 영원히 남는다

✅ DB 를 먼저 쓰고 캐시를 지운다 (cache-aside)
   1. UPDATE db          2. DEL cache
   → 사이에 읽은 요청은 옛 값을 보지만, 곧 사라진다

Even so, they drift apart on rare occasions. To prevent it completely, use delayed double deletion (after writing and deleting, delete once more a few hundred milliseconds later) or subscribe to a change log (CDC). For most services a short TTL is the cheaper answer — even if they drift, they drift only by the TTL.

What you meet in the field

Adding jitter to the TTL is also important. If you fill the cache all at once right after a deployment, the TTLs are exactly the same, so an hour later they all expire at the same time. That moment is the stampede. One line, ttl = base + random(-spread, +spread), spreads the expirations along the time axis.

You also need auditing. If you periodically run a script that iterates with SCAN and counts keys whose TTL is -1 (no expiration), you catch it right away when newly added code omits the TTL.

What you will do in the next lab

You set a TTL and check it, reproduce the trap where SET wipes out the TTL, block it with KEEPTTL, change the eviction policy and confirm that keys are actually evicted, and finally build an audit script that finds keys without a TTL.