TT Lab
Get started
Learn Learning paths Courses

Redis and Caching

Six Data Structures and Where Each Belongs

Continue in TT Lab

Summary

Using Redis well does not mean knowing many commands; it means choosing the data structure that fits the problem.

Why this was needed

Many teams use Redis with nothing but String. They serialize everything as JSON and SET and GET it. It works. But even to change a single field such as a user profile's email, you read the whole thing, parse it, modify it, and write it back. If two requests do this at the same time, one overwrites the other.

With a Hash, it is a single line, HSET user:42 email a@b.com, and it does not conflict with the other fields. One choice of data structure eliminates the race condition.

How it works

A String is a simple value and a counter. It matters that INCR is atomic — use it for view counts, rate-limit counters, and sequence numbers. SET key val NX EX 60 is the basic form of a distributed lock.

A Hash is an object with several fields. You can update and read at the field level, and when there are few fields Redis internally uses a compact encoding, which also saves memory.

A List is an ordered list, and insertion and removal at both ends are O(1). Use it for job queues and recent-activity lists. Access in the middle is O(N), so do not search a large list by index.

A Set is a collection without duplicates, and intersection, union, and difference are a single command each. Use it for tags, followers, and deduplication. The return value of SADD tells you "is this the first time seeing it", so it is also used to implement idempotency.

A ZSet (sorted set) is a set whose members have scores. It always stays sorted by score, and rank lookup is O(log N). Use it for rankings, priority queues, and time-series windows (with the score as a timestamp). It is important enough that this course devotes an entire lab to it.

A Stream is an append-only log. It has consumer groups and acknowledgements, so it is the closest thing to a real queue. It was covered in the previous course.

What you meet in the field

The command to be most careful with in production is KEYS. It scans the whole keyspace, and during that time Redis can do nothing else. Because Redis processes commands on a single thread, a single KEYS * on an instance with millions of keys causes a total freeze of several seconds. You must iterate with SCAN using a cursor.

For the same reason, be careful with FLUSHALL, SMEMBERS / LRANGE 0 -1 on large collections, and deleting huge keys with DEL (use UNLINK instead). A good share of "Redis suddenly got slow" cases are a single O(N) command.

It is also worth knowing the memory side. MEMORY USAGE <key> shows the actual usage per key, and even for the same data it can differ several-fold depending on the data structure.

A table for deciding what to choose

Data structure Where to use it Representative commands Caution
String Cache values, counters SET, INCR 512MB limit
Hash Per-field updates of an object HSET, HGETALL With many fields, HGETALL is heavy
List Queues, the latest N items LPUSH, BRPOP Insertion and lookup in the middle are O(n)
Set Deduplication, tags SADD, SINTER Intersection of large sets is expensive
Sorted Set Rankings, time-ordered indexes ZADD, ZRANGEBYSCORE The most broadly useful
Stream Event logs, consumer groups XADD, XREADGROUP Better suited to queues than a List

Sorted Set is used surprisingly widely. If you use a timestamp as the score, you get time-range queries (ZRANGEBYSCORE), and it is also easy to delete old entries (ZREMRANGEBYSCORE). If you try to build something like "events from the last 24 hours" with a List, you will soon get stuck.

For queues, Stream rather than List

If you build a queue with a List, the message disappears the moment you pop it with BRPOP. If the consumer dies during processing, that message is gone.

A Stream has the concepts of consumer groups and acknowledgement (ack).

XADD  orders * type payment amount 52000      # 발행
XREADGROUP GROUP workers w1 COUNT 10 STREAMS orders >   # 읽기(pending 으로 표시)
XACK  orders workers <id>                      # 처리 완료
XPENDING orders workers                        # 아직 확인 안 된 것
XCLAIM orders workers w2 60000 <id>            # 죽은 소비자의 것을 가져오기

XPENDING is the key. If a consumer dies, its messages remain in pending, and another consumer takes them with XCLAIM. With a List, you would have to build this yourself.

However, a Stream also grows without limit, so set a cap such as XADD ... MAXLEN ~ 100000. With ~ it trims approximately, which is much cheaper.

Problems caused by large keys

Redis processes commands on a single thread. If one takes a long time, everything else stops in the meantime.

KEYS *                      → O(n). 절대 쓰지 않는다
HGETALL (필드 10만 개)      → 응답이 크고 오래 걸린다
SMEMBERS (원소 100만 개)    → 같은 문제
DEL (큰 컬렉션)             → 삭제도 O(n) 이다. UNLINK 를 쓴다

UNLINK hands the deletion off to the background and avoids blocking. Use it when deleting large keys.

To find large keys, use redis-cli --bigkeys or --memkeys. Run them regularly to check that no single key takes up a large share of the total.

What you will do in the next lab

You work with the six data structures one by one, replace KEYS with SCAN iteration, and finally build a table of the memory difference when the same data is stored in different data structures.