The Haystack Model and Master, Volume, Filer
In one line
SeaweedFS packs millions of small files into a few large container files. There are no inodes per file, and no metadata lookups per file.
Why it was needed
Many people have experienced what happens when you put a million thumbnails on a filesystem. Disk usage is small, but inodes run out, ls takes minutes, and backup tools spend days stat-ing each file. A million 4KB files are 4GB as data but a million metadata entries to the filesystem.
Object storage does not solve all of this either. A metadata record is created for each object and requests are also per object, so the count becomes both cost and load.
SeaweedFS uses an approach that comes from Facebook's Haystack paper. It writes small files one after another inside a large volume file and remembers only their location (offset and length). If 1,000 files go into one volume file, there is one inode, and a read is a single offset seek.
How it works
There are three kinds of daemons.
The master (default 9333) manages only the placement of volumes. It decides which volume is on which volume server and where to write new files. What matters is that it does not hold file metadata — so the master does not become a bottleneck.
The volume server (default in the 8080 range) holds the actual data. It keeps a volume file (.dat) and an index (.idx) as a pair, and finds an offset by file ID to read.
The filer (default 8888) is a separate stateless server that adds directory structure and POSIX attributes on top. If you want to write and read by path, you go through the filer. The S3 gateway is also on top of this.
The basic flow is this. You request /dir/assign from the master to receive a file ID (for example 3,01637037d6) and a volume server address, and upload the file to that server. To read, you find the volume location with /dir/lookup and download from that server. These two steps are the core of SeaweedFS, and the filer and S3 gateway are convenience layers on top of them.
What it looks like in the field
Put honestly, the selection criteria are these. If small files are in large numbers and the count itself is the problem, SeaweedFS is strong. If it is small-scale with mostly large files and only an S3 API is needed, a lightweight choice such as Garage is better. If the organization is large and has a storage team, Ceph RGW is the proven choice.
You should also look honestly at the risk. SeaweedFS is on the axis with the most active releases, but its commits are highly concentrated — the founder's commits number 9,968, an order of magnitude different from second place (530). It does not mean the project is dying; it means concern about the bus factor is justified.
And there is a background in which this judgment became actually necessary. The MinIO community edition went through feature reduction in May 2025, distribution stopped in September–October, maintenance mode in December, and the repositories were archived one after another in March–July 2026. The license was AGPLv3 to the end, but the switching cost was passed to users in a form where fixes for high-severity CVEs did not come through the official images. The lesson is not "MinIO was bad" but that with open-source infrastructure where a single vendor holds the CLA, that vendor's business shift is the fate of the project. In a layer where data physically settles, this risk translates directly into migration cost.
What happens when volumes fill up
A structure that puts small files into large volumes comes with operational items unique to that structure. The first one you run into is the limit on the number and size of volumes.
A volume server does not create volumes without limit. The maximum number and the size limit of one volume are fixed, and if there is not a single writable volume, all uploads fail. The symptom that appears then is nasty: the disk has room and the server is alive, so "why doesn't it work" goes on for a long time. The shortcut to diagnosis is to check whether the placement information the master returns has any writable volumes. Since the limits are values set when you start the server, in capacity planning you must set these two values together with the disk size.
Deleted space does not come back right away. When you delete a file, it is removed from the index, but its place inside the volume file stays. To reclaim it, you have to run a job that compacts the volume, and during that time the volume becomes read-only. So in workloads with many deletions, when compaction runs determines the actual usage.
Replication is per volume. The number of copies is not set per file; the replication scheme of a volume is fixed when the volume is created, so to change it later you have to create a volume with the new setting and move things over. It is a value that needs care when you first decide it.
And you must not forget that the filer's metadata lives in a separate store. Even if you back up the volume data, if you don't take the filer's database with it, the path structure disappears entirely, and you end up in a state where the data exists but you can't tell which file is what. A backup plan must always include both, and recovery tests must also be done by restoring both together.
What we do in the next lab
You start the master, volume and filer yourself, obtain a file ID, upload and look up. You put in 500 small files to check the number of volume files, and in the last lab you put the same corpus into three places — a filesystem, S3 and SeaweedFS — and compare their characteristics in numbers.