Putting the Same Workload on Two Stores and Comparing
Goal
Put the same corpus into three places — a filesystem, S3-compatible storage (SeaweedFS S3 gateway) and the SeaweedFS filer — and measure count, time and latency yourself, to learn how to choose storage by numbers rather than preference.
Why it matters
Storage selection is not settled by copying down a benchmark table. The same product can be the best or the worst depending on the workload. A difference that splits at 2,000 files of 4KiB disappears with a single 200MiB file, and the reverse also happens. This lab makes you measure and confirm that fact yourself. In particular, when you count yourself in step 4 how many volume files the 2,000 files end up in, the sentence "it is strong with small files" turns into a concrete structural story. The final decision table is made in a form you can take out and use as it is when you make a real choice later.
One thing to note about the comparison subject. The S3 side of this lab was originally MinIO. Because MinIO keeps each object as one directory on disk plus a metadata file (xl.meta), it bears the filesystem cost measured in step 2 for every object — it was the typical opposite of the small-file problem. In 2025 the community edition ended binary distribution and was dropped from this image, and the S3 server in step 3 is now SeaweedFS's S3 gateway. So what this lab measures now is not a contest between two storage engines but two layers of small-file cost. The filesystem (step 2) shows the cost of "number of files on disk," and S3 (step 3) and the filer (step 4) knock on the same SeaweedFS through two doors to show the cost of "number of requests." In step 3 you also confirm that the number of volume files does not grow even when you put in via S3.
Steps
- Create 2,000 files of 4096 bytes under
/root/cmp/corpus/. The names run fromf0000.bintof1999.bin. - In
/root/cmp/fs.txt, write three lines:files=2000,data_bytes=8192000anddisk_kb=<du -sk 결과>(with the output of du -sk in the placeholder). disk_kb must be greater than the data size. - On the S3 server (
http://127.0.0.1:9000), create the bucketlab-cmp, upload the whole corpus, and in/root/cmp/s3.txtwriteobjects=2000 seconds=<소수>(with a decimal number in the placeholder). - Put the same corpus under the SeaweedFS filer's
/cmp/, and in/root/cmp/sw.txtwritefiles=2000 dat_files=<n> seconds=<소수>(with a decimal number in the placeholder). dat_files must be 20 or less. - Read 100 random files from each store and, in
/root/cmp/latency.csv, write astore,avg_msheader and three rows,fs,s3andfiler. - Put one 200MiB file into the three stores and read it, and in
/root/cmp/bigfile.csvwrite astore,put_ms,get_msheader and three rows. - In
/root/cmp/decision.md, write a Markdown table. The row titles are the four소파일 대량(many small files),대용량 소수(few large files),S3 API 필요(S3 API needed) and운영 인력(operations staff), and there must be a recommendation column.
Notes
- Creating files:
dd if=/dev/urandom of=f0000.bin bs=4096 count=1or a Python loop - Disk usage:
du -sk /root/cmp/corpus - Timing: the difference before and after
date +%s.%N, ortime - S3 side: the credentials are in
/opt/fixtures/s3/creds.env, and, as in step 1 ofs3-basics, you create the profilelocal. For uploading useaws s3 cp --recursive; for reading, boto3 is convenient. - The two SeaweedFS instances are different processes. The S3 server is the one the lab environment started (data in
/var/lib/lab-s3/data), and the master, volume and filer of step 4 (9333, 8180, 8888) are the ones you start yourself. - Common mistake 1: comparing after putting different data into the three stores.
- Common mistake 2: mixing measurements of a warm cache and a cold cache — keep the order constant.
Create a corpus for comparison
Create 2,000 files of 4096 bytes under /root/cmp/corpus/. The names run from f0000.bin to f1999.bin.
You have to put the same data in three places for a comparison to work. Match the size and count exactly.
Put it on the filesystem and measure the metadata cost
In /root/cmp/fs.txt, write three lines: files=2000, data_bytes=8192000 and disk_kb=<du -sk 결과> (with the output of du -sk in the placeholder). disk_kb must be greater than the data size.
The data size and the blocks actually used are different. Count the number of files as well.
Put it into S3 and measure the time taken
On the S3 server (http://127.0.0.1:9000), create the bucket lab-cmp, upload the whole corpus, and in /root/cmp/s3.txt write objects=2000 seconds=<소수> (with a decimal number in the placeholder).
Each object is a request. Check whether the count becomes the time, and whether the S3 server's volume files increase.
Put it into SeaweedFS and count the volume files
Put the same corpus under the SeaweedFS filer's /cmp/, and in /root/cmp/sw.txt write files=2000 dat_files=<n> seconds=<소수> (with a decimal number in the placeholder). dat_files must be 20 or less.
Use the daemons you started in the previous lab as they are. The number of volume files is the key thing to observe.
Compare random read latency
Read 100 random files from each store and, in /root/cmp/latency.csv, write a store,avg_ms header and three rows, fs, s3 and filer.
Read the same 100 in random order. Collect the three values in one file.
Do the same comparison with one large file
Put one 200MiB file into the three stores and read it, and in /root/cmp/bigfile.csv write a store,put_ms,get_ms header and three rows.
What happens to the differences that split with small files when the file is large is half of the conclusion.
Write a selection criteria table
In /root/cmp/decision.md, write a Markdown table. The row titles are the four 소파일 대량 (many small files), 대용량 소수 (few large files), S3 API 필요 (S3 API needed) and 운영 인력 (operations staff), and there must be a recommendation column.
Organize in a table which one to choose for each workload characteristic. The row titles are the grading criteria.