TT Lab
Get started
Learn Learning paths Courses

Object Storage and S3

Putting the Same Workload on Two Stores and Comparing

Continue in TT Lab

Goal

Put the same corpus into three places — a filesystem, S3-compatible storage (SeaweedFS S3 gateway) and the SeaweedFS filer — and measure count, time and latency yourself, to learn how to choose storage by numbers rather than preference.

Why it matters

Storage selection is not settled by copying down a benchmark table. The same product can be the best or the worst depending on the workload. A difference that splits at 2,000 files of 4KiB disappears with a single 200MiB file, and the reverse also happens. This lab makes you measure and confirm that fact yourself. In particular, when you count yourself in step 4 how many volume files the 2,000 files end up in, the sentence "it is strong with small files" turns into a concrete structural story. The final decision table is made in a form you can take out and use as it is when you make a real choice later.

One thing to note about the comparison subject. The S3 side of this lab was originally MinIO. Because MinIO keeps each object as one directory on disk plus a metadata file (xl.meta), it bears the filesystem cost measured in step 2 for every object — it was the typical opposite of the small-file problem. In 2025 the community edition ended binary distribution and was dropped from this image, and the S3 server in step 3 is now SeaweedFS's S3 gateway. So what this lab measures now is not a contest between two storage engines but two layers of small-file cost. The filesystem (step 2) shows the cost of "number of files on disk," and S3 (step 3) and the filer (step 4) knock on the same SeaweedFS through two doors to show the cost of "number of requests." In step 3 you also confirm that the number of volume files does not grow even when you put in via S3.

Steps

  1. Create 2,000 files of 4096 bytes under /root/cmp/corpus/. The names run from f0000.bin to f1999.bin.
  2. In /root/cmp/fs.txt, write three lines: files=2000, data_bytes=8192000 and disk_kb=<du -sk 결과> (with the output of du -sk in the placeholder). disk_kb must be greater than the data size.
  3. On the S3 server (http://127.0.0.1:9000), create the bucket lab-cmp, upload the whole corpus, and in /root/cmp/s3.txt write objects=2000 seconds=<소수> (with a decimal number in the placeholder).
  4. Put the same corpus under the SeaweedFS filer's /cmp/, and in /root/cmp/sw.txt write files=2000 dat_files=<n> seconds=<소수> (with a decimal number in the placeholder). dat_files must be 20 or less.
  5. Read 100 random files from each store and, in /root/cmp/latency.csv, write a store,avg_ms header and three rows, fs, s3 and filer.
  6. Put one 200MiB file into the three stores and read it, and in /root/cmp/bigfile.csv write a store,put_ms,get_ms header and three rows.
  7. In /root/cmp/decision.md, write a Markdown table. The row titles are the four 소파일 대량 (many small files), 대용량 소수 (few large files), S3 API 필요 (S3 API needed) and 운영 인력 (operations staff), and there must be a recommendation column.

Notes

Create a corpus for comparison

Create 2,000 files of 4096 bytes under /root/cmp/corpus/. The names run from f0000.bin to f1999.bin.

You have to put the same data in three places for a comparison to work. Match the size and count exactly.

Put it on the filesystem and measure the metadata cost

In /root/cmp/fs.txt, write three lines: files=2000, data_bytes=8192000 and disk_kb=<du -sk 결과> (with the output of du -sk in the placeholder). disk_kb must be greater than the data size.

The data size and the blocks actually used are different. Count the number of files as well.

Put it into S3 and measure the time taken

On the S3 server (http://127.0.0.1:9000), create the bucket lab-cmp, upload the whole corpus, and in /root/cmp/s3.txt write objects=2000 seconds=<소수> (with a decimal number in the placeholder).

Each object is a request. Check whether the count becomes the time, and whether the S3 server's volume files increase.

Put it into SeaweedFS and count the volume files

Put the same corpus under the SeaweedFS filer's /cmp/, and in /root/cmp/sw.txt write files=2000 dat_files=<n> seconds=<소수> (with a decimal number in the placeholder). dat_files must be 20 or less.

Use the daemons you started in the previous lab as they are. The number of volume files is the key thing to observe.

Compare random read latency

Read 100 random files from each store and, in /root/cmp/latency.csv, write a store,avg_ms header and three rows, fs, s3 and filer.

Read the same 100 in random order. Collect the three values in one file.

Do the same comparison with one large file

Put one 200MiB file into the three stores and read it, and in /root/cmp/bigfile.csv write a store,put_ms,get_ms header and three rows.

What happens to the differences that split with small files when the file is large is half of the conclusion.

Write a selection criteria table

In /root/cmp/decision.md, write a Markdown table. The row titles are the four 소파일 대량 (many small files), 대용량 소수 (few large files), S3 API 필요 (S3 API needed) and 운영 인력 (operations staff), and there must be a recommendation column.

Organize in a table which one to choose for each workload characteristic. The row titles are the grading criteria.