TT Lab
Get started
Learn Learning paths Courses

Designing a Log Pipeline

Why You Split the Index by Date

Continue in TT Lab

In one line

The reason to split a log index by date, as in logs-2026.08.21, is so you can delete it. Deleting documents one by one is slow, but deleting a whole index is a file deletion and finishes immediately.

Why this was needed

Deleting a document in a search engine is not really deleting it. It marks the document as "deleted," and it actually drops out later when segments are merged. So if you delete 100 million old log entries with delete_by_query, it takes hours and the cluster is slow the whole time.

If you split indexes by date, "delete August 1" is a single DELETE logs-2026.08.01. It costs the same as deleting a directory.

How it works

An index template automatically applies settings to newly created indexes.

PUT _index_template/labhub-logs
{
  "index_patterns": ["labhub-*"],
  "template": {
    "settings": {
      "number_of_shards": 1,
      "number_of_replicas": 1,
      "refresh_interval": "10s"
    },
    "mappings": {
      "dynamic": false,
      "properties": {
        "@timestamp": { "type": "date" },
        "level":      { "type": "keyword" },
        "message":    { "type": "text" },
        "trace_id":   { "type": "keyword" },
        "pod":        { "type": "keyword" }
      }
    }
  }
}

"dynamic": false is the heart of this configuration. The reason is below.

ISM (lifecycle management) automates the deleting.

hot (7일)  →  warm (23일, 복제본 0, 강제 병합)  →  delete

Mapping explosion

If you leave dynamic on, a mapping is created automatically every time a new field arrives. But if the application puts a user ID into the log as a field name (user_12345: "..."), or puts a whole error object in as JSON, the fields grow into the thousands.

As fields increase, the cluster state grows, and since it is replicated to every node, the whole cluster slows down. In severe cases the master node dies. This is called a mapping explosion.

There are three ways to prevent it.

  1. "dynamic": false — a field you have not defined is stored but not indexed (it can't be searched but can be retrieved)
  2. "dynamic": "strict" — a field you have not defined is rejected when it arrives
  3. index.mapping.total_fields.limit — sets an upper bound (default 1000)

strict looks safe but is dangerous. If the collector attaches one unexpected field, the whole document is rejected and the log is lost. In a real case, Fluent Bit attached an internal field _p, a 400 error went into an endless retry loop, and the logs were blocked entirely. That is why false is better for log indexes.

Common misconceptions

Thinking more shards make it faster. One shard is one Lucene index, and each uses file handles and memory. For logs of a few GB per day, 5 shards is a waste. The rule of thumb is 10–50GB per shard, and if it is smaller than that, reduce the shards.

Mixing up text and keyword. text is analyzed and split into tokens, so full-text search works but aggregation and sorting do not. keyword is stored whole, so aggregation and sorting work but partial search does not. Level, Pod name, and trace ID are keyword; the message body is text.

Build retention and cost into the design

Half of log index design is search performance, and the other half is when to throw what away. If you don't decide the latter, one day the disk fills up, and then you delete in a hurry and end up deleting what you need too.

Divide it into stages. Keep the recent data on fast disk, older data on slower storage, and still older data in object storage, and delete it at the end. Index lifecycle management (ILM) does this transition automatically.

Stage Period (example) What happens
hot 0–3 days Both writes and searches happen
warm 3–14 days Read-only. Segments are merged and replicas are reduced
cold 14–90 days Moved to slower nodes or object storage
delete 90+ days Deleted

Splitting shards too finely is the most common mistake. Each shard comes with memory and file handles, so thousands of small shards slow the cluster down. Aim for 10–50GB per shard, and if a day's data is smaller than that, create indexes weekly instead of daily, or leave it to the data stream's rollover condition (max_primary_shard_size).

Don't make every field searchable. For a field that is only viewed and never searched, setting index: false reduces indexing cost and storage. Conversely, a number used only for aggregation needs only doc_values.

Don't put logs of different shapes into one index. If field names overlap but the types differ, indexing is rejected, and that document is silently lost. Split indexes per service, and set a naming convention for the common fields.

The cost usually comes from indexing, not from searching. Documents per second is CPU. If you send debug-level logs to production as they are, you pay that cost every day. Apply sampling, or move toward using traces for normal requests and keeping logs only for unusual cases.

What really matters in practice

Putting the trace ID in the log is half of observability. With it, you can find "the logs left by this slow request" in one go. Without it, you have to feel your way by time and Pod name.

It is best to follow a standard for field names from the start. ECS (Elastic Common Schema) fixes names such as @timestamp, log.level, service.name, and trace.id. If you want to change later, you have to reindex all the old indexes.