TT Lab
Get started
Learn Learning paths Courses

Database Concepts

NoSQL — The Schema Did Not Vanish; the Responsibility Moved

Continue in TT Lab

In a nutshell

NoSQL is not a superset of the relational model but a trade that gives up certain constraints to gain certain capabilities, and if you choose without knowing what you gave up, you pay the price later.

Why this was needed

Relational databases are powerful but assume several things: that the schema is fixed in advance, that the data fits on one machine or at least within a range where joins are possible, and that strong consistency is needed.

In the late 2000s, as web services could no longer meet these assumptions one by one, other options appeared. The key motivation was horizontal scaling. To split across several machines instead of making one machine bigger, features such as joins and transactions became obstacles, and designs that gave them up emerged.

How it works

The characteristics by category divide as follows.

Category Data model Strength Price
Key-value One value per key Very fast simple lookups Cannot query inside a value
Document JSON-like documents Flexible schema, everything in one document for retrieval Joins and consistency between documents are the app's job
Wide column Row key + column families Bulk writes, time series Query patterns must be decided in advance for the design
Graph Nodes and edges Relationship traversal Weak at bulk aggregation

Here we must point out the most widespread misconception. "No schema" does not mean the schema has disappeared; it means the party that enforces the schema has moved from the database to the application. A document with a typo in a field name goes in silently, and months later the code that reads that field fails silently. What the database used to block must now be blocked by code review and tests.

CAP is also often misused. It is a question of which to choose between consistency and availability at the moment a network partition actually occurs, not a story about choosing two of the three in normal times. When there is no partition, most systems provide both.

What it looks like in the field

The criterion for choosing is not fashion but the query pattern. There are a few practical signals.

You must also take into account that modern relational databases have absorbed a good part of this. PostgreSQL's jsonb provides a good deal of the flexibility of a document store while keeping transactions and joins. But you must not overuse it. The practical boundary is to promote to columns the fields that you query, sort by, or need constraints on, and leave only the remaining raw data in jsonb. This is because values inside jsonb have no accurate statistics, so the planner's row count estimates go off, and as a result the choice of join method goes wrong.

Five questions to answer before choosing

Rather than memorizing the characteristics of each category, it gets you to the answer much faster to first write down the following five things about what you are about to build.

  1. What queries will you read with. Do you look up by a single key, filter by conditions, or follow relationships? Document stores and wide-column stores read well only along the access paths decided when storing, so if this answer is "I don't know yet", relational is safe.
  2. How large will the data grow. If it fits on one machine, there is no reason to pay the price of distribution. A single server today handles several terabytes.
  3. What must change together. If this list is long, it means you need transactions.
  4. Is it acceptable to show an old value for a moment. View counts and recommendation lists are fine, but balances and stock are not.
  5. Who will operate it. A new store brings new backup procedures, new metrics, and new kinds of failures along with it.

If you answer the five questions, in most cases the conclusion converges to one. Start with a single relational database, and when a specific query pattern clearly starts to hurt, split out only that part into a different store. For example, sessions move to a key-value store and full-text search to a search engine. The opposite order, that is, starting by putting in several stores from the beginning, is a choice that pays the operational burden up front for a problem that does not even exist yet.

What changes the moment you distribute

When a database that ran on one machine is split across several, several familiar properties quietly disappear. If you move without knowing this, the code is the same but only the results look strange.

First, you may not be able to read back a value right after writing it. In a store with several replicas, writes may go to one and reads to another. This is where the accident comes from in which a screen that saves a profile and immediately views it shows the old value. Most stores offer "what I just wrote, I can read" as an option, so you can turn on that option or send only the reads right after a write to the primary replica. The option name differs by product, but the price is the same: that read gets slower.

Second, changing several documents at once becomes hard. A job that finished in a single transaction in a relational database can succeed only partially if the documents are scattered across different nodes. So document stores are designed on the principle of putting what must change together in the same document. It is important that the reason for putting an order and its order items in one document is not performance but atomicity.

Third, you cannot change the shard key later. The key that decides which node data goes to is in effect an irreversible decision. If you choose wrongly, the load concentrates on particular nodes only (a hot shard), and adding more of those nodes does not solve it. The typical failure is using time as the shard key. All recent data concentrates on one node, so even if you increase the nodes to ten, nine sit idle.

To sum up: the decision to move from relational to NoSQL is not about choosing a store but about deciding which of the rules the database used to keep for you the application will now keep itself. If you move without writing that list down, you will discover them one by one as accidents after the move.

What to check in the quiz that follows

Check whether you can explain what each category gives up and what it gains, and what "no schema" actually means.