What Gets Better and What Gets Worse When You Split
Summary
Microservices are not a choice made for performance but a choice made so that teams can deploy without waiting for each other. The price is that what used to be a function call becomes a network call.
Why this was needed
You often hear that monoliths are bad. But if you look at the order in which things actually collapse, code does not come first. Deployment comes first. Eight teams commit to one repository, cutting a release branch takes two days, and one team's bug halts the whole deployment — this is where people start saying "let's split it up".
That diagnosis is generally right. The problem is what comes next. The list of things you lose the moment you split is usually not prepared. Inside a monolith, inventory.reserve(sku, 3) was a function call measured in nanoseconds and could not fail. The moment you split it into services, it becomes an HTTP call measured in milliseconds, and it can time out, can half succeed, and the other side can be restarting. And the transaction that wrapped that call no longer exists.
How it works
There are three criteria for judging.
First, is deployment independence actually needed? If two features are always deployed together, there is almost no gain from splitting them. If you split them and still always release them together, that is a distributed monolith, which adds the drawbacks of a distributed system to those of a monolith.
Second, does the data diverge? If two services keep writing to the same table, the boundary is drawn wrongly. A service boundary is not a code boundary but a data ownership boundary.
Third, is there a team? Each service needs people to be on call. If a team of three operates nine services, deployments get faster but the middle-of-the-night pages triple.
The method of transition is not a big bang but the Strangler Fig. You put a facade in front, move one feature at a time to a new service, shift traffic by percentage, and once stable, delete that code from the monolith. You repeat transform → coexist → eliminate. The key to this method is that you can always go back. You only need to revert one line of routing in the facade.
What you meet in the field
The facade is usually handled by an API gateway (Kong, Envoy, NGINX). An Anti-Corruption Layer is placed along with it. It is a layer that translates so that the legacy data model does not seep into the new service. If you omit it, within 6 months the new service comes to resemble the legacy schema exactly. Then the point of splitting is lost.
The most frequently seen failure is "we split off the inventory service but still share the inventory table". In this case deployment was separated, but a schema change still requires agreement between two teams. This is because the coupling was in the data, not the code.
A list of costs to pay before splitting
Adding one service does not only add code. These are the things that actually come with it.
| Item | What increases |
|---|---|
| Repository and CI | One pipeline, build time, secret management |
| Deployment | Manifests, rollout strategy, rollback procedure |
| Observability | Dashboards, alerts, log labels, trace propagation |
| On-call | One runbook, assigning an owner |
| Communication | Network round trips, serialization, retry and timeout design |
| Data | The transaction boundary is cut — sagas or eventual consistency |
The last row is the most expensive. What used to finish within one transaction becomes compensating transactions and a state machine. The code triples and bugs hide inside it.
Signals that splitting is nevertheless right
There are cases where you must split despite the high cost. If any one of the three applies, it is worth it.
The deployment cadences differ greatly. If a part that ships ten times a day and a part that ships once a quarter are in one repository, the slow one holds back the fast one.
The resource characteristics differ. If you put GPU inference and ordinary CRUD in the same Pod, CRUD ends up on the GPU node too. If the scaling axes differ, splitting is right.
Failure isolation is needed. If report generation uses up all the memory and even payments die, the two must be separated.
Different teams alone is not enough. Teams change, and if you split the system along the org chart, you have to split the system again every time the organization changes.
Going back is design too
Merging back what you split is right more often than you think. But most teams regard it as a failure and put it off.
These are signs you should merge.
- The two services are always deployed together
- Building one feature requires changing both sides
- The call between them is synchronous and mandatory (if one dies, the other is useless)
- The data must be one transaction but you are imitating it with a saga
If three or more apply, merging is better. Going back to a modular monolith is not a retreat but a correction.
The order of splitting — what to take out first
If you have decided to split, the next question is the order. If you take out just anything first, you end up doing the hardest thing first.
Candidates to take out first satisfy three conditions. They share almost no data, their deployment cadence differs from the main body, and the main body keeps running even if they fail. Things like sending notifications, file conversion, report generation, and search indexing fit here.
What to take out later is the center of the transactions. If you split what was bound as one transaction — orders, payments, inventory — you take on a distributed transaction problem at that moment. You need compensating transactions or the saga pattern, and that is something to do after the team is ready.
Split the data first. Splitting only the services while sharing the database is the worst. Deployment is separate but the schema changes together, so you end up in a state where both sides' deployments must be timed simultaneously. A monolith is better than that.
Create the boundary inside the code first, before taking it out. Divide into modules, narrow the calls between those modules to a single interface, and wait until that interface is stable. Once it is in that state, splitting the process becomes mechanical work. If you split the process without knowing the boundary, you end up fixing that boundary across the network.
Split in a way that can be reversed. Send traffic to the new service little by little, and if there is a problem, revert to the code in the main body. Delete the code from the main body only after several weeks without problems.
Decide in advance what to measure. Deployment frequency, change failure rate, and the number of repositories touched to change one feature. If the last one is growing, the boundary is wrong, so stop and look again before splitting more.
What you will do in the next lab
You take a monolith in which orders and inventory are in one process, fix the contract first as a document, and split it into two services. After splitting, you build a script that automatically compares whether the responses of the two implementations are the same, and finally also write one line of evidence for "this should not have been split".