TT Lab
Get started
Learn Learning paths Courses

Cloud Fundamentals

Why Lift and Shift Fails

Continue in TT Lab

In one line

If you move things as they are (lift and shift), it usually becomes more expensive. On-premises is a structure where you use up the resources you bought, and the cloud is a structure where you pay for what you use, so the design premises are opposite.

Why it was needed

The case "we moved 30 servers to the cloud as they were and the cost doubled" is very common. Looking at each part, it goes like this.

In other words, before moving you have to measure the size again (right-sizing), turn things off when they are not used, and redraw the traffic paths to get value. If you don't do that work, the cloud is just an expensive data center.

Criteria for excluding things from migration

Things better not moved

Target Reason
Large batches that run at full load 24 hours No benefit from pay-as-you-go. Buying is cheaper
Systems that need ultra-low latency Sometimes they have to sit physically next to something
Dependence on special hardware License dongles, specific cards, industrial interfaces
Data that may not be transferred abroad Impossible from the start if there is no region
Stable systems already fully depreciated Moving something that works also costs money
Workloads that frequently export large volumes Egress charges eat up the break-even

The last row is often overlooked. Incoming traffic is usually free, but outgoing traffic (egress) is expensive. For a service that keeps sending large files out, most of the cloud bill comes from here.

When hybrid is the answer

There is no need to see it as all or nothing.

If you decide to move, the order

  1. Inventory — draw what depends on what. Without this, there is nothing after it.
  2. Classification — what to move as is / rebuild / replace with managed / discard. Finding what to discard is actually the most valuable. There are always servers nobody uses.
  3. Re-sizing — based on actual usage metrics. Do not use the peak-based specs as they are.
  4. Start small — start with things that can be reversed and refine the procedure.
  5. Cost observation — look at it by tag from the first month after moving. If you add it later, you can't.

What it looks like in the field

Cost is decided in design

If you try to cut the bill after migration, there is usually not much you can do. Most cloud cost is already decided at the point where you decide what to use and how. So we write down separately the things to pin down at the design stage.

Commit in advance for resources whose lifetime you know. If you commit to 1 or 3 years, you use the same instance considerably more cheaply. Conversely, cheap instances that can be reclaimed at any time are used only for work that can tolerate interruption. Batches and test environments fit here, and if you use them for a service that holds state, an incident occurs every time one is reclaimed.

Split storage by access frequency. If you put rarely read data such as logs and backups in the same tier as frequently read data, the charge becomes several times larger. But cheap tiers charge you when you retrieve, or take time, so if you put data you would need for recovery in too cold a tier, you are in trouble just when you need it. It is safe to use lifecycle rules to move down only the old data automatically.

Draw where traffic flows. The charges differ for traffic within the same region, between availability zones, out of the region, and out to the internet. Placing the database in a different availability zone is necessary for availability, but it carries charges accordingly, so choosing it knowingly is different from it just happening.

Look at managed services at a price that includes operating costs. A self-operated database looks cheap on the price list, but people have to do backups, redundancy, version updates and night-time response. If you don't convert that time into cost and compare, self-operating always looks cheaper. The smaller the team, the heavier this item weighs.

Finally, enforce tags from the start. Resources without a tag saying which team and which service they belong to become untouchable after just a few months. You can't tell whether they are safe to delete, so you leave them, and the resources left that way come to take up a considerable part of the bill.

Next course

Permissions. This is where incidents happen most often in the cloud, and it is an area you can learn by writing policy documents directly, even without an account.