What Goes Along When You Hand It to a Managed Service
In one line
IaaS, PaaS and SaaS are not technical categories but a scale of how much you hand over to someone else. The higher you go, the more convenient it is, and you lose control by the same amount.
Why it was needed
"Should we use RDS, or install it ourselves on EC2?" is a question that almost every project asks once. The answer depends not on "which is better" but on what you can give up.
| IaaS (directly on EC2) | PaaS (RDS) | |
|---|---|---|
| Version choice | Any version | Only versions the provider supports |
| Extensions | Install anything | Only what is on the allow list |
| OS access | Anything as root | None — there is no shell |
| Patching | We plan it | The provider does it in the maintenance window |
| Incident investigation | Free use of logs and profiling | Only the exposed metrics and logs |
| Operations staff | Needed | Almost unnecessary |
| Cost | Usually cheaper | Usually more expensive |
The item no OS access hurts the most in practice. There are real cases where teams give up on managed services because they cannot use a single PostgreSQL extension or cannot touch kernel parameters. So before deciding, it is good to write down as a list "what are we doing on the server right now."
How it works
Where do containers fit
IaaS ─ VM ─ 컨테이너(직접 운영) ─ 관리형 K8s ─ 컨테이너 서버리스 ─ PaaS ─ SaaS
(EKS/AKS/GKE) (Fargate/Cloud Run)
Managed Kubernetes hands over only the control plane. The worker nodes' OS, kubelet and CNI are often still our part, and if you don't know this, the question "it's managed, so why do we patch the nodes?" comes up.
What does serverless trade away
Function-level execution (Lambda and the like) almost eliminates the operational burden, but it comes with constraints.
- Cold start — if there have been no calls for a while, the first request is slow.
- Execution time limit — you cannot run long jobs.
- No state — you cannot accumulate anything in local disk or memory.
- Vendor lock-in — the event model and runtime differ by provider.
If traffic is uneven and jobs are short, it is overwhelmingly advantageous; with a steady load and long jobs, it is actually more expensive.
Calculating lock-in as a value
If you use a managed service, you are tied to that provider. If you try to avoid this at all costs, you can use nothing, but if you use it without knowing the price, the cost of moving later explodes. The criteria are as follows.
- Is it built on a standard — a PostgreSQL-compatible managed service is easy to move, a proprietary API is hard.
- How much data accumulates — the larger the data, the more expensive the move.
- Are there substitutes — does another provider offer a service of the same kind.
A more useful attitude than "we won't be locked in" is "we know what we get in return for being locked in."
What it looks like in the field
- Moved to a managed DB but had to go back because a needed extension was missing → a missed pre-check.
- Ran a batch on serverless and hit the execution time limit → the nature of the job did not fit.
- On managed K8s, we were the ones responding to node CVEs → we misunderstood the boundary.
Write the responsibility boundary down in sentences
What causes incidents more often than choosing a model is misunderstanding the boundary. In most cases, an item left as "it's managed, so they'll take care of it" turned out to be our part. So each time you adopt a service, it is good to write down, one line each, who owns the following items.
| Item | Who does it |
|---|---|
| Physical equipment and hypervisor | Always the provider |
| Guest OS patches | On IaaS us, on PaaS the provider, and the nodes of managed K8s are usually us |
| Application vulnerabilities | Always us |
| Access permission settings | Always us |
| Whether data is encrypted | Usually we turn it on |
| Backup retention period and recovery tests | Even if the provider makes backups, the recovery test is ours |
| Availability zone placement | We decide |
The row in the table that betrays people most often is backup. People feel safe because a managed database makes a backup every day, but whether actual recovery works from that backup, whether the retention period is as long as we need, and whether it still exists when the account itself has gone wrong are all things we have to check. Having a backup and being able to recover are different facts.
The same goes for permissions. A large share of cloud incidents is "a misconfiguration," and this area remains our part whichever model we choose. What shrinks as you go up is the operational burden, not the responsibility. In fact, the more you use managed services, the more the only means of control we have is configuration, so the weight of each of those settings grows. So when you move to a managed service, you must spend the hands you have left, because fewer are needed, on reviewing configuration and testing recovery, and if you don't set aside that time, the operations time you saved turns straight into risk.
What to look at next
Where the value of a managed service comes from — we look not at the price list but at operations time.