Where "The Cloud Takes Care of It" Ends
In one line
The cloud provider is responsible for the security of the infrastructure (security of the cloud), and we are responsible for the security of what we put on top of it (security in the cloud). Most incidents happen where this boundary is misunderstood.
Why it was needed
If you open the "customer data leaked from AWS" incidents that appear in the news, almost all of them are an S3 bucket left open to the public. AWS was not breached. The party that made the configuration that way was the customer.
Conversely, a hypervisor vulnerability or a physical break-in at a data center is in an area we cannot touch, and that is the provider's responsibility.
If you don't know the boundary, you go wrong in two directions — we duplicate work the provider would do for us (waste), or we trust the provider to do what is our job (incident).
How it works
The boundary moves with the service model
IaaS(EC2) PaaS(RDS) SaaS(Workspaces)
데이터 고객 고객 고객 ← 항상 고객
접근 권한 고객 고객 고객 ← 항상 고객
애플리케이션 고객 고객 사업자
런타임 고객 사업자 사업자
OS 패치 고객 사업자 사업자
하이퍼바이저 사업자 사업자 사업자
물리 사업자 사업자 사업자
The top two lines — data and access permissions — are the customer's part in any model. This is why "it's managed, so it's safe" does not hold. Even if you use RDS, if you leave the password as admin/admin, that is our incident.
Five commonly mistaken beliefs
| Belief | Reality |
|---|---|
| "It's managed, so backups happen on their own" | Automatic backups usually have to be turned on, and we set the retention period too |
| "Encryption is the default" | Encryption at rest is often optional, and the app has to enforce encryption in transit |
| "The provider patches" | OS patches on IaaS are the customer's part. Even with managed services, we choose the maintenance window |
| "If I delete it, it's gone" | It remains in snapshots, replicas and logs. Backup retention is itself delayed deletion |
| "99.99% availability is guaranteed" | An SLA is a refund criterion, not a promise of no outages |
The last row is especially important. An SLA of 99.99% is a contract that says "about 52 minutes a year is within the normal range, and beyond that, part of the charge is returned as credit." Whether our service can withstand those 52 minutes is something we have to design.
So what we must do
- Access permissions — who can do what. This is the topic of the next course.
- Data classification and encryption — decide what is sensitive, and encrypt that first.
- Backups and recovery rehearsals — turning backups on is different from being able to recover.
- Configuration monitoring — automatically find public buckets and open security groups.
- Log retention — the material to investigate when an incident happens. The provider will not keep it for you.
The boundary differs by service type
"Shared responsibility" is one line, but the actual boundary differs for each service.
| What I do | What the provider does | |
|---|---|---|
| IaaS (EC2) | OS patches, firewall, application, data | Hypervisor, physical security, network |
| PaaS (RDS) | Schema, accounts, parameters, backup policy | OS and engine patches, hardware |
| SaaS (S3) | Permissions, encryption choices, lifecycle | Everything else |
| Serverless (Lambda) | Code, dependencies, permissions | Runtime, scaling, patching |
The higher you go, the less is my part, but data and permissions are my part at every layer. These two make up most real incidents.
The ranking of incidents that actually happen
| Rank | Cause | Whose responsibility |
|---|---|---|
| 1 | Wrong permission settings (public buckets, excessive IAM) | Mine |
| 2 | Leaked credentials (git commits, logs) | Mine |
| 3 | Unpatched application dependencies | Mine |
| 4 | No backups, or recovery not verified | Mine |
| 5 | Provider outage | The provider's (but I still feel the impact) |
Data rarely leaks because of a defect on the cloud provider's side. Most of the time it is configuration. That is why the sentence "it's the cloud, so it's safe" is dangerous — what becomes safe is only some of the work I was not doing, and the new work that appears (permission design, credential management) is larger.
You must prepare for number 5 too
You cannot sit back because it is the provider's responsibility. When an outage happens, users complain to us.
- Multi-AZ is the default — a single-AZ setup dies along with that AZ.
- Region outages only as a plan — multi-region is expensive. Decide your RTO and pick a level that matches it. "Recover in another region within a few hours" is also a valid answer.
- Multiply the SLAs of the managed services you depend on — if you use a DB at 99.95%, a queue at 99.9% and storage at 99.99% in series, the total is 99.84%. That is 70 minutes a month.
The last line affects design. To raise availability, reducing serial dependencies is more effective than making each element better.
What it looks like in the field
- An audit asked "is it encrypted" and nobody knew the answer → there was no classification.
- After an outage, compensation was demanded on the basis of the SLA → a few dollars of credit was all there was.
- A managed DB was used but recovery failed → automatic backups were turned off.
What to look at next
We look at how the service models (IaaS/PaaS/SaaS) move this boundary in concrete terms.