Four Incident Paths and the Line That Stops Each
In one line
Cloud security incidents usually come in not through a sophisticated attack but through a door we left open. There are only a few fixed paths, and each has a fixed line of defense that corresponds to it.
This content is for defense. For each path, we pair "what went wrong and what stops it."
Why it was needed
Cloud incidents look different from product to product, but they converge on a few paths: public settings, leaked credentials, excessive permissions and abuse of metadata. You must know the first line of defense for each path and the order for preserving evidence after an incident to keep the same configuration mistake from spreading into a compromise of the whole organization.
How it works
Path 1 — Public storage
This is the most common. A bucket is left open to the public and forgotten.
How it happens — it was opened to share a file temporarily and not reverted. Or, while setting up static website hosting, everything was switched to public.
Lines of defense
- Turn on public access blocking at the account level (it overrides individual bucket settings)
- Deny the public setting itself with an organization guardrail
- Automatically detect public buckets with configuration monitoring and alert
Path 2 — Leaked credentials
How it happens — repository commits, CI log output, screenshots, chats. A key uploaded to a public repository is discovered within minutes.
Lines of defense
- Do not create long-term keys in the first place (roles, OIDC)
- Hook a secret scan in before commit
- Anomaly detection — a region, time of day or API call volume different from usual
- Budget alerts — mining instances show up first as a sudden jump in charges
The last item is surprisingly effective. It does not prevent the incident, but you find out within a few hours.
Path 3 — Misuse of excessive permissions
How it happens — an insider's mistake, or a compromised account using broad permissions as they are. If a principal with s3:* is compromised, deletion is possible too.
Lines of defense
- Least privilege + permission boundaries
- Stronger conditions on high-risk actions such as deletion and permission changes (requiring MFA and so on)
- Turn on versioning and deletion protection for important data
- Keep backups in a different account — this is decisive
A backup that can be deleted from the compromised account is not a backup. You must put it across the account boundary to survive a ransomware situation.
Path 4 — Abuse of the metadata service (SSRF)
How it happens — an application has a feature that sends a request to an address when you give it a URL (such as fetching an image), and an attacker enters the instance metadata address. Then that instance's temporary credentials come out as the response.
Lines of defense
- Allow only the latest version of the metadata service (the kind that requires a token)
- Validate user-supplied URLs in the application — block private ranges
- Keep the instance role's permissions to a minimum. Even if they leak, there is little they can do
The order when an incident happens
- Set the scope — which credentials and resources are involved
- Block — revoke keys, edit role trust policies, invalidate sessions
- Preserve — secure logs and snapshots first. If you delete them, the investigation is over
- Investigate — reconstruct a timeline of what was done from the audit logs
- Prevent — add guardrails so the same path does not open again
What is often missed in step 2 is temporary credentials that have already been issued. Even if you change the role's policy, sessions that have already gone out may remain valid until expiry, so you have to invalidate the sessions separately.
Five things to do in peacetime
- Turn on audit logs and keep them in a different account
- Put MFA on the root account and do not use it day to day
- Set budget alerts — a supplementary detection line for belatedly discovering a mining-type spike in usage
- Turn on public access blocking at the account level
- Keep backups in a different account
Making the work of narrowing permissions actually run
Everyone agrees on least privilege, but it does not get carried out. Because if something stops while narrowing, that person gets blamed. So we reverse the order: first measure, then narrow.
Use the permissions actually used as the standard. Each cloud has a tool that analyzes access records and pulls out "the APIs this role actually called over the past 90 days." It is more accurate than a person imagining and writing a policy, and above all, you know in advance what will break.
Don't block all at once; leave a warning first. Split the policy into two, leave the one that really blocks as it is, and make the narrowed policy only record "this request would be denied under the new policy." Watch for a few days, and when the record goes quiet, switch then.
Eliminating long-lived keys is more effective than refining policies. Most leak incidents are caused not by policies being broad but by an access key written down somewhere. If you switch people to SSO and workloads to role delegation or OIDC federation, the very key to be stolen disappears. The place where CI connects to the cloud is especially so.
For the remaining keys, enforce expiry. Do not manage by a list of those that run and those that don't; make them deactivate automatically after 90 days. A control that depends on people's diligence surely collapses in a busy quarter.
Having one permission boundary reduces the scope of mistakes. Give developers permission to create roles, but set an upper limit on the permissions that role can have. Then the path by which they widen their own permissions is blocked, while they do not have to ask others for day-to-day work. Since requests and approvals decrease, it actually gets followed.
The last is records. Send audit logs separately to a write-only store, and separate the people who can delete that store into a different account. Since the first thing done after a compromise is to delete the logs, whether the logs remain decides whether you can investigate.
What it looks like in the field
- Permission probing continued from a region different from usual → the audit-log-based anomaly alert sounded first.
- Charges also spiked belatedly → the cost alert was useful, but it was not the first detection line of the compromise.
- Revoked the key but the attack continued → sessions that had already been issued were alive.
- Backups were also deleted after ransomware → they were in the same account.
Next course
Network. If permissions decide "who," the network decides "from where."