TT Lab
Get started
Learn Learning paths Courses

Cloud Permission Design

Access Keys Leak Because They Exist

Continue in TT Lab

In one line

It is best not to create long-term access keys at all. Temporary credentials obtained by assuming a role expire on their own after a few hours, so even if they leak, the damage is cut off by time.

Why it was needed

Incidents where an AWS key is committed to a public GitHub repository still happen every day. A bot finds it within minutes and launches cryptocurrency-mining instances. The bill comes in the tens of millions of won.

The real problem here is not the commit mistake but that the key existed. A long-term key has these properties.

How it works

What it means to assume a role

A role is a bundle of permissions, and the trust policy decides "who can assume this role."

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": { "Service": "ec2.amazonaws.com" },
    "Action": "sts:AssumeRole"
  }]
}

This role can be assumed by an EC2 instance. If you attach it to an instance, the programs running inside obtain credentials without a key. They are fetched automatically from the metadata service and refreshed automatically before they expire.

In summary:

Long-term access key Role (temporary credentials)
Expiry None Usually 1–12 hours
Distribution Copied as files or environment variables Injected automatically
Revocation Find and revoke manually Just wait for expiry
Tracking Unclear who uses it Audit logs per role

Service account keys are the same problem

It is the same in Kubernetes. Putting a cloud key into a Pod as a Secret carries exactly the same risk. Instead, use OIDC federation — assume a cloud role with a service account token issued by the cluster (AWS's IRSA, GCP's Workload Identity, Azure's Workload Identity). The key point is that the key no longer exists at all.

CI is the same. If instead of putting long-term keys into GitHub Actions or a Gitea runner you have it assume a role through OIDC, the cloud key disappears from the repository secrets.

When you still need a long-term key

There are cases where you cannot eliminate it completely (external SaaS integrations and so on). Then do at least this.

  1. Create a dedicated principal — do not share the key of a person's account.
  2. Narrow the permissions to that purpose only — down to IP and time with a Condition.
  3. Make rotation a procedure. Put it in the calendar.
  4. Look at the usage trace — if the last-used time is old, delete it.

Number 4 is the cheapest and most effective. Just deleting unused keys greatly reduces the attack surface.

Sessions and expiry time

When you assume a role, you decide the session duration. The shorter, the safer, but if it is too short, it expires in the middle of a long job. The practical guidelines are these.

What it looks like in the field

What to watch out for when assuming roles

It is true that roles are safer than long-term keys, but if you write the trust policy loosely, that safety disappears entirely. The places where incidents happen are mostly fixed.

A role left open for anyone to assume. If you write the principal of the trust policy broadly, someone in another account can assume that role. Roles opened to give permissions to an external partner are the typical case, and the other account's number alone is not enough. This is because anyone inside that account can assume it. When opening to the outside, you must also attach an identifier known only to both sides as a condition, so that a third party cannot get around it through that account.

Leaving out conditions in federation. When CI assumes a role through OIDC, if you do not attach a condition on which repository and which branch, any repository that uses that CI service can assume our role. This is the most common mistake, and one line of condition blocks it.

The permission to pass a role. When creating a resource, the action of attaching a role to it is effectively a permission to "make something do what I cannot do." When you grant this action, you must narrow it down to which roles can be attached, otherwise a low-privilege user can create a resource with an administrator role attached and escalate privileges.

Look at the end of the chain. When a structure is created in which a role assumes another role and that assumes yet another, it becomes hard for a person to trace what can ultimately be done. The evaluation tool mentioned in the previous module is especially valuable here, and limiting the depth of the chain by rule is also a method.

Finally, temporary credentials are not cancelled while they are valid. Even if you reduce the role's permissions, sessions that have already been issued stay alive. So keeping the session duration short is not mere hygiene but the response speed itself when an incident occurs.

What to look at next

The order for actually narrowing permissions. It is impossible to hit least privilege from the start.