TT Lab
Get started
Learn Learning paths Courses

Cloud Permission Design

Open Wide and Narrow, or Open Narrow and Widen

Continue in TT Lab

In one line

Least privilege is not something you design all at once but a process of measuring and narrowing. Nobody knows in advance every API they will call.

Why it was needed

You deploy a new service. The documentation does not say exactly which permissions are needed. So you end up choosing one of two approaches.

A. Open wide, then narrow — start with s3:*, get it working, then reduce later. B. Open narrow, then widen — put in only what is needed and add one at a time whenever something is blocked.

In reality, with A, the commit that narrows it never comes. Once it starts working, nobody touches it again. So the principle is B, but B is slow during development.

How it works

A compromise used in practice

1. 개발 환경에서만  넓게 연다   (운영 계정에는 절대 넣지 않는다)
2. 실제로 호출된 API 를 기록한다  (감사 로그·액세스 어드바이저)
3. 그 목록으로 정책을 생성한다
4. 스테이징에서 좁힌 정책으로 검증한다
5. 운영에는 좁힌 정책만 배포한다

The key is the environment separation in number 1. If you decide where broad permissions are allowed to exist, "open it for now and fix it later" does not lead to an incident.

The material for number 2 is provided by the cloud — there is a feature that tells you when a principal last called which service. A permission not used even once in 90 days is a candidate for removal.

Permission boundary

If you give a developer permission to create roles, that developer can create a role stronger than themselves. The device that prevents this is the permission boundary.

실효 권한 = (붙은 정책) ∩ (권한 경계)

Because it is an intersection, a permission outside the boundary does not arise even if it is written in the policy. It is a way of expressing "create roles freely, but you cannot go beyond this range."

At the organization level, there are further guardrails above this, such as service control policies (SCPs). It is a Deny that covers the whole account, so even an administrator cannot break through it.

A few commonly used guardrails

Guardrail What it prevents
No use outside approved regions Resources appearing in unexpected regions
No use of the root account Everyday use of the top-level account
No disabling audit logs Wiping out traces
No creating unencrypted storage Default-value mistakes
No public access settings Bucket exposure incidents

Applying just these five makes most common incidents disappear. And because these are rules kept by the system rather than by people's attention, they hold over time.

What is commonly missed when narrowing

What it looks like in the field

Treat permissions for people and machines differently

When narrowing least privilege, permissions used by people and permissions used by programs differ in nature, so if you handle them the same way, both go wrong.

Machine permissions are easy to narrow. What a program does is fixed, and when it changes, a deployment happens, so recording the actual calls and granting exactly that much works well. And the principle for machines is not to give them long-term credentials. If you attach a role to an instance or workload so that it automatically receives short-lived credentials, incidents such as a key being committed to a repository or printed in a log disappear altogether. Even things that run outside the cloud, like external CI, can these days build the same structure with identity federation.

People's permissions are hard to narrow. What they do differs each time, and when an outage happens, permissions they don't normally use are suddenly needed. So for people you keep standing permissions low and create a path to raise them briefly when needed. It is a method where you get a few hours of permission after approval, and that fact is left as a record. Without this, in the end everyone ends up holding high permissions all the time "because it would be a problem during an outage."

In both cases, the shorter the lifetime of a credential, the lower the value of a leak. So reducing lifetime is as important as narrowing, and actually this side is easier. Refining policies line by line takes weeks, but converting a long-term key to a role finishes in one go and the effect is immediate.

Finally, we point out the procedure for cleaning up people who leave and services that disappear. Every organization has a procedure for granting permissions, but often no procedure for revoking them. When unused accounts and roles pile up, they themselves become an attack surface, and because you can't tell what is alive, you can't even clean them up. Periodically pulling out principals that have not been used for 90 days is the cheapest way to prevent this problem.

What to look at next

We look, through cases, at what incidents actually happen when all of this breaks down.