Sharing Configuration — Splitting It Up with Slices, MIG and Quotas
Goal
You handle the three methods of making several users share one card through configuration and objects. You confirm in the node status that slots grow by the number of slices, confirm through scheduling that MIG nodes have different resource names, and go as far as pinning the share handed to a team with a quota.
Why it matters
After an accident, GPU utilization is a single digit on daily average yet the queue is always full. That is because one Pod grabs a whole card. The demand "let several users share one card" comes out of this, and there are three answers whose similar names are often mixed up.
The most important distinction is what is being divided. The replicas of time-slicing does not mean to cut the GPU into pieces but to register the same device several times. Slots increase but the VRAM is one undivided block, so if one Pod takes a lot, the rest suffer memory shortage. MPS raises throughput by running kernels concurrently but its isolation is still weak. Only MIG divides in hardware, so both memory and faults are isolated. And on MIG nodes the resource name itself changes to the profile name, so it can happen that an old manifest can never use that node.
Environment
This Pod has neither a GPU nor a device plugin. Instead, there is a real control plane brought up by kwok, so if you write a number in the node status, the real scheduler sees it and judges. Quotas are also checked by the real apiserver. The working directory is /root/gpushare, and under it you use k8s/, bin/, and out/.
Steps
- Create the
gpu-shareandteam-anamespaces and a ConfigMap holding two sets of configuration. - Choose the configuration by node label and write the advertised amounts for each.
- Make
lab-node-2a MIG mixed node. - Bring up two Pods that differ only in resource name and see them split.
- Put a quota on
team-aand capture inout/quota-denied.txtthe rejection of the request that overflows. - Screen requests for several slices in advance with
bin/check-request.sh. - Add
mpsto the ConfigMap and write an isolation table inout/isolation.txt. - Place the three workloads on the nodes that suit each and write it in
out/placement.txt.
Notes
- Do not guess the advertised amounts in step 2; multiply by the
replicasof step 1. The grader reads the ConfigMap and does the same multiplication again. - You must write
capacityandallocatabletogether. If you write only one, no slots appear for the scheduler to use. - In step 3, do not also write
nvidia.com/gpuon the MIG node. The point of the mixed strategy is that that node advertises only by profile name. - For the node names in step 8, check with
kubectl get pod -o wideand write them.
Hold two sets of device plugin configuration under names
Create the gpu-share and team-a namespaces, and with /root/gpushare/k8s/device-plugin-configs.yaml create a ConfigMap device-plugin-configs in gpu-share. It has two keys, default and shared; shared has, under sharing.timeSlicing, failRequestsGreaterThanOne: true and resources[0] as nvidia.com/gpu and replicas: 4. Do not put sharing in default.
A device plugin can hold several sets of configuration in one ConfigMap and have each node choose which to use. That way, within the same cluster, some nodes can be used whole and some shared. failRequestsGreaterThanOne is a switch that rejects a request for two or more slices. That is because even if you get two slices, it only reserves the same device twice, and neither performance nor memory doubles. The value in a ConfigMap is a multi-line string, so use |-.
Choose which configuration to use by node label
Attach nvidia.com/device-plugin.config=default to lab-node-0 and =shared to lab-node-1. Then write in the status the nvidia.com/gpu the two nodes advertise. Assume both nodes have 2 physical GPUs; lab-node-0 is 2, and lab-node-1 is 2 multiplied by the number of slices.
The device plugin looks at the node's nvidia.com/device-plugin.config label and chooses which key of the ConfigMap to use. So you can have a different sharing policy per node within the same cluster. Do not guess the advertised amount; multiply by the replicas of step 1 as it is. You must write it in both capacity and allocatable. What grows here is only the number of slots, and the VRAM is still one block.
On a MIG mixed node, the resource name itself is different
Attach the labels nvidia.com/mig.strategy=mixed and nvidia.com/gpu.product to lab-node-2, and advertise nvidia.com/mig-1g.10gb as 7 in the status. Do not write nvidia.com/gpu on this node.
MIG divides the card in hardware. In the mixed strategy, the device plugin advertises resources under a different name per profile, so the resource name a Pod requires is also a profile name such as nvidia.com/mig-1g.10gb rather than nvidia.com/gpu. That is why this node does not advertise nvidia.com/gpu at all. If you cut an A100 40GB into 1g.10gb, seven pieces come out.
If the resource names differ, they cannot go to each other's slots
Write two Pods in /root/gpushare/k8s/mig-job.yaml and apply them to gpu-share. mig-job requires 1 of nvidia.com/mig-1g.10gb and plain-gpu requires 1 of nvidia.com/gpu. Do not give either a node condition.
Even without conditions, they split by resource name alone. mig-job can go only to a node that advertises that name, and plain-gpu cannot go to the MIG node. Even if they use the same physical card, they are entirely different resources to the scheduler. So if you leave old manifests as they are when introducing MIG nodes, Pods can never use that node.
Pin the share handed to a team with a quota
With /root/gpushare/k8s/gpu-quota.yaml, create a ResourceQuota gpu-quota in team-a. Limit requests.nvidia.com/gpu to 2. Then bring up in team-a two Pods that each use one GPU to fill the limit, apply /root/gpushare/k8s/team-over.yaml, which requires one more, and save the rejection message to /root/gpushare/out/quota-denied.txt.
Extended resources too can be limited by a namespace quota. The only difference is that the key name takes the form requests.<자원이름> (the placeholder is the resource name). If you place the two Pods on the slice node with the nodeSelector nvidia.com/device-plugin.config: shared, there are plenty of slots. A request after the limit is filled is blocked not by the scheduler but at the admission stage, so the Pod is not created at all. The message goes to standard error, so capture it with 2>&1.
Screen requests for several slices before deployment
Create /root/gpushare/bin/check-request.sh <파드매니페스트> (the placeholder is the Pod manifest). If a container requires more than 1 of nvidia.com/gpu, it prints a line OVER=<컨테이너이름>=<개수> (the placeholders are the container name and the count) for each and ends with exit code 1. Otherwise it prints OK and ends with 0. A Pod that does not require a GPU at all must also pass.
failRequestsGreaterThanOne is a judgment the device plugin makes at runtime. If you make the same judgment before deployment, you can keep users from making plans with a misunderstanding of the resource. Some files contain several documents, so read with yaml.safe_load_all and look at both limits and requests. The grader actually runs this script with three kinds of manifest.
Add the third method and make the isolation table
Add an mps key to the ConfigMap device-plugin-configs. sharing.mps.resources[0] is nvidia.com/gpu with replicas: 4, and you do not put timeSlicing with it. Then write the isolation table in /root/gpushare/out/isolation.txt as five lines. They are TIMESLICING_MEMORY_ISOLATION, MPS_MEMORY_ISOLATION, MIG_MEMORY_ISOLATION, TIMESLICING_FAULT_ISOLATION, and MIG_FAULT_ISOLATION.
MPS gathers the kernels of several processes into one context and really runs them concurrently. You can also set a per-process memory cap, but that is closer to a cooperative limit, and when one process dies, other processes can be affected. The values of the five lines are one of yes, no, and partial. If the value is complete isolation, it is yes; if there is none, no; and if a cap can be set but it is not enforced isolation, partial.
Place three workloads of different natures on the nodes that suit each
With /root/gpushare/k8s/placement.yaml, create three Pods in gpu-share. prod-inference uses a profile resource on the MIG node, notebook uses nvidia.com/gpu on the slice node, and batch-train uses nvidia.com/gpu on the undivided node, one each. Choose the nodes by label. Then write where they actually went in /root/gpushare/out/placement.txt as three lines of <파드이름>=<노드이름> (the placeholders are the Pod name and the node name).
The baseline is this: MIG if you need isolation, time-slicing if utilization is the goal. Production inference must not be caught up in a neighboring Pod's memory accident, so send it to the hardware partition side; an experimental notebook has low utilization, so send it to the slice node. Batch training is faster using the whole card, so send it to the undivided node. For the node conditions, just use the labels you attached in steps 2 and 3. Do not guess the node names in the file; check with kubectl get pod -o wide and write them.