TT Lab
Get started
Learn Learning paths Courses

GPU Operator and Time-Slicing

Time-Slicing, MPS and MIG

Continue in TT Lab

In one line

The replicas: 5 of time-slicing does not mean to cut the GPU into five pieces; it means to register the same device five times. Only the slots the scheduler sees become five times as many, and the VRAM is still one undivided block.

Why you come to want to share

As soon as you buy a GPU you see a graph like this. Utilization is 8% on daily average. Yet the queue is always full. The reason is simple — one Pod grabs a whole GPU. Someone who keeps a Jupyter notebook open and reads code, and a service that receives one inference request per second, are in fact barely using the GPU while holding a whole card.

The demand that comes out of this is "let several people use one card", and there are three answers. The three are entirely different in nature, but their names are all similar, so they are often mixed up.

How it works

Method What it divides Memory isolation Fault isolation Requirement
Time-slicing Only execution time (context switching) None None Any GPU
MPS Time + concurrent kernel execution A cap can be set, but it is not enforced isolation Limited A daemon is needed
MIG A hardware partition Yes Yes A100/H100 class

Time-slicing is turned on with a single device plugin configuration.

sharing:
  timeSlicing:
    failRequestsGreaterThanOne: true
    resources:
      - name: nvidia.com/gpu
        replicas: 5

With this, the nvidia.com/gpu the node advertises becomes the number of physical devices × 5. With 4 cards it is 20. The scheduler believes there are 20 slots and sends 20 Pods. But what those 20 actually meet is still 4 GPUs, and the five Pods attached to one card overlap in using the same 40GiB of VRAM. The memory guaranteed to a slot is 0.

So this kind of thing happens. If one of the five loads a large model and takes 38GiB, the other four hit CUDA OOM while their Pods are Running. The scheduler is not at fault at all — it said it would give slots and it gave slots. It just never promised memory.

failRequestsGreaterThanOne: true is a switch that tries to reduce this misunderstanding even a little. If a Pod requests nvidia.com/gpu: 2, it is rejected. That is because getting two slices does not mean "the amount of two GPUs". It is better to leave it on.

MPS gathers the CUDA kernels of several processes into one context and really runs them concurrently. Context switching cost disappears so throughput rises, and you can also set a per-process memory cap. However, the cap is closer to a cooperative limit, and when one process dies, other processes can be affected.

Only MIG is a division at the hardware level. It physically separates the SMs and memory controllers, so each instance has its own VRAM. Whatever happens next to it, my instance is safe. In exchange, it needs supported hardware, the profiles are fixed so you cannot divide freely, and to change the profile you must empty that GPU.

What it looks like in the field

First, there is one criterion. "Is it acceptable for these workloads to kill one another?" For development notebooks, internal experiments, and demos, time-slicing is enough and the effect is large. If production inference services are mixed in, sharing without isolation will inevitably become an accident someday. Even in the same cluster, it is usual to split node pools and apply a different policy per node.

Second, do not be fooled by the utilization graph. If you turn on time-slicing, the GPU utilization metric rises nicely. But that number means "someone is running a kernel", not "work finishes faster". If five take turns, each one's work gets slower. The metrics to look at are not utilization but job completion time and the number of OOMs.

Third, write down what you promised users. The moment you turn on sharing, the meaning of the request nvidia.com/gpu: 1 changes. Until yesterday it was "one GPU", and from today it is "one slot whose turn will come". If you do not announce this change, users suspect the infrastructure without ever knowing why their jobs got slower.

What to look at in the next reading

Such sharing settings are in the end delivered to nodes by a DaemonSet rollout. In the next article, you see what shape the cluster is left in when that rollout stops on one node, and why that state is not a bug but behavior as designed.