Annotation Usage Scenarios
Find your requirement, copy the annotations. One minimal Pod template per common vGPU scenario.
Each section below answers one requirement with the smallest set of annotations that meets it. Find yours and copy the snippet into the Pod template of your workload.
Before you start
- TensorFusion is installed with a running
GPUPool, and you have deployed one workload on it. If not, follow Create Workload first. - Every snippet replaces the
template.metadatablock of this Deployment. Labels and annotations go on the Pod template, not on the Deployment's own metadata.
apiVersion: apps/v1
kind: Deployment
metadata:
name: demo
spec:
replicas: 1
selector:
matchLabels:
app: demo
template:
metadata:
labels:
app: demo
tensor-fusion.ai/enabled: "true"
annotations: {} # the snippets below go here
spec:
containers:
- name: app
image: pytorch/pytorch:2.4.1-cuda12.1-cudnn9-runtime
command: ["sh", "-c", "sleep infinity"]- Several snippets need the exact GPU model or pool name. Look them up with:
kubectl get gpupool
kubectl get gpu -o custom-columns='NAME:.metadata.name,VENDOR:.status.vendor,MODEL:.status.gpuModel,CAPACITY:.status.capacity,AVAILABLE:.status.available'| I want to... | Go to |
|---|---|
| Give a Pod part of a GPU | Share one GPU among several Pods |
| Not deal with TFLOPS | State compute as a percentage |
| Give a Pod a full GPU | Use a whole GPU |
| Give a Pod more than one GPU | Use several GPUs |
| Control where the Pod lands | Pick the pool, model, vendor or card |
| Run the Pod on a node without a GPU | Use a GPU on another node |
| Isolate tenants more strictly | Choose an isolation mode |
| Protect important workloads | Set the priority |
| Use a Pod with sidecars | Choose which containers get the GPU |
| Stop guessing the size | Let TensorFusion adjust the size |
| Start distributed training atomically | Start a group of Pods together |
| Reuse one configuration | Share settings across workloads |
| Adopt TensorFusion gradually | Enable only some replicas |
| Exclude a Pod | Keep a Pod out of TensorFusion |
Share one GPU among several Pods
Use this for inference services that each need a fraction of a card. Give every Pod a TFLOPS and VRAM request, which the scheduler reserves, and a limit, which the Pod cannot exceed.
annotations:
tensor-fusion.ai/tflops-request: "10"
tensor-fusion.ai/tflops-limit: "20"
tensor-fusion.ai/vram-request: "4Gi"
tensor-fusion.ai/vram-limit: "4Gi"Several Pods can be placed on one GPU as long as the sum of their requests fits; which card is picked depends on the pool's placement mode. Inside the container, nvidia-smi reports about 4 GiB of total memory.
State compute as a percentage
If the pool has a single GPU model, a percentage of one card is easier to reason about than TFLOPS.
annotations:
tensor-fusion.ai/compute-percent-request: "20"
tensor-fusion.ai/compute-percent-limit: "50"
tensor-fusion.ai/vram-request: "4Gi"
tensor-fusion.ai/vram-limit: "4Gi"Do not combine compute-percent-* with tflops-*. Percentages are not counted against namespace GPU quotas, so use TFLOPS in multi-tenant clusters.
Use a whole GPU
Use this for training jobs or any workload that should not share its card. isolation: shared allocates each GPU in full and injects no limiter; dedicated-gpu fills in the full capacity of the model for you.
annotations:
tensor-fusion.ai/isolation: "shared"
tensor-fusion.ai/dedicated-gpu: "true"
tensor-fusion.ai/gpu-model: "<gpu-model>" # exact value of status.gpuModel
tensor-fusion.ai/gpu-count: "1"After scheduling, the assigned GPU shows zero available TFLOPS and VRAM, and no other Pod is placed on it.
Use several GPUs
gpu-count sets the number of GPUs for each Pod. All of them come from the same node. The request and limit apply to each GPU, so this Pod reserves 2 × 8Gi.
annotations:
tensor-fusion.ai/gpu-count: "2"
tensor-fusion.ai/tflops-request: "30"
tensor-fusion.ai/tflops-limit: "60"
tensor-fusion.ai/vram-request: "8Gi"
tensor-fusion.ai/vram-limit: "8Gi"Inside the container, torch.cuda.device_count() returns 2. The allowed range is 1 to 128.
Pick the pool, model, vendor or card
Each of these annotations narrows the set of GPUs the scheduler may use. Combine them as needed. If nothing matches, the Pod stays Pending.
annotations:
tensor-fusion.ai/gpupool: "<gpupool-name>" # default: the cluster's default pool
tensor-fusion.ai/gpu-model: "<gpu-model>" # exact value of status.gpuModel
tensor-fusion.ai/vendor: "NVIDIA" # exact value of status.vendor
tensor-fusion.ai/tflops-request: "10"
tensor-fusion.ai/tflops-limit: "20"
tensor-fusion.ai/vram-request: "4Gi"
tensor-fusion.ai/vram-limit: "4Gi"To pin the Pod to specific cards on a node, list their device indices. The GPU count then equals the number of indices.
annotations:
tensor-fusion.ai/gpu-indices: "0,1"
tensor-fusion.ai/tflops-request: "10"
tensor-fusion.ai/tflops-limit: "20"
tensor-fusion.ai/vram-request: "4Gi"
tensor-fusion.ai/vram-limit: "4Gi"Node-level constraints work as usual: the TensorFusion scheduler honors nodeSelector, node affinity, taints and tolerations in the Pod spec.
Use a GPU on another node
By default a Pod runs in local GPU mode: it is scheduled onto a GPU node and uses the card directly. In remote vGPU mode the Pod can run on any node, including one without a GPU, and its GPU calls are forwarded over the network to a worker Pod on a GPU node.
annotations:
tensor-fusion.ai/is-local-gpu: "false"
tensor-fusion.ai/tflops-request: "10"
tensor-fusion.ai/tflops-limit: "20"
tensor-fusion.ai/vram-request: "4Gi"
tensor-fusion.ai/vram-limit: "4Gi"Check the result:
kubectl get tensorfusionconnection
kubectl get pods -A -l tensor-fusion.ai/component=worker,tensor-fusion.ai/workload=demo -o wideA TensorFusionConnection exists for the Pod, and a worker Pod runs on a GPU node. The tensor-fusion.ai/gpu-ids annotation is on the worker Pod, not on your Pod.
Choose local mode when latency matters, and remote mode when the application must stay on CPU nodes or GPU nodes are scarce. The default for a pool is its defaultUsingLocalGPU setting.
Choose an isolation mode
tensor-fusion.ai/isolation decides how a Pod is separated from others on the same GPU. The default is soft.
| Mode | Choose it when |
|---|---|
soft | You trust the workloads and want the most flexible sharing. Limits can be changed while the Pod runs, so autoscaling works. |
hard | You need stricter performance isolation between tenants. Limits are fixed when the Pod starts. |
shared | The Pod should own whole GPUs. See Use a whole GPU. |
partitioned | You need hardware-level partitions such as NVIDIA MIG, and the administrator has set the cluster up for it. |
Hard isolation:
annotations:
tensor-fusion.ai/isolation: "hard"
tensor-fusion.ai/tflops-request: "20"
tensor-fusion.ai/tflops-limit: "20"
tensor-fusion.ai/vram-request: "4Gi"
tensor-fusion.ai/vram-limit: "4Gi"In local mode, a hard-isolated Pod gets an extra tensorfusion-worker container. With the default installation, a GPU serves one isolation mode at a time, so soft and hard Pods land on different cards.
Partitioned isolation with a named partition template:
annotations:
tensor-fusion.ai/isolation: "partitioned"
tensor-fusion.ai/partition-id: "<partition-template-id>"
tensor-fusion.ai/gpu-model: "<gpu-model>"
tensor-fusion.ai/vendor: "<vendor>"Partitioned mode does not work with the default isolationModePolicy: dynamic. Ask your administrator whether the cluster supports it and which template IDs exist.
Set the priority
annotations:
tensor-fusion.ai/qos: "critical" # low | medium | high | critical
tensor-fusion.ai/tflops-request: "20"
tensor-fusion.ai/tflops-limit: "20"
tensor-fusion.ai/vram-request: "8Gi"
tensor-fusion.ai/vram-limit: "8Gi"high and critical Pods without their own priorityClassName get the tensor-fusion-high or tensor-fusion-critical PriorityClass and can preempt lower-priority Pods when GPUs are full. See Create Workload for how to choose a level and what happens when you leave it out.
Choose which containers get the GPU
A Pod with more than one container must say which containers use the GPU. The others are left untouched and see no GPU.
annotations:
tensor-fusion.ai/inject-container: "app" # comma-separated for several
tensor-fusion.ai/tflops-request: "10"
tensor-fusion.ai/tflops-limit: "20"
tensor-fusion.ai/vram-request: "4Gi"
tensor-fusion.ai/vram-limit: "4Gi"When several containers use GPUs and each should get its own cards, add the count for each container. The sum must not exceed gpu-count.
annotations:
tensor-fusion.ai/inject-container: "trainer,eval"
tensor-fusion.ai/gpu-count: "3"
tensor-fusion.ai/container-gpu-count: '{"trainer":2,"eval":1}'
tensor-fusion.ai/tflops-request: "30"
tensor-fusion.ai/tflops-limit: "60"
tensor-fusion.ai/vram-request: "8Gi"
tensor-fusion.ai/vram-limit: "8Gi"Without container-gpu-count, every listed container sees all GPUs of the Pod.
Let TensorFusion adjust the size
Start from an estimate and let the autoscaler correct the requests and limits from actual usage.
annotations:
tensor-fusion.ai/autoscale: "true"
tensor-fusion.ai/autoscale-target: "all" # compute | vram | all
tensor-fusion.ai/tflops-request: "10"
tensor-fusion.ai/tflops-limit: "20"
tensor-fusion.ai/vram-request: "4Gi"
tensor-fusion.ai/vram-limit: "4Gi"Nothing is adjusted during the first 30 minutes after the workload is created. For schedules, percentiles and the status fields to watch, see Configure AutoScaling.
Start a group of Pods together
For distributed training, starting three of four workers wastes GPUs. With gang scheduling, no Pod of the group is bound until enough of them can be placed.
spec:
replicas: 4
template:
metadata:
labels:
app: demo
tensor-fusion.ai/enabled: "true"
annotations:
tensor-fusion.ai/gang-enabled: "true"
tensor-fusion.ai/gang-min-members: "4" # optional, defaults to all replicas
tensor-fusion.ai/gang-timeout: "5m" # optional, defaults to waiting indefinitely
tensor-fusion.ai/tflops-request: "100"
tensor-fusion.ai/tflops-limit: "100"
tensor-fusion.ai/vram-request: "10Gi"
tensor-fusion.ai/vram-limit: "10Gi"The workload needs at least 2 replicas, and gang-min-members must be between 2 and the replica count. Progress is reported in status.gang of the TensorFusionWorkload.
Share settings across workloads
Put the common settings in a WorkloadProfile in the same namespace and reference it. Annotations on the Pod override the profile.
apiVersion: tensor-fusion.ai/v1
kind: WorkloadProfile
metadata:
name: small-inference
spec:
resources:
requests:
tflops: "10"
vram: 4Gi
limits:
tflops: "20"
vram: 4Gi
qos: medium annotations:
tensor-fusion.ai/workload-profile: "small-inference"
tensor-fusion.ai/vram-limit: "6Gi" # overrides the profileEnable only some replicas
To try TensorFusion on part of a Deployment, cap the number of Pods it takes over. The remaining Pods are created unchanged.
annotations:
tensor-fusion.ai/enabled-replicas: "1"
tensor-fusion.ai/tflops-request: "10"
tensor-fusion.ai/tflops-limit: "20"
tensor-fusion.ai/vram-request: "4Gi"
tensor-fusion.ai/vram-limit: "4Gi"For the full procedure, see Migrate Existing Workload.
Keep a Pod out of TensorFusion
A Pod without the label tensor-fusion.ai/enabled: "true" is not touched. If the administrator turned on auto migration, Pods that request nvidia.com/gpu are taken over even without the label; set the label to "false" to exclude one.
labels:
app: demo
tensor-fusion.ai/enabled: "false"Next steps
- Annotation reference: every annotation, with type, default and allowed values.
- Best practices: how to choose between the options on this page.
- Configure AutoScaling
TensorFusion Docs