LogoTensorFusion Docs
LogoTensorFusion Docs
HomepageDocumentation

Getting Started

OverviewKubernetes InstallVM/Server Install(K3S)Helm Quick InstallHost/GuestVM InstallTensorFusion Architecture

Application Operations

Create WorkloadAnnotation Usage ScenariosConfigure AutoScalingMigrate Existing WorkloadBest Practices

Customize AI Infra

Production-Grade DeploymentConfig QoS and BillingBring Your Own CloudManaging License

Maintenance & Optimization

Upgrade ComponentsSetup AlertsGPU Live MigrationPreload ModelOptimize GPU Efficiency

Troubleshooting

HandbookTracing/ProfilingQuery Metrics & Logs

Reference

Comparison

Compare with NVIDIA vGPUCompare with MIG/MPSCompare with NVIDIA Run:aiCompare with HAMi

Annotation Usage Scenarios

Find your requirement, copy the annotations. One minimal Pod template per common vGPU scenario.

Each section below answers one requirement with the smallest set of annotations that meets it. Find yours and copy the snippet into the Pod template of your workload.

Before you start

  • TensorFusion is installed with a running GPUPool, and you have deployed one workload on it. If not, follow Create Workload first.
  • Every snippet replaces the template.metadata block of this Deployment. Labels and annotations go on the Pod template, not on the Deployment's own metadata.
apiVersion: apps/v1
kind: Deployment
metadata:
  name: demo
spec:
  replicas: 1
  selector:
    matchLabels:
      app: demo
  template:
    metadata:
      labels:
        app: demo
        tensor-fusion.ai/enabled: "true"
      annotations: {}   # the snippets below go here
    spec:
      containers:
        - name: app
          image: pytorch/pytorch:2.4.1-cuda12.1-cudnn9-runtime
          command: ["sh", "-c", "sleep infinity"]
  • Several snippets need the exact GPU model or pool name. Look them up with:
kubectl get gpupool
kubectl get gpu -o custom-columns='NAME:.metadata.name,VENDOR:.status.vendor,MODEL:.status.gpuModel,CAPACITY:.status.capacity,AVAILABLE:.status.available'
I want to...Go to
Give a Pod part of a GPUShare one GPU among several Pods
Not deal with TFLOPSState compute as a percentage
Give a Pod a full GPUUse a whole GPU
Give a Pod more than one GPUUse several GPUs
Control where the Pod landsPick the pool, model, vendor or card
Run the Pod on a node without a GPUUse a GPU on another node
Isolate tenants more strictlyChoose an isolation mode
Protect important workloadsSet the priority
Use a Pod with sidecarsChoose which containers get the GPU
Stop guessing the sizeLet TensorFusion adjust the size
Start distributed training atomicallyStart a group of Pods together
Reuse one configurationShare settings across workloads
Adopt TensorFusion graduallyEnable only some replicas
Exclude a PodKeep a Pod out of TensorFusion

Share one GPU among several Pods

Use this for inference services that each need a fraction of a card. Give every Pod a TFLOPS and VRAM request, which the scheduler reserves, and a limit, which the Pod cannot exceed.

      annotations:
        tensor-fusion.ai/tflops-request: "10"
        tensor-fusion.ai/tflops-limit: "20"
        tensor-fusion.ai/vram-request: "4Gi"
        tensor-fusion.ai/vram-limit: "4Gi"

Several Pods can be placed on one GPU as long as the sum of their requests fits; which card is picked depends on the pool's placement mode. Inside the container, nvidia-smi reports about 4 GiB of total memory.

State compute as a percentage

If the pool has a single GPU model, a percentage of one card is easier to reason about than TFLOPS.

      annotations:
        tensor-fusion.ai/compute-percent-request: "20"
        tensor-fusion.ai/compute-percent-limit: "50"
        tensor-fusion.ai/vram-request: "4Gi"
        tensor-fusion.ai/vram-limit: "4Gi"

Do not combine compute-percent-* with tflops-*. Percentages are not counted against namespace GPU quotas, so use TFLOPS in multi-tenant clusters.

Use a whole GPU

Use this for training jobs or any workload that should not share its card. isolation: shared allocates each GPU in full and injects no limiter; dedicated-gpu fills in the full capacity of the model for you.

      annotations:
        tensor-fusion.ai/isolation: "shared"
        tensor-fusion.ai/dedicated-gpu: "true"
        tensor-fusion.ai/gpu-model: "<gpu-model>"   # exact value of status.gpuModel
        tensor-fusion.ai/gpu-count: "1"

After scheduling, the assigned GPU shows zero available TFLOPS and VRAM, and no other Pod is placed on it.

Use several GPUs

gpu-count sets the number of GPUs for each Pod. All of them come from the same node. The request and limit apply to each GPU, so this Pod reserves 2 × 8Gi.

      annotations:
        tensor-fusion.ai/gpu-count: "2"
        tensor-fusion.ai/tflops-request: "30"
        tensor-fusion.ai/tflops-limit: "60"
        tensor-fusion.ai/vram-request: "8Gi"
        tensor-fusion.ai/vram-limit: "8Gi"

Inside the container, torch.cuda.device_count() returns 2. The allowed range is 1 to 128.

Pick the pool, model, vendor or card

Each of these annotations narrows the set of GPUs the scheduler may use. Combine them as needed. If nothing matches, the Pod stays Pending.

      annotations:
        tensor-fusion.ai/gpupool: "<gpupool-name>"   # default: the cluster's default pool
        tensor-fusion.ai/gpu-model: "<gpu-model>"    # exact value of status.gpuModel
        tensor-fusion.ai/vendor: "NVIDIA"            # exact value of status.vendor
        tensor-fusion.ai/tflops-request: "10"
        tensor-fusion.ai/tflops-limit: "20"
        tensor-fusion.ai/vram-request: "4Gi"
        tensor-fusion.ai/vram-limit: "4Gi"

To pin the Pod to specific cards on a node, list their device indices. The GPU count then equals the number of indices.

      annotations:
        tensor-fusion.ai/gpu-indices: "0,1"
        tensor-fusion.ai/tflops-request: "10"
        tensor-fusion.ai/tflops-limit: "20"
        tensor-fusion.ai/vram-request: "4Gi"
        tensor-fusion.ai/vram-limit: "4Gi"

Node-level constraints work as usual: the TensorFusion scheduler honors nodeSelector, node affinity, taints and tolerations in the Pod spec.

Use a GPU on another node

By default a Pod runs in local GPU mode: it is scheduled onto a GPU node and uses the card directly. In remote vGPU mode the Pod can run on any node, including one without a GPU, and its GPU calls are forwarded over the network to a worker Pod on a GPU node.

      annotations:
        tensor-fusion.ai/is-local-gpu: "false"
        tensor-fusion.ai/tflops-request: "10"
        tensor-fusion.ai/tflops-limit: "20"
        tensor-fusion.ai/vram-request: "4Gi"
        tensor-fusion.ai/vram-limit: "4Gi"

Check the result:

kubectl get tensorfusionconnection
kubectl get pods -A -l tensor-fusion.ai/component=worker,tensor-fusion.ai/workload=demo -o wide

A TensorFusionConnection exists for the Pod, and a worker Pod runs on a GPU node. The tensor-fusion.ai/gpu-ids annotation is on the worker Pod, not on your Pod.

Choose local mode when latency matters, and remote mode when the application must stay on CPU nodes or GPU nodes are scarce. The default for a pool is its defaultUsingLocalGPU setting.

Choose an isolation mode

tensor-fusion.ai/isolation decides how a Pod is separated from others on the same GPU. The default is soft.

ModeChoose it when
softYou trust the workloads and want the most flexible sharing. Limits can be changed while the Pod runs, so autoscaling works.
hardYou need stricter performance isolation between tenants. Limits are fixed when the Pod starts.
sharedThe Pod should own whole GPUs. See Use a whole GPU.
partitionedYou need hardware-level partitions such as NVIDIA MIG, and the administrator has set the cluster up for it.

Hard isolation:

      annotations:
        tensor-fusion.ai/isolation: "hard"
        tensor-fusion.ai/tflops-request: "20"
        tensor-fusion.ai/tflops-limit: "20"
        tensor-fusion.ai/vram-request: "4Gi"
        tensor-fusion.ai/vram-limit: "4Gi"

In local mode, a hard-isolated Pod gets an extra tensorfusion-worker container. With the default installation, a GPU serves one isolation mode at a time, so soft and hard Pods land on different cards.

Partitioned isolation with a named partition template:

      annotations:
        tensor-fusion.ai/isolation: "partitioned"
        tensor-fusion.ai/partition-id: "<partition-template-id>"
        tensor-fusion.ai/gpu-model: "<gpu-model>"
        tensor-fusion.ai/vendor: "<vendor>"

Partitioned mode does not work with the default isolationModePolicy: dynamic. Ask your administrator whether the cluster supports it and which template IDs exist.

Set the priority

      annotations:
        tensor-fusion.ai/qos: "critical"   # low | medium | high | critical
        tensor-fusion.ai/tflops-request: "20"
        tensor-fusion.ai/tflops-limit: "20"
        tensor-fusion.ai/vram-request: "8Gi"
        tensor-fusion.ai/vram-limit: "8Gi"

high and critical Pods without their own priorityClassName get the tensor-fusion-high or tensor-fusion-critical PriorityClass and can preempt lower-priority Pods when GPUs are full. See Create Workload for how to choose a level and what happens when you leave it out.

Choose which containers get the GPU

A Pod with more than one container must say which containers use the GPU. The others are left untouched and see no GPU.

      annotations:
        tensor-fusion.ai/inject-container: "app"   # comma-separated for several
        tensor-fusion.ai/tflops-request: "10"
        tensor-fusion.ai/tflops-limit: "20"
        tensor-fusion.ai/vram-request: "4Gi"
        tensor-fusion.ai/vram-limit: "4Gi"

When several containers use GPUs and each should get its own cards, add the count for each container. The sum must not exceed gpu-count.

      annotations:
        tensor-fusion.ai/inject-container: "trainer,eval"
        tensor-fusion.ai/gpu-count: "3"
        tensor-fusion.ai/container-gpu-count: '{"trainer":2,"eval":1}'
        tensor-fusion.ai/tflops-request: "30"
        tensor-fusion.ai/tflops-limit: "60"
        tensor-fusion.ai/vram-request: "8Gi"
        tensor-fusion.ai/vram-limit: "8Gi"

Without container-gpu-count, every listed container sees all GPUs of the Pod.

Let TensorFusion adjust the size

Start from an estimate and let the autoscaler correct the requests and limits from actual usage.

      annotations:
        tensor-fusion.ai/autoscale: "true"
        tensor-fusion.ai/autoscale-target: "all"   # compute | vram | all
        tensor-fusion.ai/tflops-request: "10"
        tensor-fusion.ai/tflops-limit: "20"
        tensor-fusion.ai/vram-request: "4Gi"
        tensor-fusion.ai/vram-limit: "4Gi"

Nothing is adjusted during the first 30 minutes after the workload is created. For schedules, percentiles and the status fields to watch, see Configure AutoScaling.

Start a group of Pods together

For distributed training, starting three of four workers wastes GPUs. With gang scheduling, no Pod of the group is bound until enough of them can be placed.

spec:
  replicas: 4
  template:
    metadata:
      labels:
        app: demo
        tensor-fusion.ai/enabled: "true"
      annotations:
        tensor-fusion.ai/gang-enabled: "true"
        tensor-fusion.ai/gang-min-members: "4"   # optional, defaults to all replicas
        tensor-fusion.ai/gang-timeout: "5m"      # optional, defaults to waiting indefinitely
        tensor-fusion.ai/tflops-request: "100"
        tensor-fusion.ai/tflops-limit: "100"
        tensor-fusion.ai/vram-request: "10Gi"
        tensor-fusion.ai/vram-limit: "10Gi"

The workload needs at least 2 replicas, and gang-min-members must be between 2 and the replica count. Progress is reported in status.gang of the TensorFusionWorkload.

Share settings across workloads

Put the common settings in a WorkloadProfile in the same namespace and reference it. Annotations on the Pod override the profile.

apiVersion: tensor-fusion.ai/v1
kind: WorkloadProfile
metadata:
  name: small-inference
spec:
  resources:
    requests:
      tflops: "10"
      vram: 4Gi
    limits:
      tflops: "20"
      vram: 4Gi
  qos: medium
      annotations:
        tensor-fusion.ai/workload-profile: "small-inference"
        tensor-fusion.ai/vram-limit: "6Gi"   # overrides the profile

Enable only some replicas

To try TensorFusion on part of a Deployment, cap the number of Pods it takes over. The remaining Pods are created unchanged.

      annotations:
        tensor-fusion.ai/enabled-replicas: "1"
        tensor-fusion.ai/tflops-request: "10"
        tensor-fusion.ai/tflops-limit: "20"
        tensor-fusion.ai/vram-request: "4Gi"
        tensor-fusion.ai/vram-limit: "4Gi"

For the full procedure, see Migrate Existing Workload.

Keep a Pod out of TensorFusion

A Pod without the label tensor-fusion.ai/enabled: "true" is not touched. If the administrator turned on auto migration, Pods that request nvidia.com/gpu are taken over even without the label; set the label to "false" to exclude one.

      labels:
        app: demo
        tensor-fusion.ai/enabled: "false"

Next steps

  • Annotation reference: every annotation, with type, default and allowed values.
  • Best practices: how to choose between the options on this page.
  • Configure AutoScaling

Table of Contents

Before you start
Share one GPU among several Pods
State compute as a percentage
Use a whole GPU
Use several GPUs
Pick the pool, model, vendor or card
Use a GPU on another node
Choose an isolation mode
Set the priority
Choose which containers get the GPU
Let TensorFusion adjust the size
Start a group of Pods together
Share settings across workloads
Enable only some replicas
Keep a Pod out of TensorFusion
Next steps