Run TAO training jobs on your GPU cluster.
GPU job execution
Submits TAO container training and inference jobs to Kubernetes clusters with NVIDIA GPU scheduling.
When to use it
Use it on EKS, GKE, AKS, or on-premises clusters with NVIDIA GPU support, or when integrating TAO into a Kubernetes-native ML platform.
Give it a container image, command, and GPU count; it submits a Kubernetes Job and returns job status and logs through the SDK.
What you provide
This skill
Kubernetes Jobs
Submits Kubernetes Jobs (irreversible)
S3 job-scoped result prefixes
Deletes S3 result prefixes (irreversible)
Kubernetes Job status and logs
Reads Job status and logs
setup-nvidia-gpu-host.sh
GPU host checks
Python is required to import and run the TAO and Kubernetes SDK clients.
The nvidia-tao-sdk package with its kubernetes extra must be installed; the documented pin is 7.1.0rc42.
The Kubernetes Python client must be installed.
An authenticated cluster reachable through kubeconfig or an in-cluster service account is required.