- BcmInspect Base Command Manager (BCM) inventory and health evidence, and make explicitly requested changes when Base Command Manager edit mode is enabled.NVIDIA/AI-Factory-Operations-Agent24
- GrafanaCreate and open Grafana dashboards for metrics, Loki logs, and mixed panels. Use when the user asks for a dashboard, visualization, chart, panel, or graph of GPU, Base Command Manager (BCM), node, network, InfiniBand, Prometheus metrics, or Loki logs.NVIDIA/AI-Factory-Operations-Agent24
- Hardware AgentTriage hardware faults on NVIDIA GPU compute nodes through the Hardware Agent. Starting a triage is read-only log collection and hardware_analyze_dut is available in View mode; never require Edit or Auto. Before asking to start any fresh NVDebug collection, you MUST say: 'A fresh NVDebug collection can take 5-30 minutes. Do you want me to start it?' Stop and wait for the next user response. An ordinary affirmative such as yes, let's go, do it, or proceed is sufficient; start immediately without NVIDIA/AI-Factory-Operations-Agent24
- IraopResearch NVIDIA DGX, SuperPOD, GB200, NVLink, BlueField, Base Command Manager (BCM), and on-prem infrastructure documentation through the IRA/Sequoia retrieval agent. Use when the user asks to research docs, deployment policy, operational guidance, or newly uploaded knowledge-base documents.NVIDIA/AI-Factory-Operations-Agent24
- ObservabilityAnalyze cluster observability data and create Grafana dashboards. Use when the user asks about metrics, logs, log labels, cluster health history, dashboards, charts, panels, or visualizations.NVIDIA/AI-Factory-Operations-Agent24
- SlurmInspect mounted Slurm accounting exports and job logs for failed workload root cause analysis.NVIDIA/AI-Factory-Operations-Agent24
- Nvidia Ontology ManagementModel and publish semantic definitions in Auto Ontology. Use for terms, relationships, measures, imports, and governed results—not deployment or querying.NVIDIA/auto-ontology21
- Nvidia Ontology QueryQuery Auto Ontology and validate generated SQL, rows, and answers. Use for MCP or REST access, readiness, authentication, conversations, and grounded questions.NVIDIA/auto-ontology21
- Nvidia Ontology SetupSet up or troubleshoot the Auto Ontology runtime. Use for Helm (the official install), Docker Compose, developer setup, and MCP connection to an existing deployment.NVIDIA/auto-ontology21
- K8s Launch Kit CleanRemove an NVIDIA Network Operator deployment from a Kubernetes cluster with l8k clean. Use when the user explicitly asks to uninstall, tear down, remove, reset, or clean a Network Operator installation, delete its custom resources, or keep the Helm chart while clearing Network Operator CRs.NVIDIA/k8s-launch-kit16
- K8s Launch Kit ConfigUse this skill when the user needs help understanding, creating, or editing a k8s-launch-kit (l8k) configuration file (l8k-config.yaml or cluster-config.yaml). Activate for: config file questions, parameter tuning, subnet configuration, NV-IPAM setup, DOCA driver settings, maintenance concurrency, NIC configuration operator settings, changing MTU, VFs, resource names, or understanding what any config field does.NVIDIA/k8s-launch-kit16
- K8s Launch Kit DeployUse this skill when the user wants to deploy generated NVIDIA networking manifests to a Kubernetes cluster using k8s-launch-kit (l8k). Activate for: applying manifests, deploying to cluster, the `l8k deploy` subcommand or the legacy --deploy flag on `l8k generate`, applying generated files, or any mention of pushing l8k output to a live cluster. Even if the user just says 'apply these' or 'push to cluster' after generating manifests, use this skill.NVIDIA/k8s-launch-kit16
- K8s Launch Kit DiscoverUse this skill when the user wants to discover their Kubernetes cluster's network hardware capabilities using k8s-launch-kit (l8k). Activate for: cluster discovery, hardware detection, NIC detection, finding what GPUs or NICs are in a cluster, creating a cluster config file, or when the user says 'discover' in the context of l8k or NVIDIA networking.NVIDIA/k8s-launch-kit16
- K8s Launch Kit DryrunUse this skill when the user wants to preview what k8s-launch-kit (l8k) would deploy without making changes, or wants to safely validate their configuration before applying. Activate for: dry-run, preview, validation, 'what would happen if', testing configurations, schema discovery, checking generated manifests, or any cautious pre-deployment step. Also use when the user asks 'is my config valid' or 'show me what would be created' -- even without mentioning dry-run explicitly.NVIDIA/k8s-launch-kit16
- K8s Launch Kit GenerateUse this skill when the user wants to generate Kubernetes YAML manifests for NVIDIA networking deployment using k8s-launch-kit (l8k). Activate for: manifest generation, profile selection, choosing between SR-IOV/host-device/RDMA-shared/IPoIB/MacVLAN/Spectrum-X, creating deployment files, or when the user asks 'which profile should I use' or needs help choosing a network configuration.NVIDIA/k8s-launch-kit16
- K8s Launch Kit PipelineUse this skill when the user wants to run the full k8s-launch-kit (l8k) pipeline end-to-end: discover cluster hardware, select a profile, generate manifests, and deploy them all in one command. Also activate for CI/CD integration, automation pipelines, 'one-liner', 'complete workflow', or end-to-end NVIDIA networking deployment.NVIDIA/k8s-launch-kit16
- K8s Launch Kit Sharedk8s-launch-kit (l8k) CLI: Shared patterns for binary location, global flags, output formatting, exit codes, and error handling. Read this before using any other k8s-launch-kit skill.NVIDIA/k8s-launch-kit16
- K8s Launch Kit TroubleshootUse this skill when the user has problems with NVIDIA Network Operator on Kubernetes, or wants to analyze a sosreport diagnostic dump. Activate for: OFED driver crashes, SR-IOV pods failing, NicClusterPolicy errors, network operator pod issues, RDMA not working, NIC configuration failures, pods stuck in CrashLoopBackOff or ContainerCreating with network annotations, VF allocation issues, or when the user mentions 'troubleshoot', 'debug', 'sosreport', 'diagnose', or describes any NVIDIA networkinNVIDIA/k8s-launch-kit16
- K8s Launch Kit ValidateUse this skill when the user wants to verify that an NVIDIA networking deployment matches the configuration that produced it. Activate for: 'is my deployment correct', 'are all the manifests applied', 'does the network operator version match', 'verify deployment', 'check cluster state against config', or any question about whether the cluster reflects what l8k generated. Wraps the `l8k validate` subcommand.NVIDIA/k8s-launch-kit16
- K8s Network EngineerEmbody a senior NVIDIA Networking Engineer who is an expert on deploying cloud-native networking on Kubernetes with k8s-launch-kit (l8k). Activate whenever the user mentions NVIDIA network profiles, SR-IOV, RDMA, Spectrum-X, BlueField, ConnectX, NIC configuration, Network Operator, DOCA drivers, multirail networking, l8k, k8s-launch-kit, or any Kubernetes networking topic involving NVIDIA hardware. Also activate when the user asks general questions about high-performance networking, GPU interconNVIDIA/k8s-launch-kit16
- AnomalygenUse when running the PAIDF AnomalyGen pipeline over a defect dataset — mask placement, fine-tuning, synthetic defect-image generation (SDG), evaluation, quality refinement, and pseudo-labeling — even when the user names only one stage.NVIDIA/paidf-anomalygen14
- Anomalygen ReleaseUse when building the paidf-anomalygen Docker image (develop / product / air-gapped) from docker/Dockerfile — build the chosen target from source and run it. Not for training or synthetic defect-image generation (SDG).NVIDIA/paidf-anomalygen14
- CommitCommit current changes following the NVMesh Management commit message format (JIRA + type/scope + title + body). Use when the user asks to commit current changes with a JIRA reference, or mentions this commit format.NVIDIA/nvmesh-management13
- Digital Health Clinical Asr BuildStage 2 of the Clinical ASR Flywheel. Use when curating clinical terms, tagging IPA, and synthesizing a NeMo manifest. NOT for scoring (use /digital-health-clinical-asr-eval).NVIDIA/digital-health-skills10
- Digital Health Clinical Asr EvalStage 3 of Clinical ASR Flywheel. Score a NeMo manifest, produce the five-section KER leaderboard (by-ipa_source diagnostic). Not for ASR auth (/riva-asr).NVIDIA/digital-health-skills10
- Digital Health Clinical Asr FinetuneStage 4 of the Clinical ASR Flywheel. Use when priority KER is above 0.3 to run stock NeMo SFT on Parakeet TDT v2 and offline cycle N+1 re-eval. NOT for generic word boosting (use /finetune-asr).NVIDIA/digital-health-skills10
- Digital Health Clinical Asr SetupStage 1 of Clinical ASR Flywheel. Use when bootstrapping a cycle: NVCF+MW disclosure, NVIDIA_API_KEY check, deps install, TTS+ASR smoke test.NVIDIA/digital-health-skills10
- Nemo Clinical Data DesignerUse when generating synthetic tabular datasets via Data Designer — sampler columns, LLM columns, custom generators. Not for ASR audio.NVIDIA/digital-health-skills10
- Riva AsrUse when the user wants to deploy, run, or test an ASR (speech-to-text) Riva NIM — cloud-hosted (build.nvidia.com) or self-hosted Parakeet/Canary/Whisper.NVIDIA/digital-health-skills10
- Riva Asr CustomUse when the user wants to deploy a custom-trained ASR model as a Riva NIM, or convert a NeMo model via nemo2riva / riva-build / riva-deploy / RMIR.NVIDIA/digital-health-skills10
- Riva Nim SetupUse when getting started with NVIDIA Riva Speech NIMs: NGC access, Docker login for nvcr.io, NVIDIA Container Toolkit, GPU verification, Riva Python client.NVIDIA/digital-health-skills10
- Riva TtsUse when the user wants to deploy, run, or test a TTS (speech-synthesis) Riva NIM — cloud-hosted (build.nvidia.com) or self-hosted Magpie / voice cloning.NVIDIA/digital-health-skills10
- Ai Inference RecipeCreates single-point srt-slurm recipes from NVIDIA AI Inference benchmark rows. Use when the user mentions ai-inference, NVIDIA inference performance pages, benchmark rows, or asks to create a non-sweep recipe from model/GPU/framework/sequence/concurrency details.NVIDIA/srt-slurm-recipes8
- Paidf Auto LabelingUse when a user needs to get started with PAIDF Auto-Labeling, plan a scenario, run or debug a shipped cookbook, author prompts or cookbooks, migrate a pipeline, or configure a stage. Confirm critical inputs (data path, output path, endpoints) and ask when any are missing. This is a router: read the matching reference instead of inventing a workflow.NVIDIA/paidf-auto-labeling5
- Fleet Health ReportGenerate a standalone fleet-wide HTML health snapshot from live nvfleetint data, including node health, capacity, active-alert impact, recent errors, and machines needing immediate attention. Use for fleet dashboards, executive summaries, or scoped fleet reports. Do not use for a single-node root-cause investigation.NVIDIA/fleet-intelligence-client3
- Node Rca RccaInvestigate one NVIDIA Fleet Intelligence node and generate an evidence-backed HTML RCA/RCCA from live current and historical alerts plus authoritative corrective-action research. Use for node incident analysis, root-cause analysis, corrective actions, or post-incident reports.NVIDIA/fleet-intelligence-client3
- NvfleetintQuery NVIDIA Fleet Intelligence with the nvfleetint CLI. Use for ad hoc questions about fleets, nodes, GPUs, node groups, compute zones, alerts, agent health, firmware, verification, inventory, errors, or authentication. For a fleet-wide HTML snapshot use fleet-health-report; for a single-node RCA/RCCA use node-rca-rcca.NVIDIA/fleet-intelligence-client3
- Update Non NormativeUse when updating or reviewing yaml-sigil-spec non-normative companion material after specification changes, including diagrams, SVG and PNG images, schema-adjacent documentation, conformance indexes, examples, and cross-references.NVIDIA/yaml-sigil-spec2
- Yaml Sigil Traits Spec UpdateUse when updating the pinned yaml-sigil-spec submodule in yaml-sigil-traits or reconciling its public trait and DTO vocabulary after YamlSigil specification changes.NVIDIA/yaml-sigil-traits2
- Paidf Orchestration SetupAudit, prepare, and deploy PAIDF Orchestration on a Kubernetes GPU cluster - single-GPU H100/L40S hosts, managed Kubernetes, kubeadm, and similar. Select for requests to set up, install, deploy, configure, or check a PAIDF Orchestration environment; run a workflow on a new or unverified GPU host; connect via kubeconfig; validate GPU compute; deploy the Airflow controller; or choose external versus in-cluster model services. A plain SSH host is not a supported backend.NVIDIA/paidf-orchestration1
- Paidf Orchestration Write DagUse when a user describes a custom PAIDF Orchestration pipeline — a specific ordered combination of stages such as augmentation only, auto-labeling only, detection+captioning only, or image attribute augmentation without full auto-labeling — that no existing DAG in airflow/dags/workflows/ covers, and asks for a new Kubernetes DAG. Also use to check that a generated or existing DAG's model/container/prompt choices match an external spec document (e.g. a PAIDF `launchable.md`).NVIDIA/paidf-orchestration1
- Physical Ai Event Video GenerationRun the PAIDF Orchestration Event Video Generation DAG on Kubernetes - image-to-video anomaly generation, auto-labeling, and anomaly dataset generation. Select for requests about event video generation, anomaly video generation, image-to-video synthesis, Cosmos3 image2video, anomaly dataset creation, safety/surveillance SDG, or generating person-falling, person-climbing, person-running, fighting, smoking/vaping, fire/smoke, or shoplifting video clips from a seed image. Runs environment setup firNVIDIA/paidf-orchestration1
- Physical Ai Image Attribute AugmentationRun the PAIDF Orchestration Image Attribute Augmentation DAG on Kubernetes - person-crop clothing augmentation, attribute search, and augmented dataset generation. Select for requests about image attribute augmentation, person attribute search, person re-identification data, clothing augmentation, attribute captions, augmentation payloads, run status, or result retrieval. Runs environment setup first when controller readiness is unknown. Not for video or defect-image generation.NVIDIA/paidf-orchestration1
- Paidf Curation And RetrievalUse when operating PAIDF Curation and Retrieval or NVIDIA Cosmos Curator pipelines (split, filter, caption, embed, dedup, shard, image annotate) or PAIDF Data Mining nearest-neighbor matching on Curator embeddings. Activate for Make or CLI pipeline config, GPU run preflight, FFmpeg sidecar, SAM3 keys, or Curator-to-TAO handoff. Do not use for generic ETL, vector-database RAG, model training, orchestration, or embeddings outside Cosmos Curator and PAIDF Data Mining.NVIDIA/paidf-curation-and-retrieval