- Vss Search ArchiveUse this skill to run top-level VSS fusion search on archived video, or to ingest video files / RTSP streams for search. Do NOT use for ad-hoc visual Q&A (use vss-ask-video), live captioning (use vss-deploy-dense-captioning), or video summarization and reports (use vss-summarize-video).NVIDIA/skills3,503
- Vss Setup Behavior AnalyticsUse to deploy the vss-behavior-analytics service standalone (entrypoint, config-source, optional calibration). Not for the full warehouse deploy.NVIDIA/skills3,503
- Vss Setup Video Analytics ApiUse to deploy the vss-video-analytics-api REST service standalone (config-source, data-log bind, Elasticsearch, optional Kafka). Not for full warehouse deploy.NVIDIA/skills3,503
- Vss Summarize VideoUse to summarize a recorded video via the LVS summarization microservice (HITL-gated) with a VLM fallback. Not for report generation or live RTSP captioning.NVIDIA/skills3,503
- Warp Compile Time OptimizerUse when compile time or startup time is the problem in code that uses Warp: a request to improve, optimize, or cut compile times; an app that is slow to start or stalls at the first wp.launch; seconds of compiling before real work begins; JIT modules recompiling on every run or every CI job. Only applies when the code being optimized uses Warp kernels. Not for steady-state kernel runtime, memory, correctness, building Warp itself from source, or nvcc/C++ build times.NVIDIA/skills3,503
- Warp Debug GradientsUse to diagnose and fix incorrect gradients in differentiable Warp programs. Anything trained, optimized, calibrated, or fit through Warp kernels depends on wp.Tape gradients, so treat any misbehavior of such a workflow as a gradient problem until proven otherwise — use this when training diverges or NaNs, won't train at all, stalls or plateaus above the expected loss, converges to a wrong or biased answer, is worse than a reference implementation, works at small scale but fails at production scNVIDIA/skills3,503
- Warp EvalEvaluate whether an existing hot path is a credible NVIDIA Warp candidate. Use for irregular or spatial queries, particle or geometry simulation, branch-heavy loops, many small launches, host fallbacks, or large intermediates. CPU-only code and absent GPU dependencies are normal unless NVIDIA is prohibited. Exclude required cross-vendor or CPU-only deployment, vendor-lowered dense or NN layers, general Warp API questions, and already-selected Warp kernels. Contribution policy alone is not exclusNVIDIA/skills3,503
- Physicsnemo DiscoverOfficial NVIDIA-authored guidance for navigating PhysicsNeMo — pick the model, datapipe, or example for a SciML/AI4Science task (surrogates, forecasting, downscaling, physics-informed, inverse, generative). Points at existing files via live repo search; never writes code. Do NOT use for installation or environment setup, training-loop or other code authoring/scaffolding, contributor/CI/packaging questions, repo-specific questions in physicsnemo-sym/-cfd/-curator, or general (non-physics) ML/PyToNVIDIA/physicsnemo3,323
- Physicsnemo Shard TensorOfficial NVIDIA-authored guidance for PhysicsNeMo ShardTensor domain parallelism — integrate domain parallelism into training/inference scripts (new or existing) with DDP or FSDP2, write and register shard patches to enable new layers/ops, and bootstrap multi-GPU correctness tests. Use when working with ShardTensor, scatter_tensor, domain parallelism, sequence/spatial sharding, ring attention, DeviceMesh + DDP/FSDP2 hybrid parallelism, or physicsnemo.domain_parallel. Do NOT use for generic PyTorNVIDIA/physicsnemo3,323
- Nemo RetrieverUse when searching, extracting, ingesting, or querying a document collection with the NeMo Retriever 26.8.1 CLI, including local LanceDB indexes and deployed Retriever services. Use for PDFs, images, Office files, HTML, text, audio, and video; not for editing documents or web search.NVIDIA/NeMo-Retriever2,980
- Nemo Retriever McpUse when a task needs to search or add documents through NeMo Retriever MCP.NVIDIA/NeMo-Retriever2,980
- Nat Agent ConfigurationUse when selecting, configuring, composing, or troubleshooting NeMo Agent Toolkit agents and control-flow components, including ReAct, tool-calling, ReWOO, reasoning, router, sequential, parallel, and sub-agent patterns.NVIDIA/NeMo-Agent-Toolkit2,659
- Nat EvaluationUse when designing, configuring, running, or troubleshooting NeMo Agent Toolkit evaluations, datasets, evaluator selection, ATIF surfaces, quality gates, custom evaluators, and `nat eval`.NVIDIA/NeMo-Agent-Toolkit2,659
- Nat InstallationUse when installing or configuring NVIDIA NeMo Agent Toolkit, verifying the `nat` CLI, setting up optional extras, or creating a first hello-world workflow.NVIDIA/NeMo-Agent-Toolkit2,659
- Nat Mcp And ServingUse when serving NeMo Agent Toolkit workflows, exposing workflows through FastAPI, configuring MCP clients or servers, or troubleshooting transport and server setup.NVIDIA/NeMo-Agent-Toolkit2,659
- Nat OptimizationUse when configuring or running NeMo Agent Toolkit optimization with `nat optimize`, including Optuna parameter tuning, prompt evolution, optimizer sizing, output interpretation, and optimizer datasets.NVIDIA/NeMo-Agent-Toolkit2,659
- Nat Path ChecksUse when fixing NeMo Agent Toolkit documentation path-check failures, especially failed `ci/scripts/path_checks.py` output, slash-delimited text mistaken for paths, relative path references, Markdown code escaping, and path-check allowlist decisions.NVIDIA/NeMo-Agent-Toolkit2,659
- Nat TelemetryUse when adding, configuring, or troubleshooting NeMo Agent Toolkit logging, tracing, telemetry exporters, OpenTelemetry, Langfuse, LangSmith, Weave, Phoenix, profiling, or observability provider integrations.NVIDIA/NeMo-Agent-Toolkit2,659
- Nat Tools And FunctionsUse when authoring, registering, composing, or testing custom NeMo Agent Toolkit tools, functions, function groups, Python components, custom agents, custom evaluators, or advanced extension patterns.NVIDIA/NeMo-Agent-Toolkit2,659
- Nat User RulesUse first for general NVIDIA NeMo Agent Toolkit coding-agent behavior, task routing, naming conventions, component discovery rules, and cross-skill guidance.NVIDIA/NeMo-Agent-Toolkit2,659
- Nat Workflow CreationUse when creating, editing, validating, running, or troubleshooting NeMo Agent Toolkit workflow YAML, component discovery, LLM configuration, and common `nat` CLI commands.NVIDIA/NeMo-Agent-Toolkit2,659
- Skill EvolutionUse before creating, editing, or deciding whether to update any AI coding agent skill in this repository, including corrections to existing skill behavior, references, or routing.NVIDIA/NeMo-Agent-Toolkit2,659
- Cccl Add Review RuleUse when adding or extending a rule in docs/cccl/development/review_guidelines.md; describes how to research, word, place, and validate a new review guideline.NVIDIA/cccl2,524
- Cccl ReviewUse when reviewing CCCL code changes (a working-tree diff, PR, or commit range); loads the cccl-style and cccl-test skills and the CCCL review guidelines.NVIDIA/cccl2,524
- Cccl StyleUse when editing or reviewing CCCL code for style conventions; read common CCCL guidance and the path-specific references named by this skill.NVIDIA/cccl2,524
- Cccl TestUse when writing, updating, reviewing, or validating CCCL tests; read common CCCL test guidance and the path-specific references named by this skill.NVIDIA/cccl2,524
- Sass DiffUse when asked to check for SASS (or PTX) changes between commits, branches, or a local changeset; guides normalization, comparison, and reporting of CUDA disassembly diffs.NVIDIA/cccl2,524
- Trt Cpp Runtime QuickstartLoad and run a TensorRT engine (.plan / .engine) from C++ using the TensorRT 11 / 10.x **modern Runtime API**, avoiding the deprecated TRT 8.x binding-index APIs that older guidance still promotes. Use whenever the user asks about loading or running a TensorRT .plan/.engine from C++, even on "minimal example" requests — without this skill the default reply uses deprecated enqueueV2-style code. Also use when the user hits "Engine plan file is generated on an incompatible device", deserializeCudaENVIDIA/trt-samples-for-hackathon-cn1,672
- Trt Onnx QuickstartBuild and verify a TensorRT engine from a Hugging Face model ID or ONNX file, with numerical parity checked against ONNX Runtime. Use when the user imports a non-LLM model to TensorRT, needs a verified engine from ONNX, hits trtexec "unsupported operator", must verify the engine matches ONNX numerically, debugs a polygraphy parity failure (large max abs diff at FP16), or configures multi-input dynamic shapes. Triggers: convert ONNX to TensorRT, Hugging Face to TensorRT, trtexec onnx, trtexec unsNVIDIA/trt-samples-for-hackathon-cn1,672
- Trt Perf AnalysisValidate and analyze TensorRT performance data from paired layer-info JSON and profile/latency JSON files. Use when asked to inspect TensorRT, TRT, torch-tensorrt, or ONNX-TensorRT perf reports, verify that layer/profile JSON files are valid and from the same model, infer basic model information, find likely fusion or latency optimization opportunities, and produce a concise Markdown performance report or structured JSON data.NVIDIA/trt-samples-for-hackathon-cn1,672
- Trt Strong Typing MigrationMigrate a TensorRT build from weak typing (deprecated 10.12, removed 11.0) to strong typing — across Python INetworkDefinition builders, the trtexec CLI, and C++ builder code. Use when a TRT 11 upgrade breaks a weakly-typed build. Triggers: weakly typed to strongly typed, kSTRONGLY_TYPED, weak typing deprecated, kFP16/kINT8 removed, setPrecision rejected, setComputePrecision deprecated, do I still need --stronglyTyped, how to add the kSTRONGLY_TYPED flag, ModelOpt autocast, INT8 on TRT 11. NOT fNVIDIA/trt-samples-for-hackathon-cn1,672
- Trt Torch QuickstartCompile a PyTorch model to a TensorRT engine via Torch-TensorRT — AOT or JIT — under the new strong-typing default. Use when the user compiles PyTorch to TensorRT without ONNX, hits "enabled_precisions should not be used when use_explicit_typing=True", sees Dynamo graph breaks or PyTorch fallback, debugs ABI errors at import torch_tensorrt, or needs the compatible torch / torch_tensorrt / tensorrt-cu13 version pins for TensorRT 11. Triggers: torch_tensorrt, torch_tensorrt.dynamo.compile, torch.cNVIDIA/trt-samples-for-hackathon-cn1,672
- Kaizen UiKaizen UI (KUI) component library and design-pattern advisor for NVIDIA applications. Use whenever someone describes a UI to build, modify, or improve — even without explicit mention of KUI, "design pattern," or "best practice." Covers forms, filters, search, sort, cards, settings, activity feeds, navigation, data visualization, dashboards, loading states, empty states, feedback, iconography, element visibility, and page layout decisions. Also use when modifying existing UI (adding form fields, NVIDIA/Personal-AI-Router1,583
- Pair Github PrFills GitHub pull request descriptions with the required PAIR pair-release-intent:v1 block so the release-intent check passes. Use when opening, updating, or rewriting a PR description, when the user mentions release intent, version bumps, changelog for a PR, or when a workflow run fails on release-intent-check.NVIDIA/Personal-AI-Router1,583
- Pair Test StyleWrite, refactor, or review tests in Personal AI Router using focused cases, descriptive subtests, explicit error handling, and structured assertions. Use when adding or changing Go service or desktop tests; apply to the tests in scope rather than starting a repository-wide cleanup.NVIDIA/Personal-AI-Router1,583
- Sync BackendRefresh services/ from upstream, resolve JSON-RPC drift, update contract docs, rebuild cli-bin, and verify. Use when the backend changed or the user asks to sync dhc-modular / services.NVIDIA/Personal-AI-Router1,583
- Deploy AirgappedPrepare mirrors and transfer artifacts, configure DeepOps, deploy Slurm or Kubernetes GPU clusters without Internet access, and validate them with machine-readable gates. Use for disconnected, restricted-egress, offline, or air-gapped DeepOps installations and for diagnosing missing package, file, chart, or container artifacts.NVIDIA/deepops1,476
- Deploy K8s Gpu ClusterDeploy a Kubernetes GPU cluster with DeepOps (Kubespray + GPU Operator) and prove it schedules GPU pods. Use when asked to deploy or rebuild Kubernetes on GPU servers with this repository.NVIDIA/deepops1,476
- Deploy Slurm ClusterDeploy a Slurm GPU cluster with DeepOps and prove it works. Use when asked to deploy, install, or rebuild Slurm on one or more GPU servers with this repository.NVIDIA/deepops1,476
- Diagnose Driver InstallDiagnose NVIDIA driver installation failures on DeepOps-managed nodes — nvidia-smi errors, "No devices were found", DKMS build failures, or GPU pods crash-looping. Use before reinstalling anything.NVIDIA/deepops1,476
- Provision With MaasProvision or reinstall bare-metal servers and test VMs through Canonical MAAS, map deployed machines into DeepOps Ansible inventory with MAAS tags, validate access, or release them safely. Use when operating DeepOps with a MAAS-owned machine lifecycle.NVIDIA/deepops1,476
- Validate Gpu ClusterCheck whether a DeepOps-deployed Slurm or Kubernetes GPU cluster is healthy and report a machine-readable verdict. Use for health checks, post-deploy verification, "is the cluster working?" questions, and after any node or driver change.NVIDIA/deepops1,476
- Analysis MethodsTeaches the analyst agent how to write correct, robust Python analysis code for FHIR clinical data using pandas, matplotlib, and scipy.NVIDIA/dgx-spark-playbooks1,393
- Analysis MethodsTeaches the analyst agent how to write correct, robust Python analysis code for FHIR clinical data using pandas, matplotlib, and scipy.NVIDIA/dgx-spark-playbooks1,393
- Case SummaryPrepare a complete clinical case summary for a patient from FHIR endpoints. Use when asked to summarize a patient, compile a case, or prepare for tumor board.NVIDIA/dgx-spark-playbooks1,393
- Case SummaryPrepare a complete clinical case summary for a patient from FHIR endpoints. Use when asked to summarize a patient, compile a case, or prepare for tumor board.NVIDIA/dgx-spark-playbooks1,393
- Clinical DelegationHow to delegate clinical tasks to specialist agents. Always use sub-agent runtime with explicit agentId — never ACP. Never call FHIR via web_fetch.NVIDIA/dgx-spark-playbooks1,393
- Clinical DelegationHow to delegate clinical tasks to specialist agents. Always use sub-agent runtime with explicit agentId — never ACP. Never call FHIR via web_fetch.NVIDIA/dgx-spark-playbooks1,393
- Clinical KnowledgeTeaches agents clinical reference ranges, condition codes, quality measure definitions, drug classifications, and regulatory context so they can flag abnormal values and identify care gaps.NVIDIA/dgx-spark-playbooks1,393
- Clinical KnowledgeTeaches agents clinical reference ranges, condition codes, quality measure definitions, drug classifications, and regulatory context so they can flag abnormal values and identify care gaps.NVIDIA/dgx-spark-playbooks1,393
- Cohort CompareAnalyze a cohort of patients from FHIR endpoints to find care gaps and patterns. Use when asked to compare patients, find quality gaps, or analyze a population.NVIDIA/dgx-spark-playbooks1,393
- Cohort CompareAnalyze a cohort of patients from FHIR endpoints to find care gaps and patterns. Use when asked to compare patients, find quality gaps, or analyze a population.NVIDIA/dgx-spark-playbooks1,393
- Dgx DiagnoseDiagnose common DGX Station GB300 issues — CUDA crashes, wrong-GPU targeting, vLLM/SGLang container bugs, MIG state problems, NVLink/Fabric Manager errors, X/Vulkan failures, HuggingFace auth, and port conflicts. Use when the user reports a GPU error, inference server crash, MIG problem, or any unexplained DGX Station failure.NVIDIA/dgx-spark-playbooks1,393
- Dgx StationInspect and guide NVIDIA DGX Station GB300 development using the local dgx-assist CLI and pinned NVIDIA playbooks. Use for general Station platform questions, Software 1.0 or 2.0 compatibility, GB300 or RTX GPU selection, UUID ordering, mixed ATS/HMM coherency, CDMM, general containers, CDI, CUDA visibility, or vsloshd power-sloshing behavior. Do not use for vLLM or SGLang container selection or tuning, serving a named model, changing MIG, or troubleshooting a reported failure when the dedicatedNVIDIA/dgx-spark-playbooks1,393
- Dgx StationInspect and guide NVIDIA DGX Station GB300 development using the local dgx-assist CLI and pinned NVIDIA playbooks. Use for general Station platform questions, Software 1.0 or 2.0 compatibility, GB300 or RTX GPU selection, UUID ordering, mixed ATS/HMM coherency, CDMM, general containers, CDI, CUDA visibility, or vsloshd power-sloshing behavior. Do not use for vLLM or SGLang container selection or tuning, serving a named model, changing MIG, or troubleshooting a reported failure when the dedicatedNVIDIA/dgx-spark-playbooks1,393
- Dgx Station DiagnoseRun and interpret the complete read-only dgx-assist diagnostic suite for NVIDIA DGX Station GB300, correlate findings with pinned NVIDIA playbooks, export a redacted support bundle, and apply one separately approved allowlisted fix. Use when the user reports a Station, CUDA, GPU health, coherency, vsloshd, Docker, CDI, MIG, cache, port, or owned inference-service failure.NVIDIA/dgx-spark-playbooks1,393
- Dgx Station DiagnoseRun and interpret the complete read-only dgx-assist diagnostic suite for NVIDIA DGX Station GB300, correlate findings with pinned NVIDIA playbooks, export a redacted support bundle, and apply one separately approved allowlisted fix. Use when the user reports a Station, CUDA, GPU health, coherency, vsloshd, Docker, CDI, MIG, cache, port, or owned inference-service failure.NVIDIA/dgx-spark-playbooks1,393
- Dgx Station InferenceResolve, tune, preflight, launch, verify, inspect, and stop exact-model inference on NVIDIA DGX Station through dgx-assist. Use for vLLM or SGLang container selection, NGC versus upstream, GPU memory utilization, CPU or KV offload, HBM fit, KV-cache sizing, ISL or context length, prefix caching, chunked prefill, batching, concurrency, performance tuning, serving or deploying a named model, an OpenAI-compatible endpoint, Station recipe models, or an owned inference service. Require an exact modelNVIDIA/dgx-spark-playbooks1,393
- Dgx Station InferenceResolve, tune, preflight, launch, verify, inspect, and stop exact-model inference on NVIDIA DGX Station through dgx-assist. Use for vLLM or SGLang container selection, NGC versus upstream, GPU memory utilization, CPU or KV offload, HBM fit, KV-cache sizing, ISL or context length, prefix caching, chunked prefill, batching, concurrency, performance tuning, serving or deploying a named model, an OpenAI-compatible endpoint, Station recipe models, or an owned inference service. Require an exact modelNVIDIA/dgx-spark-playbooks1,393
- Dgx Station MigInspect NVIDIA DGX Station GB300 MIG state and installed-driver profiles, create a digest-bound layout plan, disclose disruption and restoration, and apply an approved still-valid plan. Use when the user asks to enable, disable, partition, reconfigure, inspect, or troubleshoot MIG instances or needs MIG UUIDs. Never assume static profile IDs or terminate GPU clients.NVIDIA/dgx-spark-playbooks1,393