- Nvshmem Configure Nic Pe MappingRecommend NVSHMEM NIC-to-PE mappings and environment exports. Use for HCA selection, multi-NIC configuration, or topology-based mapping diagnostics.NVIDIA/nvshmem594
- Nvshmem DocsFind version-aware official NVSHMEM and NVSHMEM4Py documentation for releases, installation, APIs, runtime settings, transports, containers, and troubleshooting.NVIDIA/nvshmem594
- Nvshmem Enable TmaPrepare or review NVSHMEM CUDA kernels for TMA SMEM registration and direct-SMEM transfers. Do not use for unrelated CUDA tuning.NVIDIA/nvshmem594
- Nvshmem Get StartedUse when NVSHMEM beginners want a tutorial-style overview on assessment, mental models, first C/C++ or Python NVSHMEM programs, compilation, launching, and next steps.NVIDIA/nvshmem594
- Nvshmem InstallPlan and validate NVSHMEM and NVSHMEM4Py installations. Use for package, container, or source deployments.NVIDIA/nvshmem594
- Nvshmem Select Remote TransportSelect an NVSHMEM remote transport from target system and kernel evidence. Use for inter-node selection, compatibility checks, or configuration.NVIDIA/nvshmem594
- Nvshmem Troubleshoot And Report BugsDiagnose NVSHMEM runtime failures and prepare bug reports for launch, crashes, hangs, correctness, transport, or topology issues.NVIDIA/nvshmem594
- Nvshmem Tune PerformanceRoute NVSHMEM tuning to data collection, remote transport, NIC-to-PE mapping, or TMA. Do not use for unrelated CUDA, NCCL, or application tuning.NVIDIA/nvshmem594
- Cosmos3 Codebase NavNavigate the Cosmos3 package codebase to find where parameters, configs, defaults, scripts, and documentation live. Use when the user asks "where is X in cosmos3", "how do I find the config for Y", "where are the defaults", "where do I change a parameter", or any question about locating files, modules, or settings. Also use when the user opens or edits files and needs orientation.NVIDIA/cosmos-framework555
- Cosmos3 Codebase NavNavigate the Cosmos3 package codebase to find where parameters, configs, defaults, scripts, and documentation live. Use when the user asks "where is X in cosmos3", "how do I find the config for Y", "where are the defaults", "where do I change a parameter", or any question about locating files, modules, or settings. Also use when the user opens or edits files and needs orientation.NVIDIA/cosmos-framework555
- Cosmos3 Env TroubleshootDiagnose and fix Cosmos3 environment, installation, and runtime errors. Use when the user encounters an ImportError, ModuleNotFoundError, CUDA error, Docker error, checkpoint download failure, or any traceback during setup or inference.NVIDIA/cosmos-framework555
- Cosmos3 Env TroubleshootDiagnose and fix Cosmos3 environment, installation, and runtime errors. Use when the user encounters an ImportError, ModuleNotFoundError, CUDA error, Docker error, checkpoint download failure, or any traceback during setup or inference.NVIDIA/cosmos-framework555
- Cosmos3 InferenceGuide users through running Cosmos3 inference — offline batch generation, online serving with Ray and Gradio, parallelism options, input formats, sampling parameters, and prompt upsampling. Use when the user asks "how do I run inference", "how do I generate a video", "how do I serve the model", "what parameters should I use", or any question about running the model to produce outputs.NVIDIA/cosmos-framework555
- Cosmos3 InferenceGuide users through running Cosmos3 inference — offline batch generation, online serving with Ray and Gradio, parallelism options, input formats, sampling parameters, and prompt upsampling. Use when the user asks "how do I run inference", "how do I generate a video", "how do I serve the model", "what parameters should I use", or any question about running the model to produce outputs.NVIDIA/cosmos-framework555
- Cosmos3 Post TrainingGuide users through Cosmos3 supervised fine-tuning (SFT) post-training: preparing the example dataset and Wan2.2 VAE, converting the base checkpoint to DCP, launching distributed training (paired launch shell recommended, raw `torchrun` as an alternative), running T2V/I2V/V2V inference with the trained DCP checkpoint, and optionally exporting it to Hugging Face safetensors. Use when the user asks how to post-train Cosmos3, fine-tune on a custom video dataset, export a trained checkpoint, or invoNVIDIA/cosmos-framework555
- Cosmos3 Post TrainingGuide users through Cosmos3 supervised fine-tuning (SFT) post-training: preparing the example dataset and Wan2.2 VAE, converting the base checkpoint to DCP, launching distributed training (paired launch shell recommended, raw `torchrun` as an alternative), running T2V/I2V/V2V inference with the trained DCP checkpoint, and optionally exporting it to Hugging Face safetensors. Use when the user asks how to post-train Cosmos3, fine-tune on a custom video dataset, export a trained checkpoint, or invoNVIDIA/cosmos-framework555
- Cosmos3 SetupGuide users through Cosmos3 installation, environment setup, checkpoint downloading, and verification. Use when the user asks "how do I install cosmos3", "how do I set up the environment", "how do I download checkpoints", "how do I use Docker", or any question about getting the package running for the first time.NVIDIA/cosmos-framework555
- Cosmos3 SetupGuide users through Cosmos3 installation, environment setup, checkpoint downloading, and verification. Use when the user asks "how do I install cosmos3", "how do I set up the environment", "how do I download checkpoints", "how do I use Docker", or any question about getting the package running for the first time.NVIDIA/cosmos-framework555
- Api CallerCall any REST API dynamically. Make GET, POST, PUT, DELETE requests to any endpoint with custom headers and JSON body.NVIDIA/SkillEvaluator541
- CalculatorEvaluate mathematical expressions and unit conversions. Handles arithmetic, percentages, exponents, and common unit conversions (temperature, distance, weight). No external dependencies.NVIDIA/SkillEvaluator541
- Create Custom GraderUse when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.NVIDIA/SkillEvaluator541
- Task ListRequired for 4+ step requests; add tasks at start and update status after each step.NVIDIA/SkillEvaluator541
- Text AnalyzerAnalyze text content and produce statistics including word count, line count, character count, most frequent words, and readability metrics. Works on any plain text input provided inline or from a file path.NVIDIA/SkillEvaluator541
- Apply Inference OptimizationsApply FlashDreams-style inference speedups to model integrations after a baseline exists: bounded windows and fixed K/V caches, cache/decode overlap, `torch.compile`, CUDA graph capture, attention backend checks, decoder layout or replacement, transfer/materialization changes, and ordered presentation tuning. Use when porting known optimizations into a runner, demo, serving adapter, or downstream integration while preserving quality and reset behavior.NVIDIA/flashdreams510
- Flashdreams PostprocessingAdd or modify FlashDreams video post-processing processors, sessions, presets, and runner stream wiring. Use when implementing a new VideoPostProcessorConfig / VideoPostProcessor / VideoPostProcessorSession, registering a --postprocess.preset entry point, changing VideoPostprocessStream behavior, or reasoning about streaming buffering, layouts, per-view processing, distributed execution, or postprocess tests.NVIDIA/flashdreams510
- Flashdreams PreparationUse when changing FlashDreams V2 downloads, compilation, native loading, model construction, runtime validation, or application preloading.NVIDIA/flashdreams510
- Integrate A ModelEnd-to-end workflow for porting an external video diffusion model into a flashdreams integration — scope the architecture, scaffold a workspace-member plugin, reuse an existing recipe, write the checkpoint key-remap, layer model-specific conditioners, wire the runner, and verify with checkpoint weight-equality + upstream parity + a GPU rollout. Use when integrating a new model (e.g. a HuggingFace/research release) into flashdreams or a downstream repo, porting upstream weights, or reproducing anNVIDIA/flashdreams510
- Maintaining Oss StateMaintain FlashDreams's OSS-release state — the LICENSE / NOTICE / THIRD-PARTY-NOTICES / REUSE.toml / LICENSES/ / CONTRIBUTING.md collateral that satisfies OSRB Bug 6107043, the per-file SPDX headers, the third-party dependency manifest in THIRD-PARTY-NOTICES, and the pyproject.toml + uv.lock dependency pins. Use when adding or upgrading a runtime dependency, vendoring third-party source into the repo, adding a new first-party source file (any .py / .pyx / .pyi / .c / .cc / .cpp / .h / .hpp / .cuNVIDIA/flashdreams510
- Profile Model PerformanceInspect and baseline performance for FlashDreams-style model integrations and interactive demos: map the generation path, add trustworthy timing splits, build focused probes, and identify whether decode, model/denoise, cache, data transfer, or presentation dominates. Use when starting performance work on an existing model runner, demo, serving path, or downstream integration before implementing speedups. Pair with `apply-inference-optimizations` after the bottleneck is known and `validate-perforNVIDIA/flashdreams510
- Python Docstring StyleWrite Python docstrings and inline comments matching the flashdreams house style — SPDX header, one-line module docstring, Google-style function docstrings (Args/Returns/Raises), PEP 257 attribute docstrings on dataclass/class fields *and on module-level constants*, double-backticks for code references, imperative first sentences, and signpost-style inline block comments (kept, not stripped, on a tightening pass). Use when authoring or editing any .py file under flashdreams/, when adding a new mNVIDIA/flashdreams510
- Validate Performance QualityDesign benchmark, quality, and documentation validation for FlashDreams-style performance changes. Use when adding or updating sweep commands, profiler probes, decoder-quality comparisons, compile/cache probes, manual GPU validation, performance summaries, model cards, or README guidance after optimizing a model integration, demo, or serving path.NVIDIA/flashdreams510
- Ncu ReportAnalyze NVIDIA Nsight Compute (ncu) profiling reports (.ncu-rep files). Extract metrics, performance data, SASS/CUDA source, and identify bottlenecks. TRIGGER when: user asks to analyze, profile, or look at an ncu report, .ncu-rep file, Nsight Compute report, kernel performance/profiling data from ncu, or asks to generate/collect an ncu profile for a tilus kernel or example script. DO NOT TRIGGER when: user is writing unrelated profiling code.NVIDIA/tilus495
- Write DocsConvention and format for writing instruction docstrings and RST tutorials in tilus documentation. TRIGGER when: user asks to add, update, or write documentation for tilus instructions, instruction groups, or tutorials.NVIDIA/tilus495
- Cuopt DebuggingTroubleshoot cuOpt LP/MILP problems including errors, wrong results, infeasible solutions, performance issues, and status codes. Use when the user says something isn't working, gets unexpected results, or needs help diagnosing issues.NVIDIA/cuopt-examples472
- Cuopt Model MapperMap interpreted optimization problems into cuOpt-native models for the fast path with minimal clarifying questions.NVIDIA/cuopt-examples472
- Cuopt SandboxRun cuOpt in the NemoClaw sandbox — probe/smoke gates, prefer cancelable Python gRPC jobs, use legacy remote execution only when that API is unavailable, then vendored cuOpt skills.NVIDIA/cuopt-examples472
- Generic Max SupplyMulti-period supply chain planning model: data files, BOM structure, variable/constraint reference for the max-supply base model.NVIDIA/cuopt-examples472
- Optimization From Data OrchestratorCoordinate uploaded data plus a natural-language question into interpretation, clarification, cuOpt solve, and a user-facing answer.NVIDIA/cuopt-examples472
- Optimization Intent RouterClassify whether a data-backed request is LP, MILP, QP, routing, or non-optimization analytics.NVIDIA/cuopt-examples472
- Optimization Mode RouterChoose fast direct-to-cuOpt solve versus replayable or auditable model artifact mode.NVIDIA/cuopt-examples472
- Tabular Optimization IngestionInfer optimization structure from uploaded tables and identify minimal clarifications before cuOpt modeling.NVIDIA/cuopt-examples472
- CuequivarianceDefine custom groups (Irrep subclasses), build segmented tensor products with CG coefficients, create equivariant polynomials and IrDictPolynomials, and use built-in descriptors (linear, tensor products, spherical harmonics). Use when working with cuequivariance group theory, irreps, or segmented polynomials.NVIDIA/cuEquivariance441
- Cuequivariance JaxExecute equivariant polynomials in JAX using segmented_polynomial (naive/uniform_1d), the ir_dict workflow with IrDictPolynomial and dict[Irrep, Array], and Flax NNX layers (IrrepsLinear, SphericalHarmonics, IrrepsIndexedLinear). Use when writing JAX code with cuequivariance.NVIDIA/cuEquivariance441
- Cuequivariance TorchExecute equivariant tensor products in PyTorch using SegmentedPolynomial (naive/uniform_1d/fused_tp/indexed_linear), high-level operations (ChannelWiseTensorProduct, FullyConnectedTensorProduct, Linear, SymmetricContraction, SphericalHarmonics, Rotation), and layers (BatchNorm, FullyConnectedTensorProductConv). Use when writing PyTorch code with cuequivariance.NVIDIA/cuEquivariance441
- Aicr Analyzing SnapshotsUse when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster assessment report from a snapshot. Triggers on: snapshot analysis, cluster review, provider comparison, GPU topology, node health, snapshot report.NVIDIA/aicr439
- Aicr Auditing DocsUse when reviewing AICR's Markdown documentation for duplication, drift, bloat, and gaps — to keep docs high-value as the project evolves. Triggers on "audit the docs", "review documentation", "docs cleanup", "/aicr-auditing-docs", or any request to find redundant/stale/missing docs across README, docs/, demos/, and the root governance files. Produces a prioritized findings report (research, not edits) grouped by the five audit dimensions, anchored to the project's canonical sources of truth.NVIDIA/aicr439
- Aicr Creating Guided DemosScaffolds an interactive guided demo script (demos/*.sh), live or self-paced, with the Frame → Tell → Show → Close pattern. Triggers on "demo script", "guided walkthrough", "demos/*.sh", "live demo".NVIDIA/aicr439
- Aicr Creating Slide DecksUse when building a self-contained HTML slide deck or visual talking-point for a technical concept or workflow (e.g. a demos/*.html) — shown full-screen or projected and narrated, opening in any browser with no build step or dependencies.NVIDIA/aicr439
- Aicr Cross ReviewMulti-agent PR review using Claude Code, Codex, and CodeRabbit. Runs parallel reviews with integration impact analysis, then one cross-review round to a 2-of-3 consensus, with every confirmed finding adversarially verified by a fresh agent. Never runs the reviewed commit's code, and never posts unless explicitly asked. Use when asked for a thorough cross-review or multi-reviewer analysis. Requires the Codex plugin; CodeRabbit is best-effort. Claude Code only — uses the Workflow and Agent tools, NVIDIA/aicr439
- Aicr Managing OpenvexUse when adding, updating, or removing CVE/GHSA suppressions in `.openvex.json` — the OpenVEX document consumed by the weekly image vulnerability scan workflow. Triggers on "VEX", "OpenVEX", ".openvex.json", "suppress CVE", "ignore CVE", "vulnerability suppression", "aiperf-bench CVE", or any request to act on findings reported by `Weekly Image Vulnerability Scan` for the aiperf-bench image. Keeps the file current: adds reachability-evidenced statements for new HIGH+ findings, drops statements tNVIDIA/aicr439
- Aicr Release NotesUse when drafting the human-readable GitHub release notes summary for an upcoming AICR release. Triggers on "release notes", "draft release notes", "/aicr-release-notes", or any request to summarize commits since the last tag into a polished release announcement. Runs tools/changelog, groups commits into thematic highlights, mirrors the style of the previous release, and writes a Markdown draft to a temp file for hand-editing before publishing.NVIDIA/aicr439
- Aicr Reviewing Component DriftUse when reviewing the weekly AICR component drift report — the Slack digest and drift-report.json artifact produced by Registry Drift Report (registry-drift.yaml) listing which recipes/registry.yaml chart pins have moved upstream. Triggers on "review this week's drift", "component drift", "/aicr-reviewing-component-drift", "should we bump <component>", "what changed in <chart>", a pasted drift digest from Slack, or release prep that needs to know which chart pins are safe to advance. Gathers vaNVIDIA/aicr439
- Aicr TriageUse when the user runs `/aicr-triage` or asks to triage, review, or clean up a GitHub org-level Projects v2 board (default NVIDIA AICR project 248). Reviews active non-Done issues, then promotes P2 issues to P1, demotes Ready items to Backlog, closes superseded issues, and classifies unclassified ones — applying only user-confirmed changes via `gh` CLI. Also backfills Priority on Done issues that opened and closed between runs, and asserts that every issue on the board carries both Status and PrNVIDIA/aicr439
- Aicr Uat ReportUse when reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB200) x intent combinations are passing or failing in the UAT Run workflow (uat-run.yaml). Triggers on "UAT report", "/aicr-uat-report", "which UAT combos are failing", "UAT pass rate", "download the UAT debug bundle", "why did the UAT run fail", or RC/release-candidate validation prep that needs the combinations to test manually. Runs the bundled uat_report.py, classifies failures as prodNVIDIA/aicr439
- Check Mongodb Migration ReadinessReadiness check for migrating an NVSentinel installation from the Bitnami MongoDB backend to the Percona Operator backend. Runs the preflight script, interprets the verdict table, and captures the operator decisions (data handling, quarantined nodes, GitOps) that the migration skill requires. Use this first, before any migration step.NVIDIA/NVSentinel394
- Migrate Mongodb To PerconaExecute the NVSentinel MongoDB backend migration from Bitnami to the Percona Operator after readiness is confirmed: data dump (default preserve path), removal of the installation, cleanup of surviving objects, values preparation, and the Percona deployment. Destructive: on the opt-out clean path all health event data is wiped. Use only after check-mongodb-migration-readiness reports READY and the operator confirmed the decisions.NVIDIA/NVSentinel394
- Verify Mongodb Percona MigrationPost-migration verification for the NVSentinel Percona MongoDB backend: waits on the five install gates, performs the optional data restore and consumer restarts, and walks the operator through the aftermath expectations. Use after migrate-mongodb-to-percona completes, or on its own to health-check an existing Percona-backed installation.NVIDIA/NVSentinel394
- Fault Injection LoopClosed-loop fault injection and attribution accuracy benchmark. Draws from a prioritized pool of (fault_type, rank, iter, nodes) experiments and submits them 2 at a time via sbatch — waiting for each pair to finish before submitting the next — to bound filesystem load. GPU-related faults are front-loaded in the pool. After all jobs complete, runs /log-analysis and /fr-analysis on every experiment, scores attribution vs. ground truth, aggregates gaps, and iterates on attribution modules to close NVIDIA/nvidia-resiliency-ext338
- Fr AnalysisAnalyze PyTorch NCCL flight-recorder (FR) dumps to identify collective operation hangs and isolate the responsible ranks using CollectiveAnalyzer. Use when a distributed training job hangs due to an NCCL collective timeout and FR dump files are available. Detects the wavefront process group where collectives diverge and returns the root-cause suspect ranks.NVIDIA/nvidia-resiliency-ext338
- Log AnalysisAnalyze a SLURM job log file for failure root-cause attribution and restart decisions using the packaged NVRx restart-agent. Use when you have a SLURM training job log and need to determine why the job failed and whether it should be restarted. Performs deterministic evidence extraction plus optional LLM enrichment.NVIDIA/nvidia-resiliency-ext338