- Nemo Relay Debug Runtime IntegrationUse this skill when NeMo Relay is installed or imported but application-side runtime behavior is missing or incorrect, including load failures, inactive scopes, missing events, and plugin or adaptive wiring problems.NVIDIA/NeMo-Relay190
- Nemo Relay Get StartedUse this skill when first-time NeMo Relay users want to try Relay, choose the least-complex supported quick start, or verify initial value through the CLI, a maintained integration, or direct Python, Node.js, or Rust instrumentation before production setup.NVIDIA/NeMo-Relay190
- Nemo Relay InstallUse this skill when choosing or running NeMo Relay installation for the CLI, Python, Node.js, Rust, OpenClaw, or maintained framework integrations, or when explaining Hermes Agent's built-in Relay integration.NVIDIA/NeMo-Relay190
- Nemo Relay Instrument CallsUse this skill when an application owns tool or LLM/provider call sites and needs to wrap them with NeMo Relay scopes and managed execution APIs for lifecycle events, middleware, or guardrails.NVIDIA/NeMo-Relay190
- Nemo Relay Instrument Context IsolationUse this skill when concurrent requests, async tasks, threads, workers, goroutines, or agents need independent NeMo Relay scope stacks and correct ancestry propagation.NVIDIA/NeMo-Relay190
- Nemo Relay Instrument Typed WrappersUse this skill when adding NeMo Relay typed wrappers, domain types, or provider codecs while preserving JSON middleware semantics and caller-visible behavior.NVIDIA/NeMo-Relay190
- Nemo Relay Integrate UpstreamUse this skill when assessing, extending, or implementing NeMo Relay support in an agent harness or agent framework, including coding agents and orchestration runtimes, when the host lacks Relay support or an existing integration needs deeper coverage. It identifies the host's execution boundaries and extension points, selects an appropriate attachment method for each boundary, and verifies the resulting coverage.NVIDIA/NeMo-Relay190
- Nemo Relay Migrate From FlowUse this skill when migrating applications, examples, integrations, documentation, manifests, or repository code from NeMo Flow to NeMo Relay across Python, Rust, Node.js, Go, C FFI, CLI, configuration, and observability surfaces.NVIDIA/NeMo-Relay190
- Nemo Relay Plugin Adaptive TuningUse this skill when baseline NeMo Relay instrumentation exists and the user wants to configure or evaluate adaptive plugin behavior, including telemetry, state, adaptive_hints, tool_parallelism, ACG, hint consumption, or measured rollout.NVIDIA/NeMo-Relay190
- Nemo Relay Plugin BuildUse this skill when building or packaging reusable NeMo Relay runtime behavior as an embedded configuration component or a manifest-backed `rust_dynamic` native or `worker` gRPC plugin, with deterministic validation and rollback-safe registration.NVIDIA/NeMo-Relay190
- Nemo Relay Plugin ObservabilityUse this skill when choosing or configuring NeMo Relay 0.6 or 0.7 observability through the built-in plugin, subscribers, or exporters, including raw ATOF events, ATIF trajectories, OpenTelemetry, OpenInference, or custom event handling.NVIDIA/NeMo-Relay190
- Prepare Code FreezeExecute a requested NeMo Relay code freeze by cutting a release branch, updating nightly alpha configuration, bumping main, and preparing the required PR. Do not use for an ordinary version bump or release-note draft.NVIDIA/NeMo-Relay190
- Prepare PrPrepare, create, publish, or edit a NeMo Relay pull request or its body using the repository template and contributor requirements. Do not use for code review or implementation without PR preparation.NVIDIA/NeMo-Relay190
- Rename SurfacesRename a NeMo Relay repository, package, crate, module, public symbol, import path, or brand surface across multiple consumers. Do not use for a file-local private rename.NVIDIA/NeMo-Relay190
- Review Doc StyleReview NeMo Relay documentation, examples, or public text for NVIDIA technical writing style, terminology, and repository accuracy. Do not use for ordinary implementation review.NVIDIA/NeMo-Relay190
- Update Project VersionPerform a NeMo Relay project release-version change across project-owned manifests, dependencies, lockfiles, and generated attribution surfaces while keeping just set-version coverage complete. Do not use for dependency-only updates or adding a package when no project version bump is requested.NVIDIA/NeMo-Relay190
- Nvalchemi Data StorageHow to write, read, compose, and load atomic data using nvalchemi's composable Zarr-backed storage pipeline (Writer, Reader, Dataset, MultiDataset, DataLoader). Use when saving simulation outputs or trajectories to disk, converting structures (e.g. ASE / extxyz) into Zarr stores, assembling datasets for training or inference, or wiring a DataLoader to stream batches to the GPU.NVIDIA/nvalchemi-toolkit188
- Nvalchemi Data StructuresHow to use AtomicData and Batch, the core graph-based data structures for representing atomic systems and batching them for GPU computation. Use when building systems from positions, cells, and atomic numbers, converting from ASE Atoms, batching or unbatching structures, reading per-atom vs per-graph tensors, grouping graphs within a batch (e.g. NEB path images or ensemble members) via group_idx/GroupLayout, or debugging shape, dtype, or device errors in model inputs.NVIDIA/nvalchemi-toolkit188
- Nvalchemi DistillationHow to distill a large teacher MLIP into a small student with DistillationStrategy — teacher signals and offline dataset labeling, the teacher_* loss targets, the on-policy segment loop (propagator, replay buffer, mixed loader), checkpoints and restart, and accuracy/stability evaluation with acceptance thresholds. Use when training a small student to reproduce a big model's energies, forces, stress, or per-atom energies, generating training frames from the student's own trajectories, or gating aNVIDIA/nvalchemi-toolkit188
- Nvalchemi DistributedHow to run domain-decomposed (multi-GPU) MLIP simulations with DomainParallel — choose between the halo and graph-partition strategies, author a distribution_spec so a bring-your-own model runs under domain decomposition, and write a custom dynamics integrator that stays correct across ranks.NVIDIA/nvalchemi-toolkit188
- Nvalchemi Dynamics ApiHow to configure and run dynamics simulations, compose multi-stage pipelines (FusedStage, DistributedPipeline), use inflight batching, and manage data sinks. Use when writing any simulation script — molecular dynamics (NVE/NVT), structure relaxation or geometry optimization (e.g. FIRE2), equation-of-state or adsorption scans — or orchestrating many structures through a batched GPU pipeline.NVIDIA/nvalchemi-toolkit188
- Nvalchemi Dynamics HooksHow to use and write dynamics hooks — callbacks that observe or modify batch state at specific points during each simulation step. Use when a simulation needs neighbor-list rebuilds, convergence checks or early stopping, temperature control, per-step logging or trajectory capture, or any custom per-step behavior attached to a dynamics run.NVIDIA/nvalchemi-toolkit188
- Nvalchemi Dynamics ImplementationHow to implement a dynamics integrator by subclassing BaseDynamics and overriding pre_update() and post_update() methods. Use when creating a custom integrator, optimizer, or sampler that the built-in stages do not provide; for configuring existing dynamics, see nvalchemi-dynamics-api.NVIDIA/nvalchemi-toolkit188
- Nvalchemi Fine TuningHow to fine-tune nvalchemi-compatible models with FineTuningStrategy, pretrained checkpoint initialization, module patches, trainable-parameter filters, conservative optimizer defaults, validation, restart checkpoints, and model-agnostic MACE, AIMNet2, custom BaseModelMixin, or PyTorch inputs. Use when adapting a pretrained MLIP (e.g. MACE-MP) to new reference data, freezing or patching submodules during training, or resuming an interrupted fine-tune from a checkpoint.NVIDIA/nvalchemi-toolkit188
- Nvalchemi Loss ApiHow to use built-in loss functions and implement custom losses using the BaseLossFunction template-method pattern — residual types, per-atom normalization, masking, and graph-balanced reductions. Use when choosing or weighting energy, force, or stress objectives for training or fine-tuning, masking atoms or graphs out of the loss, or writing a custom loss term.NVIDIA/nvalchemi-toolkit188
- Nvalchemi MepHow to compute minimum-energy paths (MEPs) and transition-state estimates with nvalchemi.dynamics.mep: building initial paths (interpolation, alignment, IDPP) and running batched nudged elastic band (NEB) with the NEB strategy or by building it manually with hooks. Use when setting up reaction paths between reactant and product structures, running regular or climbing-image NEB, reading NEB diagnostics, or writing a custom NEB force method or spring policy.NVIDIA/nvalchemi-toolkit188
- Nvalchemi Model WrappingHow to wrap an arbitrary MLIP (Machine Learning Interatomic Potential) using the BaseModelMixin interface to standardize inputs, outputs, and embeddings. Use when integrating a model such as MACE or AIMNet2 (e.g. MACEWrapper, loading pretrained checkpoints) so dynamics, training, or fine-tuning stages can call it, or when exposing energies, forces, or embeddings from a custom PyTorch model.NVIDIA/nvalchemi-toolkit188
- Nvalchemi ReportingHow to add observability to nvalchemi dynamics and training workflows using ReportingOrchestrator, RichReporter, TensorBoardReporter, scalar extraction, custom reporter callbacks, and dynamics LoggingHook. Use when showing live progress, writing TensorBoard summaries, preserving dynamics CSV rows, adding rank-safe distributed reporting, previewing Rich dashboards, or deciding between logging and reporting for training or molecular dynamics runs.NVIDIA/nvalchemi-toolkit188
- Nvalchemi Training ApiHow to configure nvalchemi training workflows with TrainingStrategy, custom training functions, standalone or composed losses, loss-weight schedules, optimizer and scheduler configs, validation, hooks, restartable checkpoints, model-agnostic inputs, and scaling to multiple GPUs or nodes with DistributedManager and DDPHook. Use when training a model from scratch, setting up optimizers, schedulers, validation, or checkpointing, or scaling a run across GPUs or nodes (DDP); for adapting a pretrainedNVIDIA/nvalchemi-toolkit188
- Nvalchemi Zarr PerfPerformance tuning for nvalchemi's Zarr-backed Reader, Dataset, and DataLoader pipeline. Use when configuring AtomicDataZarrReader, Dataset, DataLoader, ZarrWriteConfig, or nvalchemi-io-test for training/inference throughput, especially shuffled access, graph-like random access, fused prefetch, pinned memory, validation overhead, or Zarr chunk/shard choices.NVIDIA/nvalchemi-toolkit188
- Nim Operator InstallInstall NVIDIA NIM Operator on Kubernetes with prerequisite checks, optional NVIDIA GPU Operator dependency installation, public or local Helm chart selection, optional Dynamo support, and optional KServe compatibility verification. Use when a customer wants to install or upgrade the NIM Operator itself, with or without Dynamo and KServe, but does not want to deploy a NIM inference model yet.NVIDIA/k8s-nim-operator160
- Nim Operator UninstallSafely uninstall NVIDIA NIM Operator from Kubernetes with inventory checks, explicit approval gates for destructive actions, optional custom resource cleanup, optional CRD removal, and post-uninstall validation. Use when a customer wants to remove or clean up the NIM Operator itself, not the GPU Operator or unrelated cluster dependencies.NVIDIA/k8s-nim-operator160
- Physicsnemo Cfd Create Custom MetricCreate a custom metric for the PhysicsNeMo CFD benchmarking workflow. Use when the user wants to add a new evaluation metric, implement a custom error measure, compute force coefficients, or extend the benchmark with domain-specific quantities.NVIDIA/physicsnemo-cfd149
- Physicsnemo Cfd Create Dataset AdapterCreate a new dataset adapter for the PhysicsNeMo CFD benchmarking workflow. Use when the user wants to add a new CFD dataset, write a DatasetAdapter, integrate a new mesh format, or benchmark models on custom data.NVIDIA/physicsnemo-cfd149
- Physicsnemo Cfd Create Model WrapperCreate a new model wrapper for the PhysicsNeMo CFD benchmarking workflow. Use when the user wants to add a new CFD model, write a CFDModel wrapper, integrate a new neural network architecture, or run a custom model through the benchmarking pipeline.NVIDIA/physicsnemo-cfd149
- Compileiq Author ObjectiveUse when writing the objective_function= passed to Search(). Covers the two legal signatures (compiler-only str vs mixed list), the baseline-knockout branch, per-eval cache busting, framework-specific --apply-controls injection (raw PTXAS, NVCC, Triton, Helion, cuTeDSL/FA4, FlashInfer), correctness-before-timing, INVALID_SCORE handling, and the Debug-pack O0/O3 ACF-injection canary that must pass before launching a search. Triggers on "objective function", "apply-controls", "INVALID_SCORE", "savNVIDIA/CompileIQ137
- Compileiq Booster PackUse BEFORE running a full CompileIQ search. Walks through downloading a Booster Pack from NVIDIA/CompileIQ GitHub Releases, applying ACF candidates one at a time to the user's compiler (raw PTXAS, NVCC, Triton, Helion, FlashInfer), and keeping only candidates that compile, pass correctness, and beat the no-ACF baseline. Includes the mandatory Debug-pack O0/O3 ACF-injection canary that proves the ACF is reaching PTXAS. Triggers on "booster pack", "ACF", "apply-controls", "speed up without searchiNVIDIA/CompileIQ137
- Compileiq BootstrapUse when starting a fresh CompileIQ project, hitting a socket timeout, or before running any other compileiq-* skill. Verifies CUDA 13.3+, ptxas, GPU access, that `from compileiq.ciq import Search` and friends resolve, and that `PtxasSearchSpace().retrieve()` returns a real path. Documents the env vars that control timeouts, caching, and search-space mirroring. Triggers on "set up compileiq", "compileiq doesn't work", "socket timeout", "where do search spaces come from", "air-gapped compileiq".NVIDIA/CompileIQ137
- Compileiq DebugUse when something is wrong: Search() hangs, all evaluations return INVALID_SCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is too high, or a winning ACF candidate needs NCU profiling to explain. Symptom-indexed table on top. Triggers on "compileiq hang", "socket timeout", "INVALID_SCORE", "not converging", "every score is the same", "TypeError fromhex", "ncu profile", "register spill", "ptxas error", "not in expected format", "high cv".NVIDIA/CompileIQ137
- Compileiq Run SearchUse when composing the Search(...) call and calling .start(). Covers the four worker classes (MultiProcessWorker / IsoMultiProcessWorker / RayWorker / AsyncWorker) and when to pick each, SearchConfiguration sizing rules, dump_results checkpointing, tracker_config choice (Disabled / Loguru / MLflow), num_workers/task_timeout semantics, and GPU clock locking for stable measurements. Triggers on "Search()", "tuner.start()", "pool_size", "num_workers", "task_timeout", "IsoMultiProcessWorker", "RayWoNVIDIA/CompileIQ137
- Compileiq Search SpaceUse when picking the search_space= argument for Search(). Covers the three provider classes (PtxasSearchSpace, NvccSearchSpace, LocalSearchSpaceBin), how to pin a version/variant/tag, the attention-focused 'att' variant for attention kernels (FlashAttention, GQA, MHA, MLA, FlashInfer Batch Decode), air-gapped mirroring via CIQ_SEARCH_SPACES_DIR, and custom user-defined search spaces built from compileiq.search_spaces.base primitives. Triggers on "search space", "PtxasSearchSpace", "NvccSearchSpaNVIDIA/CompileIQ137
- Compileiq Validate ResultUse AFTER a Search has completed and BEFORE claiming any speedup or shipping an ACF. Loads the dump_results CSV, extracts top-K candidates (single-objective) or the Pareto front (multi-objective), re-measures each against the no-ACF baseline with 100+ trials on fresh caches, runs Welch's t-test plus Cohen's d, rejects three classic false-positive patterns (lucky-min / higher-variance / multiple-comparisons-of-N), and saves the validated winner as best.acf. Triggers on "validate result", "extractNVIDIA/CompileIQ137
- Tripy CompilationWork with the nvtripy compilation pipeline. Use when: using tp.compile, creating InputInfo or DimensionInputInfo, understanding the Trace → MLIR → TensorRT flow, configuring optimization levels, working with Executable objects, debugging compilation, using dynamic shapes or NamedDimension.NVIDIA/TensorRT-Incubator125
- Tripy ConstraintsAuthor input/output constraints for nvtripy operations using the declarative constraint DSL. Use when: defining input_requirements or output_guarantees, writing @wrappers.interface decorators, auto-casting dtypes, using GetInput/GetReturn/OneOf/If/Equal, debugging constraint validation errors.NVIDIA/TensorRT-Incubator125
- Tripy DebuggingDebug and diagnose errors in nvtripy code. Use when: interpreting TripyException stack traces, enabling MLIR/TensorRT debug output, understanding error reporting with stack_info, using raise_error, configuring debug environment variables, tracing compilation failures.NVIDIA/TensorRT-Incubator125
- Tripy DocumentationWrite API documentation for nvtripy following project conventions. Use when: writing docstrings for ops or modules, adding code examples, using @export.public_api document_under paths, creating Sphinx RST cross-references, understanding the docs build pipeline.NVIDIA/TensorRT-Incubator125
- Tripy New ModuleAdd a new neural network module to nvtripy. Use when: creating an nn layer, implementing a Module subclass, adding a new layer like Linear/LayerNorm/Conv, defining parameters with DefaultParameter or OptionalParameter, using constant_fields decorator.NVIDIA/TensorRT-Incubator125
- Tripy New OperationAdd a new operation to nvtripy. Use when: implementing a new op, adding a frontend op, creating a trace op, registering an op in the API. Covers the full Frontend → Trace → MLIR pipeline including export decorators, constraint definitions, and init registration.NVIDIA/TensorRT-Incubator125
- Tripy TestingWrite tests for nvtripy following project conventions. Use when: adding tests for ops, modules, trace operations, or compilation, using pytest parametrize, testing error cases with helper.raises, testing dtype combinations, understanding test directory structure.NVIDIA/TensorRT-Incubator125
- Axe A11yAutomated web accessibility scanning (axe-core) driven by patchright + real Google Chrome — audits sites behind bot walls and behind authNVIDIA/nemoclaw-community120
- Blackwall Payment GateScreen x402 payments with an advisory Blackwall verdict, then submit a payment intent to the host-side release gate. You prepare payments; only the gate can settle them.NVIDIA/nemoclaw-community120
- Blender Host Sandbox BoundaryRoute Blender, USD, OVRTX, and OVPhysX work when Hermes terminal and file tools run in an OpenShell sandbox but Blender MCP runs inside a host Blender process. Use before any task that inspects the active scene, mentions host or sandbox paths, exports or validates files, or needs artifacts to cross the boundary.NVIDIA/nemoclaw-community120
- Blender Python Api VerificationVerify Blender Python properties, enums, operators, animation APIs, and add-on interfaces against the running Blender build before writing or executing uncertain bpy code. Use for any Blender scene edit, render setup, animation, import/export, or add-on task where an API detail could vary by version.NVIDIA/nemoclaw-community120
- Coach Nemoclaw HermesCoach a NemoClaw Hermes sub-agent through Blender, OVRTX, OVPhysX, USD, rendering, simulation, and live Blender-control tasks. Use when Codex should delegate specialized execution to Hermes while managing sandbox boundaries, preventing overlapping runs, allowing tool iteration, intervening only on evidence, and independently validating artifacts.NVIDIA/nemoclaw-community120
- Coordinate Nemoclaw BlenderLegacy compatibility entry point for prompts that explicitly request coordinate-nemoclaw-blender. Use coach-nemoclaw-hermes for new Codex-coached Blender, OVRTX, OVPhysX, USD, rendering, or simulation tasks.NVIDIA/nemoclaw-community120
- Cross Source Gap AnalysisCompare findings across Slack, GitHub, NVIDIA forums, and Outlook to identify alignment gaps, missing coverage, and follow-ups.NVIDIA/nemoclaw-community120
- Cross Source Gap AnalysisCompare findings across Slack, GitHub, NVIDIA forums, and Outlook to identify alignment gaps, missing coverage, and follow-ups.NVIDIA/nemoclaw-community120
- Deep ResearchQueue deep, multi-step research and analysis tasks to the DeepAgents worker. Supports execution depth presets, request-specific rubrics, SubAgent delegation, and RubricMiddleware cross-validation.NVIDIA/nemoclaw-community120
- Github Readonly LiveRead the configured live GitHub repository through authenticated, policy-scoped GitHub REST GET requests.NVIDIA/nemoclaw-community120
- Github Readonly LiveRead an allowed live GitHub repository through authenticated, policy-scoped GitHub REST GET requests.NVIDIA/nemoclaw-community120