- Nemoclaw Contributor Create PrPublish or update a NemoClaw pull request and follow its CI and automated reviews to completion.NVIDIA/NemoClaw22,573
- Nemoclaw Contributor Implement IssueImplement an accepted NemoClaw issue or repair a classified PR finding, with focused validation. Use for requested code or test changes.NVIDIA/NemoClaw22,573
- Nemoclaw Contributor OnboardSet up, repair, or check a NemoClaw contributor checkout. Also handles requested development CLI exposure, runtime onboarding, and pinned-agent launch.NVIDIA/NemoClaw22,573
- Nemoclaw Contributor Plan IssuePlan or divide a named NemoClaw issue into independently useful changes with acceptance evidence. Use for planning requests before implementation.NVIDIA/NemoClaw22,573
- Nemoclaw Contributor Update DependenciesAudit and implement a NemoClaw dependency version upgrade, including Hermes and base images. Use when the dependency itself is changing.NVIDIA/NemoClaw22,573
- Nemoclaw Contributor Update DocsUpdate NemoClaw documentation for merged behavior changes. Use for post-merge catch-up or a requested documentation drift audit.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer Analyze Ci PerformanceAnalyze retained NemoClaw CI timings for slow CLI tests, runner queues, or base-image publication.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer Analyze Pr Value StreamAnalyze one NemoClaw PR lifetime to identify contributor, review, and automation delays or compare it with a latency target.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer Classify Ci FailureClassify one failed NemoClaw GitHub Actions job using bounded, redacted logs and optional retained artifacts.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer Cross Issue SweepFind open issues that a NemoClaw PR may also fix or conflict with. Use when related-issue analysis is requested.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer Cut Release TagPrepare and cut one signed NemoClaw semver release tag, then follow release workflows and draft the Announcement.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer DayRun a NemoClaw daytime maintainer pass over release-targeted work. Use for the maintainer queue or a requested recurring pass.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer E2eRun or inspect NemoClaw live E2E evidence. Routes requested local execution, trusted GitHub dispatch, and release-evidence inspection.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer EveningComplete the NemoClaw end-of-day documentation and release handoff. Cut a release tag only when requested.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer Find Review PrFind high-priority open NemoClaw security PRs to review, including competing or superseded candidates.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer Fix E2e FailuresContinuously maintain automatic NemoClaw main E2E results through coordinated repairs. Use for ongoing maintenance, not one-time dispatch or diagnosis.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer MorningPrepare the NemoClaw morning maintainer plan: triage the backlog, select a target version, and identify release candidates and stragglers.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer Normalize Title TagsRemove bracketed NemoClaw tags from GitHub issue and PR titles. Use for a requested title cleanup.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer PoliciesAnswer NemoClaw maintainer policy questions about triage, labels, Project fields, releases, or competing contributions. Read-only.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer Pr ComparatorCompare competing NemoClaw PRs for one issue and recommend a merge or salvage candidate from review evidence.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer Refactor DocsReorganize NemoClaw documentation pages, navigation, or content ownership while preserving published routes. Use for structural documentation refactors.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer Release NotesDraft a post-tag NemoClaw Announcement from the verified release range and shipped PRs. Use when summarizing a release.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer Runtime ProviderImplement or review a native managed NemoClaw runtime provider and its activation or qualification. Excludes the portable experimental profile.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer Security Code ReviewPerform a requested security review of a NemoClaw PR or a PR linked to an issue. Use for vulnerability or trust-boundary assessment.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer TriagePropose and apply authorized Issue Type, Project fields, and labels for NemoClaw issues or PRs, individually or in a batch.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer Validate LaunchableValidate the staging NemoClaw Brev Launchable through its web journey or a deployed environment. Use for advisory validation outside automated E2E.NVIDIA/NemoClaw22,573
- Nemoclaw Maintainer Verify StaleReproduce stale NemoClaw bug reports on the reported and newest releases, then propose evidence-backed triage. Never auto-closes issues.NVIDIA/NemoClaw22,573
- Nemoclaw Skills GuideFind the repository skill for a NemoClaw task or browse the skill catalog. Use when skill selection needs help.NVIDIA/NemoClaw22,573
- Nemoclaw User GuideFind official NemoClaw documentation for installation, configuration, operation, or troubleshooting. Also handles requested docs MCP setup.NVIDIA/NemoClaw22,573
- Nemoclaw User GuideFind official NemoClaw documentation for installation, configuration, operation, or troubleshooting. Also handles requested docs MCP setup.NVIDIA/NemoClaw22,573
- Mcore Build And DependencyContainer-based dev environment setup and dependency management for Megatron-LM. Covers acquiring and launching the CI container, uv package management, and updating uv.lock.NVIDIA/Megatron-LM18,042
- Mcore Bump Base ImageBump the NVIDIA PyTorch base image (`nvcr.io/nvidia/pytorch:YY.MM-py3`) used by Megatron-LM CI. Covers the two pin sites (GitHub CI in `docker/.ngc_version.dev` and GitLab CI in `.gitlab/stages/01.build.yml`), the post-bump CI loop (re-run functional tests, refresh golden values, mark broken tests), and the gotchas that bit PRsNVIDIA/Megatron-LM18,042
- Mcore CicdCI/CD reference for Megatron-LM. Covers CI pipeline structure, PR scope labels, triggering internal GitLab CI (which force-pushes the current branch to a pull-request/BRANCH ref — always dry-run and verify the destination first; never run against shared or protected branches), and CI failure investigation.NVIDIA/Megatron-LM18,042
- Mcore Create IssueInvestigate a failing GitHub Actions run or job and create a GitHub issue for the failure.NVIDIA/Megatron-LM18,042
- Mcore Linting And FormattingLinting and formatting for Megatron-LM. Covers running autoformat.sh, tools (ruff, black, isort, pylint, mypy), and code style rules.NVIDIA/Megatron-LM18,042
- Mcore Migrate Gpt To HybridMigration guide for moving Megatron Core GPTModel checkpoints, model providers, training commands, and layer mappings to HybridModel, including the mechanical steps for transferring an existing pretrain_gpt.py launch script.NVIDIA/Megatron-LM18,042
- Mcore Onboard Gb200 1node TestsOnboard 1-node GitHub MR functional tests for GB200 from existing mr-scoped 2-node tests.NVIDIA/Megatron-LM18,042
- Mcore Run On SlurmHow to launch distributed Megatron-LM training jobs on a SLURM cluster. Covers a minimal sbatch skeleton, environment-variable setup for torch.distributed.run, CUDA_DEVICE_MAX_CONNECTIONS rules across hardware and parallelism modes, container conventions, monitoring, and per-rank failure diagnosis.NVIDIA/Megatron-LM18,042
- Mcore Split PrSplit a PR into multiple PRs to reduce the number of required CODEOWNERS reviewer groups.NVIDIA/Megatron-LM18,042
- Mcore TestingTest system for Megatron-LM. Covers test layout, recipe YAML structure, adding and running unit and functional tests, golden values, marker filters, and CI parity.NVIDIA/Megatron-LM18,042
- Nightly SyncDomain knowledge for the nightly main-to-dev sync workflow. Covers merge strategy, CI architecture, failure investigation, and known issues.NVIDIA/Megatron-LM18,042
- Pr ReviewReview rubric for the `/review` pull-request command. The formal reviewer reads it as a file and it is not an interactive skill — do not load it to answer questions or to review code outside that command.NVIDIA/Megatron-LM18,042
- Respond To IssueResearch and draft a response to a GitHub issue or question from an external contributor.NVIDIA/Megatron-LM18,042
- Update Golden ValuesRefresh golden values from a GitHub Actions workflow run (failing-only or all jobs), calculate signed per-model percentage changes, and produce a PR-ready summary. Use when the user asks to update goldens for a CI run, refresh golden values from a workflow ID, or generate a golden-value diff summary for a PR description.NVIDIA/Megatron-LM18,042
- Exec Env CheckCheck the local execution environment for GPU availability, Docker support, and Slurm access. Returns the execution scenario (`satisfied, local, docker`, `satisfied, local, direct`, `satisfied, slurm, local`, or `not_satisfied`), the number of available GPUs, and the GPU type. On Slurm login nodes without local GPUs, the cluster is identified by delegating the hostname to internal-env-info (hostname-based mode), which owns the hostname → cluster_name patterns; GPU type and gpus_per_node then comNVIDIA/TensorRT-LLM14,764
- Exec Local CompileCompile TensorRT-LLM on a compute node inside a Docker container. Use this when already on a compute node with GPUs visible.NVIDIA/TensorRT-LLM14,764
- Exec Local DockerExecute a TensorRT-LLM workload locally in Docker. Runs a fully-resolved Docker command in background, monitors completion, reads logs, and reports results. Workflow-agnostic — does not need to know if the workload is pytest, eval, benchmark, or a custom script.NVIDIA/TensorRT-LLM14,764
- Exec Local SlurmSubmit and monitor a Slurm job on a local cluster. Supports two modes: (1) Persistent allocation (default) — allocates nodes once via nohup salloc, imports the container once, installs once, and reuses across runs by setting SLURM env vars and running the sbatch script via bash. (2) One-shot sbatch — submits a fully-generated Slurm script via sbatch, polls job status, reads logs on completion, and reports results. Workflow-agnostic — handles pytest, eval, benchmark, and custom scripts identicallNVIDIA/TensorRT-LLM14,764
- Exec Remote SlurmRemote SLURM cluster development via SSH. Use when running jobs, profiling, or developing on a remote SLURM cluster with pyxis/enroot containers. Covers SSH connection management, srun/sbatch/salloc job patterns, tmux-based allocation persistence, file transfer, and safe remote file access. Works with any SLURM cluster accessible via SSH.NVIDIA/TensorRT-LLM14,764
- Exec Slurm CompileCompile TensorRT-LLM on a SLURM cluster. Covers submitting a batch job with a container image, monitoring the job, and verifying the build. Use when the user wants to compile TRT-LLM remotely via SLURM rather than on a local compute node.NVIDIA/TensorRT-LLM14,764
- Kernel Cute WritingWrite and implement GPU kernels using NVIDIA CuTe DSL (CUTLASS 4.x Python API) — NOT for Triton, CUDA C++, or conceptual explanations. Trigger only when the user wants to write or implement a kernel, not when asking questions about CuTe DSL concepts or layouts. CuTe DSL uses cute.jit/cute.kernel decorators and cutlass.cute imports. Covers element-wise kernels, GEMM patterns, reductions, memory hierarchy (global/shared/register/TMA), MMA tensor core operations, software pipelining, and framework NVIDIA/TensorRT-LLM14,764
- Kernel Tileir OptimizationOptimize existing Triton kernels for NVIDIA TileIR backend on Blackwell GPUs (sm_100+). Adds TileIR-specific autotune configs: occupancy, num_ctas, TMA descriptors. Covers kernel classification (dot-related, norm-like, elementwise, reduction), type-specific transformations, and PTX-vs-TileIR benchmarking. Triggered by: "optimize for TileIR", "add TileIR configs", "Blackwell optimization", "TMA descriptors", "2CTA mode", "occupancy tuning". Kernels use standard `import triton`; TileIR activates vNVIDIA/TensorRT-LLM14,764
- Kernel Triton WritingONLY for OpenAI Triton (@triton.jit) kernel development. NEVER use for CUDA C++ kernels, TileIR, or profiling tools (ncu, nsys). The user's request must involve Triton explicitly. Covers Triton-specific patterns: fused elementwise, reductions (softmax, LayerNorm, RMSNorm), tiled GEMM with triton.autotune, and flash attention. Workflow: design, write, verify (with fast-path for explicit requests).NVIDIA/TensorRT-LLM14,764
- Perf AnalysisPerformance analysis coordination workflow. Guides profiling delegation, bottleneck classification (compute/memory/launch/communication/sync), and structured report generation. Use when the user asks to analyze performance, profile a workload, check MFU/SOL, or diagnose bottlenecks.NVIDIA/TensorRT-LLM14,764
- Perf AnalyzeLaunch and operate this repo's perf-analyze workflow, which DIAGNOSES a TensorRT-LLM serving deployment without applying changes — benchmark at one concurrency or a Pareto curve of them (tok/s/user vs tok/s/gpu), analytical SOL projection on by default (via the internal-perf-sol-analysis skill), nsys + ncu per-kernel deep dive (via the perf-nsight-compute-analysis skill), and a report naming the single dominant bottleneck. Use when the user wants to analyze / profile / diagnose trtllm-serve perfNVIDIA/TensorRT-LLM14,764
- Perf Host AnalysisAnalyze host/CPU overhead in TensorRT-LLM inference from nsys traces. Detect whether host overhead is the bottleneck using GPU idle ratio, host prep exposed ratio, and per-phase evidence. For regressions, isolate forward steps via allreduce/NVTX patterns, compare host operation breakdowns across versions, and identify scheduling or request-management overhead. Supports optional inter-kernel gap, eager-vs-graph, pattern mapping, and multi-rank straggler drill-down. Use standalone or within perf-aNVIDIA/TensorRT-LLM14,764
- Perf Host OptimizationProfiles and optimizes TensorRT-LLM host/CPU overhead using line_profiler (with nsys support planned). Runs iterative profile-analyze-optimize-validate rounds. Use when GPU utilization is low or optimizing PyExecutor throughput.NVIDIA/TensorRT-LLM14,764
- Perf Nsight Compute AnalysisAnalyze ncu (NVIDIA Nsight Compute) profiling output: SOL% bottleneck classification, roofline analysis, occupancy diagnosis, memory hierarchy analysis, warp stall analysis, metric interpretation, and programmatic .ncu-rep report analysis. NOT for kernel writing or code generation, Nsight Systems (nsys), host-side profiling, or system-level profiling.NVIDIA/TensorRT-LLM14,764
- Perf Nsight SystemsNsight Systems (nsys) CLI for system-level timeline profiling. Use when the user wants to run nsys profile, analyze .nsys-rep reports, use nsys stats/analyze/recipe commands, diagnose GPU idle time from timeline traces, or profile distributed training with NCCL overlap analysis. NOT for kernel-level metrics like SOL%, occupancy, or roofline (use perf-nsight-compute-analysis for ncu). NOT for writing or generating kernels. NOT for applying optimizations like CUDA Graphs.NVIDIA/TensorRT-LLM14,764
- Perf OptimizationPerformance optimization coordination playbook. Contains specialist routing table, TileIR two-step pipeline, kernel generation specialist selection, prioritization criteria, and safe modification workflow. Use when the user asks to apply optimizations, write kernels, or improve performance. Covers both user-specified optimization and autopilot-driven iterative optimization.NVIDIA/TensorRT-LLM14,764