- Perf Optimization CasebookCasebook of past successful and classic TensorRT-LLM optimizations (runtime/execution and kernel level) recorded as reusable decision precedents. Consult when deciding which optimization to apply for a classified bottleneck or a given config/model/hardware, to find prior art and adapt a proven approach instead of guessing. Each case records applicability signals, mechanism, how to apply, expected effect, accuracy risk, verification, and rollback.NVIDIA/TensorRT-LLM14,764
- Perf OptimizeLaunch and operate this repo's perf-optimize workflow, which iteratively APPLIES TensorRT-LLM serving optimizations — baseline benchmark at one concurrency or a Pareto curve of them (tok/s/user vs tok/s/gpu), analytical SOL projection on by default (via the internal-perf-sol-analysis skill) sizing the headroom the campaign chases, profile-ranked roadmap.yaml (nsys + ncu per-kernel analysis via the perf-nsight-compute-analysis skill), a fixed budget of rounds applying items one at a time gated onNVIDIA/TensorRT-LLM14,764
- Perf Torch Cuda GraphsApply CUDA Graphs to PyTorch workloads — API selection (torch.compile, PyTorch make_graphed_callables, TE make_graphed_callables, MCore CudaGraphManager, FullCudaGraphWrapper, manual torch.cuda.graph), code compatibility, capture workflows, dynamic pattern handling, and troubleshooting. Triggers: CUDA graph, torch.cuda.graph, make_graphed_callables, reduce-overhead, graph capture, graph replay, kernel launch overhead, CudaGraphManager, FullCudaGraphWrapper, full-iteration graph, stream capture.NVIDIA/TensorRT-LLM14,764
- Perf Torch Sync FreeIdentify and eliminate host-device synchronizations in PyTorch code. Detects sync points (.item(), .cpu(), boolean indexing, torch.tensor on CUDA), classifies false vs true dependencies, provides sync-free alternatives. Triggers: sync-free, synchronization, .item(), .cpu(), host-device sync, eliminate syncs, CPU stall, non_blocking, set_sync_debug_mode, cudaStreamSynchronize, cudaEventSynchronize, remove syncs, async GPU.NVIDIA/TensorRT-LLM14,764
- Perf Workload ProfilingCode instrumentation for timing workloads. Two scenarios: (1) Training loop — inject manual timing to report per-iteration latency, throughput (samples/sec), and data load time. (2) Standalone kernel/op — write CUDA event timing code with warmup, per-iteration statistics, and anti-pattern avoidance. Also covers NVTX annotation for labeling profiler timelines. NOT for: running or analyzing profiler tools (nsys, ncu, Nsight Systems, Nsight Compute), writing kernels (Triton, CuTe, CUDA), applying oNVIDIA/TensorRT-LLM14,764
- Trtllm Case ExecutorRun TensorRT-LLM test cases, benchmarks, evaluations, or custom scripts by checking the environment (local GPU or Slurm), selecting the appropriate Docker image, and executing either locally or via Slurm job submission. Accepts pre-built command strings — command construction for trtllm-bench, trtllm-eval, and test_perf_sanity is handled upstream by the caller (e.g. trtllm-test-specialist using build_test_command.py).NVIDIA/TensorRT-LLM14,764
- Trtllm Code ContributionBest practices for contributing code to TensorRT-LLM. Covers the official contribution process (issue tracking, fork workflow, DCO signing), coding guidelines, implementation workflow, common mistakes, testing strategy, commit hygiene, and review readiness. Incorporates rules from CONTRIBUTING.md and CODING_GUIDELINES.md plus lessons distilled from real PR retrospectives. Use when implementing new features, optimizations, or bug fixes in the TensorRT-LLM codebase.NVIDIA/TensorRT-LLM14,764
- Trtllm Codebase ExplorationSystematic approach to exploring the TensorRT-LLM codebase before implementing new features or optimizations. Teaches how to discover existing infrastructure, trace code paths, and avoid reimplementing what already exists. Derived from real mistakes where ~250 lines of code were written and deleted because existing forward methods weren't discovered upfront. Use when starting any new feature, optimization, or code modification in TRT-LLM.NVIDIA/TensorRT-LLM14,764
- Trtllm Flashinfer UpgradeUpgrade flashinfer-python version in TensorRT-LLM. Fetches the latest releases from GitHub (stable and nightly), compares with the current pinned version, lets the user pick a target version, and updates all version references across the repo. Use when the user wants to bump or upgrade flashinfer.NVIDIA/TensorRT-LLM14,764
- Trtllm Model Onboard MultimodalOnboard a HuggingFace multimodal model (vision/audio/video + text) to the TensorRT-LLM PyTorch backend. Use when writing a new `tensorrt_llm/_torch/models/modeling_<vlm>.py` plus its input processor and weight mapper, or extending an existing VLM.NVIDIA/TensorRT-LLM14,764
- Trtllm Moe DevelopReview, design, and refactor TensorRT-LLM PyTorch MoE code for architecture fit, clean code, maintainability, and testability. Always use for any modification, review, refactor, or design planning that touches MoE modules, including tensorrt_llm/_torch/moe/fused_moe, ConfigurableMoE, MoE backends, MoEScheduler/moe_scheduler.py, forward execution/chunking, communication strategies, EPLB, quantization/weight handling, routing, factories, MoE docs, or MoE tests. Also use when the user asks whether NVIDIA/TensorRT-LLM14,764
- Trtllm Serve Config GuideGenerate a source-backed starting `trtllm-serve --config` YAML for basic aggregate single-node PyTorch serving, aligned with checked-in TensorRT-LLM configs and deployment docs. Preserves explicit latency / balanced / throughput objectives. Excludes disaggregated, multi-node, and non-MTP speculative configs.NVIDIA/TensorRT-LLM14,764
- Trtllm Test Script BuilderBuild Slurm scripts or Docker commands for TensorRT-LLM workloads. Resolves all parameters (docker image, mounts, parallelism, MPI mode), generates the complete script from Category templates, and writes both the script and a job_spec.json manifest to the work directory.NVIDIA/TensorRT-LLM14,764
- Trtllm Test SpecialistRuns model-level and module-level tests for TensorRT-LLM. First classifies the test scope (module test or model test), then dispatches to the appropriate workflow. Model tests are further classified by type (functionality/smoke test, benchmark, or evaluation). Prompts the user for parallelism parameters (tp, ep, dp), dataset paths, device type, or a config file as needed. All test execution is delegated to trtllm-case-executor.NVIDIA/TensorRT-LLM14,764
- Visual Gen Component TestDesign, implement, calibrate, and run deterministic TensorRT-LLM VisualGen L1 operation tests and L2 component-by-feature parity tests. Use for DiT, VAE, text-encoder, scheduler, parallelism, quantization, caching, compile, or model-specific component correctness; not for perceptual model evaluation or full-pipeline LPIPS quality gates.NVIDIA/TensorRT-LLM14,764
- Build From IssuePlan and implement work described in a GitHub issue, including verification, documentation, and a PR that closes the issue.NVIDIA/OpenShell14,461
- Build Openshell Mxc WindowsMaintain and validate OpenShell's build-only Windows MSVC lane for x64 and ARM64. Use when working on Windows compilation, `windows:*` mise tasks, unsupported Windows compute-driver contracts, or Windows build reports. This skill does not implement Docker, Kubernetes, Podman, VM, MXC driver, policy translation, MSI, service, or supervisor runtime support on Windows.NVIDIA/OpenShell14,461
- Create Github IssueCreate GitHub issues using the gh CLI. Use when the user wants to create a new issue, report a bug, request a feature, or create a task in GitHub. Trigger keywords - create issue, new issue, file bug, report bug, feature request, github issue.NVIDIA/OpenShell14,461
- Create Github PrCreate GitHub pull requests using the gh CLI. Use when the user wants to create a new PR, submit code for review, or open a pull request. Trigger keywords - create PR, pull request, new PR, submit for review, code review.NVIDIA/OpenShell14,461
- Create RfcCreate OpenShell RFC proposals in rfc/ from a design request. Use when the user asks to write, draft, start, create, or update an RFC, Request for Comments, architecture proposal, API proposal, process proposal, or cross-cutting design proposal that should follow the OpenShell RFC process and template.NVIDIA/OpenShell14,461
- Create SpikeInvestigate an OpenShell problem and create a structured issue with technical findings for human disposition.NVIDIA/OpenShell14,461
- Debug InferenceDebug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM. Use for provider attachment, endpoint policy, credential substitution, topology, and migration from the removed managed inference endpoint. Trigger keywords - debug inference, managed inference endpoint, local inference, ollama, lm studio, vllm, sglang, trtllm, NIM, inference failing, model server unreachable, credential_endpoint_miNVIDIA/OpenShell14,461
- Debug Openshell ClusterDebug why an OpenShell gateway deployment is unhealthy, unreachable, or unable to create sandboxes. Use for gateway health failures, Docker/Podman runtime issues, Helm failures, Kubernetes scheduling, TLS or auth, gateway interceptors, supervisor middleware startup or runtime failures, external compute-driver sockets, VM drivers, or sandbox startup. Trigger keywords - debug gateway, gateway failing, deployment failing, helm install failing, cluster health, gateway health, gateway not starting, hNVIDIA/OpenShell14,461
- Fix Security IssueImplement an authorized fix for a reviewed security issue and open a PR that closes its issue.NVIDIA/OpenShell14,461
- Gator GateValidate and monitor OpenShell GitHub issues and PRs using the gator:* state machine. Use when asked to triage issues/PRs for project validity, gate PRs, run gator, validate submissions, or monitor PRs toward merge readiness.NVIDIA/OpenShell14,461
- Generate Sandbox PolicyGenerate sandbox security policies from plain-language requirements and optional REST API documentation. Produces L4 or fine-grained L7 network policies and ordered network middleware configuration. Use for API access rules, middleware host selection, failure behavior, or built-in and operator-run middleware attachment. Trigger keywords - generate policy, create policy, update policy, change policy, sandbox policy, network policy, API policy, security policy, allow API, restrict API, network midNVIDIA/OpenShell14,461
- Helm Dev EnvironmentStart up, tear down, and configure the local Kubernetes development environment for OpenShell. Uses k3d (Docker-backed k3s) + Skaffold + Helm. Covers cluster lifecycle, optional add-ons (Keycloak OIDC, Envoy Gateway), HA testing, and port mappings. Trigger keywords - local k8s, local cluster, k3d, skaffold, helm dev, start cluster, stop cluster, tear down cluster, delete cluster, create cluster, helm:k3s, helm:skaffold, local dev environment, dev cluster, k8s dev, envoy gateway local, keycloak lNVIDIA/OpenShell14,461
- Launch Openshell GatorLaunch and supervise OpenShell gator agents. Use when starting gator on issues or PRs, checking gator sandboxes, building the gator sandbox image, restarting stuck gators, inspecting gator logs, or experimenting with gator harness/model overrides. Trigger keywords - launch gator, start gator, run gator, gator sandbox, supervised gator, gator logs, restart gator.NVIDIA/OpenShell14,461
- Openshell CliGuide agents through using the OpenShell CLI (openshell) for sandbox management, gateway registration, provider configuration and refresh, profile management, policy iteration, settings, service exposure, BYOC workflows, and attached-provider inference. Covers basic through advanced multi-step workflows. Trigger keywords - openshell, sandbox create, sandbox exec, sandbox connect, logs, provider create, profile list, profile describe, provider refresh, policy set, policy get, settings, service exNVIDIA/OpenShell14,461
- Openshell Policy AdvisorUse when an OpenShell sandbox returns policy_denied, mentions policy.local, or needs a narrow network policy proposal.NVIDIA/OpenShell14,461
- Review Github PrReview a GitHub pull request by summarizing its diff and key design decisions. Use when the user wants to review a PR, understand changes in a branch, or get a code review summary. Trigger keywords - review PR, review pull request, summarize PR, summarize diff, code review, review branch, PR summary, diff summary.NVIDIA/OpenShell14,461
- Review Security IssueReview an authorized security issue for validity, severity, and a remediation plan.NVIDIA/OpenShell14,461
- SbomGenerate and manage Software Bill of Materials (SBOMs) for the OpenShell project. Covers SBOM generation with Syft, license resolution via public registries, and CSV export for compliance review. Trigger keywords - SBOM, sbom, bill of materials, license audit, license resolution, generate sbom, sbom csv, dependency license, supply chain, license scan.NVIDIA/OpenShell14,461
- Sync Agent InfraReconcile contributor skills, AGENTS.md, CONTRIBUTING.md, issue and PR templates, and workflow references after repository workflow changes.NVIDIA/OpenShell14,461
- Test Release CanaryManually dispatch and iterate on the Release Canary workflow that smoke-tests published OpenShell artifacts (install.sh on macOS/Ubuntu/Fedora, Helm chart on kind) after each Release Dev publish. Use when changing `.github/workflows/release-canary.yml`, validating a release before tagging, debugging a canary failure, or reproducing a canary job locally. Trigger keywords - release canary, release-canary, canary failed, canary dispatch, test release canary, post-release smoke, install.sh canary, hNVIDIA/OpenShell14,461
- Triage IssueAssess community issues, validate technical claims, request missing evidence, and prepare valid work for human disposition.NVIDIA/OpenShell14,461
- Tui DevelopmentGuide for developing the OpenShell TUI — a ratatui-based terminal UI for the OpenShell platform. Covers architecture, navigation, data fetching, theming, UX conventions, and development workflow. Trigger keywords - term, TUI, terminal UI, ratatui, openshell-tui, tui development, tui feature, tui bug.NVIDIA/OpenShell14,461
- Update Docs From CommitsScan recent git commits for changes that affect user-facing behavior, then draft or update the corresponding documentation pages. Use when docs have fallen behind code changes, after a batch of features lands, or when preparing a release. Trigger keywords - update docs, draft docs, docs from commits, sync docs, catch up docs, doc debt, docs behind, docs drift.NVIDIA/OpenShell14,461
- Watch Github ActionsWatch and monitor GitHub Actions workflow runs using the gh CLI. Use when the user wants to check workflow status, watch a running workflow, view CI/CD jobs, or monitor build progress. Trigger keywords - watch pipeline, pipeline status, CI status, check build, monitor CI, view pipeline, pipeline progress, workflow status, actions status.NVIDIA/OpenShell14,461
- Deprecate ApiMark a Polygraphy function, class, module, or alias as deprecated so it warns at runtime and is scheduled for removal. Use when asked to deprecate an API, replace one API with another while keeping backwards compatibility, or add a deprecation warning.NVIDIA/TensorRT13,363
- Headless ScreenshotsUse this skill when asked to take browser screenshots of web pages or web-based tools in a headless/automated way — especially when the page uses HTML canvas (e.g. Cytoscape.js, WebGL, Chart.js) and the screenshots need to show canvas-rendered content like graph nodes, edge labels, or drawn shapes. Also use when asked to automate multi-step browser interactions (clicks, drags, form fills) before screenshotting.NVIDIA/TensorRT13,363
- Release PolygraphyPrepare a Polygraphy release by auditing all changes since the previous release, finalizing or adding the versioned CHANGELOG section, updating polygraphy/__init__.py, validating the release-only diff, and opening a GitLab merge request targeting develop. Use when asked to create, prepare, cut, or publish a new Polygraphy version or release MR.NVIDIA/TensorRT13,363
- Trt Cpp Runtime QuickstartLoad and run a TensorRT engine (.plan / .engine) from C++ using the TensorRT 11 / 10.x **modern Runtime API**, avoiding the deprecated TRT 8.x binding-index APIs that older guidance still promotes. Use whenever the user asks about loading or running a TensorRT .plan/.engine from C++, even on "minimal example" requests — without this skill the default reply uses deprecated enqueueV2-style code. Also use when the user hits "Engine plan file is generated on an incompatible device", deserializeCudaENVIDIA/TensorRT13,363
- Trt Onnx QuickstartBuild and verify a TensorRT engine from a Hugging Face model ID or ONNX file, with numerical parity checked against ONNX Runtime. Use when the user imports a non-LLM model to TensorRT, needs a verified engine from ONNX, hits trtexec "unsupported operator", must verify the engine matches ONNX numerically, debugs a polygraphy parity failure (large max abs diff at FP16), or configures multi-input dynamic shapes. Triggers: convert ONNX to TensorRT, Hugging Face to TensorRT, trtexec onnx, trtexec unsNVIDIA/TensorRT13,363
- Trt Perf AnalysisGenerate an interactive HTML performance report from existing TensorRT (TRT) layer-info and profile/latency JSON pairs, commonly named `layers_*.json` and `profile_*.json`. Use when the user has such dumps and wants to analyze or diagnose inference performance, including hot layers, per-layer-type latency breakdowns, fusion and other optimization opportunities, or comparison across engines, builds, or configurations.NVIDIA/TensorRT13,363
- Trt Strong Typing MigrationMigrate a TensorRT build from weak typing (deprecated 10.12, removed 11.0) to strong typing — across Python INetworkDefinition builders, the trtexec CLI, and C++ builder code. Use when a TRT 11 upgrade breaks a weakly-typed build. Triggers: weakly typed to strongly typed, kSTRONGLY_TYPED, weak typing deprecated, kFP16/kINT8 removed, setPrecision rejected, setComputePrecision deprecated, do I still need --stronglyTyped, how to add the kSTRONGLY_TYPED flag, ModelOpt autocast, INT8 on TRT 11. NOT fNVIDIA/TensorRT13,363
- Trt Torch QuickstartCompile a PyTorch model to a TensorRT engine via Torch-TensorRT — AOT or JIT — under the new strong-typing default. Use when the user compiles PyTorch to TensorRT without ONNX, hits "enabled_precisions should not be used when use_explicit_typing=True", sees Dynamo graph breaks or PyTorch fallback, debugs ABI errors at import torch_tensorrt, or needs the compatible torch / torch_tensorrt / tensorrt-cu13 version pins for TensorRT 11. Triggers: torch_tensorrt, torch_tensorrt.dynamo.compile, torch.cNVIDIA/TensorRT13,363
- Accelerated Computing CudfOfficial NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.NVIDIA/cudf9,771
- Build Test CudfUse this skill to build and test code changes inside a cudf devcontainer.NVIDIA/cudf9,771
- Build Test Cudf JavaBuild and test cudf Java bindings (cudf-java) inside a cudf devcontainer. Use when the user asks to build, compile, or test Java code in the cudf repository.NVIDIA/cudf9,771
- Cudf Query Engine BuilderBuilds a native cuDF proof of concept for a query engine with no working GPU path. Use when an engineer supplies target-engine or adapter code plus a query, operator sequence, or CPU implementation selected for the POC and needs working libcudf C++ or cuDF Java code, a runnable single-GPU program, and a CPU/reference correctness comparison. Do not use for standalone dataframe analysis, cuDF mainline changes, production or deployment integration, production fallback or I/O policy beyond the singNVIDIA/cudf9,771
- Debug Cudf PandasDebug and fix pandas test suite failures under the cudf.pandas compatibility layer. Use when given pytest node IDs of failing pandas tests that need to be fixed for cudf.pandas compatibility.NVIDIA/cudf9,771
- Perf Compare CudfBenchmark a cuDF branch, WIP changes, or a PR against the `main` branchNVIDIA/cudf9,771
- Reproduce CiReproduce cudf CI failures locally. Provide a GitHub Actions job URL to auto-discover parameters, or supply the container image and CI script directly.NVIDIA/cudf9,771
- Review CudfUse this skill to review GitHub pull requests for cudfNVIDIA/cudf9,771
- Review Cudf Polars ExpressionsUse when implementing or reviewing support for Polars Expressions in cudf-polarsNVIDIA/cudf9,771
- Update Narwhals TestingUse when updating Narwhals for cudf-polars and cudf testing in CINVIDIA/cudf9,771
- Warp Changelog AuditUse when auditing and recovering Warp changelog fragments, finalizing a release changelog, or synchronizing a tagged release back to main.NVIDIA/warp7,166
- Warp Changelog AuditUse when auditing and recovering Warp changelog fragments, finalizing a release changelog, or synchronizing a tagged release back to main.NVIDIA/warp7,166
- Warp Closing IssueUse when the user provides Warp commit SHA(s) and GitHub issue number(s) to assess, draft issue comments, post progress updates, or recommend whether issue threads should stay open or close.NVIDIA/warp7,166