Official Provider SkillView repo
Launch and operate this repo's perf-optimize workflow, which iteratively APPLIES TensorRT-LLM serving optimizations — baseline benchmark at one concurrency or a Pareto curve of them (tok/s/user vs tok/s/gpu), analytical SOL projection on by default (via the internal-perf-sol-analysis skill) sizing the headroom the campaign chases, profile-ranked roadmap.yaml (nsys + ncu per-kernel analysis via the perf-nsight-compute-analysis skill), a fixed budget of rounds applying items one at a time gated on
Files1 files
SKILL.md366 lines
Loading editor…
Install
RecommendedOne command — your agent picks it up automatically.
Select an AI agent above to see the install command.
or
Manual Install
More stepsDownload the file and paste it into your agent's system prompt.
Skill details
Versionv1.0.0
AuthorNVIDIA
Categoryai-ml
Skill IDNVIDIA/TensorRT-LLM/agent-flow/.claude/skills/perf-optimize
Related skills
Exec Env CheckCheck the local execution environment for GPU availability, Docker support, and Slurm access. Returns the execution scenario (`satisfied, local, docker`, `satisfied, local, direct`, `satisfied, slurm, local`, or `not_satisfied`), the number of available GPUs, and the GPU type. On Slurm login nodes without local GPUs, the cluster is identified by delegating the hostname to internal-env-info (hostname-based mode), which owns the hostname → cluster_name patterns; GPU type and gpus_per_node then comExec Local CompileCompile TensorRT-LLM on a compute node inside a Docker container. Use this when already on a compute node with GPUs visible.Exec Local DockerExecute a TensorRT-LLM workload locally in Docker. Runs a fully-resolved Docker command in background, monitors completion, reads logs, and reports results. Workflow-agnostic — does not need to know if the workload is pytest, eval, benchmark, or a custom script.Exec Local SlurmSubmit and monitor a Slurm job on a local cluster. Supports two modes: (1) Persistent allocation (default) — allocates nodes once via nohup salloc, imports the container once, installs once, and reuses across runs by setting SLURM env vars and running the sbatch script via bash. (2) One-shot sbatch — submits a fully-generated Slurm script via sbatch, polls job status, reads logs on completion, and reports results. Workflow-agnostic — handles pytest, eval, benchmark, and custom scripts identicallExec Remote SlurmRemote SLURM cluster development via SSH. Use when running jobs, profiling, or developing on a remote SLURM cluster with pyxis/enroot containers. Covers SSH connection management, srun/sbatch/salloc job patterns, tmux-based allocation persistence, file transfer, and safe remote file access. Works with any SLURM cluster accessible via SSH.Exec Slurm CompileCompile TensorRT-LLM on a SLURM cluster. Covers submitting a batch job with a container image, monitoring the job, and verifying the build. Use when the user wants to compile TRT-LLM remotely via SLURM rather than on a local compute node.