Trt Perf Analysis
NVIDIA/TensorRT/.agents/skills/trt-perf-analysisai-mlOfficial
Official Provider SkillView repo
Generate an interactive HTML performance report from existing TensorRT (TRT) layer-info and profile/latency JSON pairs, commonly named `layers_*.json` and `profile_*.json`. Use when the user has such dumps and wants to analyze or diagnose inference performance, including hot layers, per-layer-type latency breakdowns, fusion and other optimization opportunities, or comparison across engines, builds, or configurations.
Files18 files
SKILL.md58 lines
Loading editor…
Install
RecommendedOne command — your agent picks it up automatically.
Select an AI agent above to see the install command.
or
Manual Install
More stepsDownload the archive and add the files to your project manually.
Skill details
Versionv1.0.0
AuthorNVIDIA
Categoryai-ml
Skill IDNVIDIA/TensorRT/.agents/skills/trt-perf-analysis
Files18 files
Related skills
Deprecate ApiMark a Polygraphy function, class, module, or alias as deprecated so it warns at runtime and is scheduled for removal. Use when asked to deprecate an API, replace one API with another while keeping backwards compatibility, or add a deprecation warning.Headless ScreenshotsUse this skill when asked to take browser screenshots of web pages or web-based tools in a headless/automated way — especially when the page uses HTML canvas (e.g. Cytoscape.js, WebGL, Chart.js) and the screenshots need to show canvas-rendered content like graph nodes, edge labels, or drawn shapes. Also use when asked to automate multi-step browser interactions (clicks, drags, form fills) before screenshotting.Release PolygraphyPrepare a Polygraphy release by auditing all changes since the previous release, finalizing or adding the versioned CHANGELOG section, updating polygraphy/__init__.py, validating the release-only diff, and opening a GitLab merge request targeting develop. Use when asked to create, prepare, cut, or publish a new Polygraphy version or release MR.Trt Cpp Runtime QuickstartLoad and run a TensorRT engine (.plan / .engine) from C++ using the TensorRT 11 / 10.x **modern Runtime API**, avoiding the deprecated TRT 8.x binding-index APIs that older guidance still promotes. Use whenever the user asks about loading or running a TensorRT .plan/.engine from C++, even on "minimal example" requests — without this skill the default reply uses deprecated enqueueV2-style code. Also use when the user hits "Engine plan file is generated on an incompatible device", deserializeCudaETrt Onnx QuickstartBuild and verify a TensorRT engine from a Hugging Face model ID or ONNX file, with numerical parity checked against ONNX Runtime. Use when the user imports a non-LLM model to TensorRT, needs a verified engine from ONNX, hits trtexec "unsupported operator", must verify the engine matches ONNX numerically, debugs a polygraphy parity failure (large max abs diff at FP16), or configures multi-input dynamic shapes. Triggers: convert ONNX to TensorRT, Hugging Face to TensorRT, trtexec onnx, trtexec unsTrt Strong Typing MigrationMigrate a TensorRT build from weak typing (deprecated 10.12, removed 11.0) to strong typing — across Python INetworkDefinition builders, the trtexec CLI, and C++ builder code. Use when a TRT 11 upgrade breaks a weakly-typed build. Triggers: weakly typed to strongly typed, kSTRONGLY_TYPED, weak typing deprecated, kFP16/kINT8 removed, setPrecision rejected, setComputePrecision deprecated, do I still need --stronglyTyped, how to add the kSTRONGLY_TYPED flag, ModelOpt autocast, INT8 on TRT 11. NOT f