devopsResmi
Resmi Sağlayıcı Skill'iView repo
Research NVIDIA DGX, SuperPOD, GB200, NVLink, BlueField, Base Command Manager (BCM), and on-prem infrastructure documentation through the IRA/Sequoia retrieval agent. Use when the user asks to research docs, deployment policy, operational guidance, or newly uploaded knowledge-base documents.
Dosyalar1 dosya
SKILL.md30 satır
Loading editor…
Kurulum
ÖnerilenTek komut — ajanınız otomatik olarak devreye alır.
Kurulum komutunu görmek için yukarıdan bir AI aracı seçin.
veya
Manuel Kurulum
Daha fazla adımDosyayı indirin ve ajanınızın sistem istemine yapıştırın.
Skill detayları
Versiyonv1.0.0
YazarNVIDIA
Kategoridevops
Skill IDNVIDIA/AI-Factory-Operations-Agent/helm/mosaic-stack/files/openclaw-seed/skills/iraop
İlgili skill'ler
Ai Factory Operations Agent HeadlessUse AI Factory Operations Agent without opening the browser UI. Invoke it through the headless HTTP API, packaged CLI, or stdio MCP server, and deploy it first when no service is available.BcmInspect Base Command Manager (BCM) inventory and health evidence, and make explicitly requested changes when Base Command Manager edit mode is enabled.GrafanaCreate and open Grafana dashboards for metrics, Loki logs, and mixed panels. Use when the user asks for a dashboard, visualization, chart, panel, or graph of GPU, Base Command Manager (BCM), node, network, InfiniBand, Prometheus metrics, or Loki logs.Hardware AgentTriage hardware faults on NVIDIA GPU compute nodes through the Hardware Agent. Starting a triage is read-only log collection and hardware_analyze_dut is available in View mode; never require Edit or Auto. Before asking to start any fresh NVDebug collection, you MUST say: 'A fresh NVDebug collection can take 5-30 minutes. Do you want me to start it?' Stop and wait for the next user response. An ordinary affirmative such as yes, let's go, do it, or proceed is sufficient; start immediately without ObservabilityAnalyze cluster observability data and create Grafana dashboards. Use when the user asks about metrics, logs, log labels, cluster health history, dashboards, charts, panels, or visualizations.SlurmInspect mounted Slurm accounting exports and job logs for failed workload root cause analysis.