ProfessorGPTProfessorGPT
Resmi Sağlayıcı Skill'iView repo

LLM judge agent for grading AI agent eval transcripts. Checks deterministic assertions (transcript_contains, tool_called) and uses LLM reasoning for behavioral assertions (llm_judge). Returns structured JSON grades.

Dosyalar1 dosya
SKILL.md108 satır
Loading editor…

Kurulum

Önerilen

Tek komut — ajanınız otomatik olarak devreye alır.

Kurulum komutunu görmek için yukarıdan bir AI aracı seçin.

veya

Manuel Kurulum

Daha fazla adım

Dosyayı indirin ve ajanınızın sistem istemine yapıştırın.

Skill detayları

Versiyonv1.0.0
YazarAWS Labs
Kategoriai-ml
Skill IDawslabs/agent-builder-toolkit-aws-transform/evaluation/src/eval_runner/execution/data/skills/eval-judge

İlgili skill'ler

Scenario RunnerSimulated human agent for eval scenarios. Interacts with the agent under test via the ACP bridge, following the scenario goal and guidance to respond to agent questions, approve tool calls, and drive the multi-turn flow to completion.Benefits AdvisorExplain an Acme employee benefit, including eligibility, employee cost, coverage, and key detailsAgui AuthorAuthor live dashboard UI from an agent via the `emit_ui` MCP tool. Emit one of six allow-listed components (approval_card, choice_prompt, diff_summary, progress, metric, agent_card) with JSON props and it renders in any AG-UI client watching the fleet. Use when you want the operator to see a decision, a diff, or a status readout instead of scrolling terminal text. Arbitrary HTML/markup is refused.Agui AuthorAuthor live dashboard UI from an agent via the `emit_ui` MCP tool. Emit one of six allow-listed components (approval_card, choice_prompt, diff_summary, progress, metric, agent_card) with JSON props and it renders in any AG-UI client watching the fleet. Use when you want the operator to see a decision, a diff, or a status readout instead of scrolling terminal text. Arbitrary HTML/markup is refused.Cao ContributingContribute changes to the CAO (CLI Agent Orchestrator) codebase — the local dev loop, the CI gate map, and the pre-PR checklist. Use when the user says "open a PR", "why did CI fail", "run the checks before I push", "the mypy/Code Quality job is red", "add a test and verify coverage", or when making any code change intended to land on a branch/PR. Covers uv-based build/test/lint, the ci.yml jobs and their pass/fail semantics, and the golden rules that stop a green-locally / red-in-CI surprise. NCao ContributingContribute changes to the CAO (CLI Agent Orchestrator) codebase — the local dev loop, the CI gate map, and the pre-PR checklist. Use when the user says "open a PR", "why did CI fail", "run the checks before I push", "the mypy/Code Quality job is red", "add a test and verify coverage", or when making any code change intended to land on a branch/PR. Covers uv-based build/test/lint, the ci.yml jobs and their pass/fail semantics, and the golden rules that stop a green-locally / red-in-CI surprise. N