Agents CLI: Agent Evaluation

Part of Google's agents-cli skills suite. Documents the Agent Platform's evaluation methodology: preparing an EvaluationDataset (single-turn, multi-turn, and multi-agent JSON schemas), running `agents-cli eval run` to execute the agent over the dataset and grade traces, analyzing failed metrics, and iterating via prompt/tool/instruction fixes. Also covers opt-in commands for user-simulated multi-turn datasets, LLM-based failure clustering, and ADK GEPA prompt optimization. Applies to any agents-cli project regardless of framework. Does not cover ADK API code patterns (google-agents-cli-adk-code), deployment (google-agents-cli-deploy), or scaffolding (google-agents-cli-scaffold).

This skill hasn't been reviewed yet.

What it does

Evaluation methodology for agents built with Google's agents-cli — dataset schema, built-in metrics, LLM-as-judge scoring, and the failure-analysis loop (the Quality Flywheel).

How to use this

Once it is installed, this skill picks itself up. You do not need to name it.

“evaluate my agent”
“evaluate my ADK agent”
“write an eval dataset for my agent”
“analyze eval failures”
Created by
Google LLC
Category
Engineering

Run this on Skills and Agents

You can run the skill on Skills and Agents by creating an account. You don't need to install anything. We provide the model and tokens. Your connected tools and company brain are ready to use.

Create an account

Skills give you best practices and templates to get the most out of your AI