AI Agent Evaluation and Benchmarking SOP: How to Scientifically Measure an Agent
An AI agent evaluation and benchmarking SOP: four metric categories (success rate/efficiency/safety/robustness), building an eval set (typical/edge/adversarial 60:25:15, programmatic scoring first), multi-run sampling with baseline comparison, and latency/cost logging. References the site's V4-Flash hands-on harness method.