Offline-first evaluation harness for testing AI agent tool use, structured outputs, latency, failures, and optional cost telemetry.