A multi-agent AI system that benchmarks and ranks language models based on project-specific requirements, quality, latency, and cost.