Benchmark-driven evaluation harness for testing accuracy, groundedness and hallucination behaviour in a document-grounded AI agent.