Benchmarking the ability of large language models to detect semantic conflicts across domains, documents, and evolving knowledge bases.