Developed as part of a litigation platform I've started, but not sharing this benchmark to promote that. Just thought the findings were interesting--namely, Anthropic's models (except for Haiku) producing zero hallucinated cases.
Also, if anyone would like to see additional tests / benchmarks or models tested, let me know and I'll incorporate them into v3 if it makes sense.