We(a team of ex-EY employees) built FAB - Finance Agents Benchmark, testing AI Agents on the work behind financial due diligence. Agents can find the relevant facts, but still struggle to carry them through to a complete, reliable analysis.
We’ll keep expanding FAB to more companies and testing more models. Building and running this benchmark isn’t cheap, so we’re scaling it in stages.
The benchmark is public.
GitHub: Hugging Face: https://github.com/SecondState-ai/finance-agents-benchmark data room and tasks: https://huggingface.co/datasets/secondstate/finance-agents-b...