Hugging Face has released FinanceComplexQA, a comprehensive benchmark designed to evaluate agentic reasoning on complex, real-world financial documents. The benchmark addresses the growing need for reliable AI agents capable of handling industrial-grade financial analysis.
FinanceComplexQA is built on a novel skill called Finance-LaTeX SKILL, which synthesizes financial documents with complex layouts using expert knowledge. An agent workflow based on this skill generated 2,000 professional documents and 6,000 high-quality question-answer pairs. The final benchmark comprises 2,026 deep research tasks targeting 1,009 financial documents.
Key features of FinanceComplexQA include bilingual support (English and Chinese), coverage of six mainstream scenarios and seven task types, expert-level document reasoning questions, deep research on complex layouts, stable and permanent reference answers, and precise evaluation through an Agent-as-a-Judge with multiple metrics.
The researchers used FinanceComplexQA to evaluate leading RAG systems and agentic reasoning tools for financial document QA. By analyzing failure cases, they studied capabilities in numerical computation, multi-hop reasoning, content summarization, and industry analysis.
The benchmark is part of a growing ecosystem of financial AI benchmarks, including AGORA, MoCA-Agent, ICBCBench, FORCE-Bench, LakeQA, CM-LRS, and DocArena, as noted by the Librarian Bot recommendations.