
MultiHiertt: Benchmarking Numerical Reasoning Over Multi-Hierarchical Financial Tables
MultiHiertt shows models score 38% F1 against 87% for humans on 10,440 financial QA pairs, with a 15-point drop on cross-table questions.
#financial-statements
Balance sheet, income statement, and cash-flow generation research

MultiHiertt shows models score 38% F1 against 87% for humans on 10,440 financial QA pairs, with a 15-point drop on cross-table questions.

FinanceBench tests 16 AI setups on 10,231 real SEC filing questions: shared-vector-store RAG answers only 19% right, so retrieval is not the bottleneck.

FinMaster Benchmark: top LLMs hit 96% on financial literacy but only 3% on statement generation. Error propagation costs 21 points on consulting tasks.