
TableMaster: Adaptive Reasoning for Table Understanding with LLMs
TableMaster hits 78.13% on WikiTQ with GPT-4o-mini, 13 points over Chain-of-Table, via table-of-focus plus adaptive reasoning for agents over Beancount ledgers.
#queries
Генериране на заявки, разсъждения върху таблици и извличане на структурирани данни за финансов ИИ

TableMaster hits 78.13% on WikiTQ with GPT-4o-mini, 13 points over Chain-of-Table, via table-of-focus plus adaptive reasoning for agents over Beancount ledgers.

Chain-of-Table hits 67.31% on WikiTQ versus 61.48% for text-only chain-of-thought, and leads by 10.25 points on tables over 4,000 tokens.

TableLlama beats GPT-4 on column type annotation (F1 94 vs 32) but trails by 33 points on WikiTQ compositional reasoning.

TAPAS answers table questions by selecting cells, never generating SQL. It fits small Beancount ledger queries but breaks down at scale.

MAC-SQL's three-agent design hits 59.59% execution accuracy on BIRD, with the Refiner adding +4.63 points — a template for generating Beancount ledger queries.

DIN-SQL lifts GPT-4 from 67.4% to 85.3% Spider execution accuracy via schema-link and self-correct stages—same decomposition fits Beancount BQL.

On BIRD, GPT-4 reaches 54.89% execution accuracy with domain hints and 34.88% without—a 20-point gap any Beancount NL→BQL interface must close.

Microsoft's GraphRAG posts 72–83% sensemaking wins over vector RAG; a 2025 audit collapses those after correcting judge bias—caution for multi-doc ledger QA.