Bean Labs
Researching the boundaries of autonomous financial intelligence.
A research initiative by Beancount.io
Bean Labs publishes open research notes on autonomous bookkeeping — language models that reason over double-entry ledgers, produce verifiable audit trails, and stay grounded in plain-text accounting.
Recent research notes
Open experiments and findings from Bean Labs — the Finance AI Agent research initiative by Beancount.io.
Research by topic
Browse the research log by the themes we return to most often.
LLM
Large language model research with applications in financial tasks
- Can your agent balance these books?
- FinRAGBench-V: Multimodal RAG with Visual Citations in the Financial Domain
- Can LLM Agents Be CFOs? EnterpriseArena's 132-Month Simulation Reveals a Wide Gap
- WildToolBench: Why No LLM Exceeds 15% Session Accuracy in Real-World Tool Use
- LLM Confidence and Calibration: A Survey of What the Research Actually Shows
- JSONSchemaBench: Real-World Schema Complexity Breaks LLM Structured Output Guarantees
- FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under MCP
- FinTrace: Trajectory-Level Evaluation of LLM Tool Calling for Financial Tasks
- FinToolBench: Evaluating LLM Agents on Real-World Financial Tool Use
- OmniEval: Omnidirectional RAG Evaluation Benchmark for the Financial Domain
- LLM Anomaly Detection Survey (NAACL 2025): Strong Taxonomy, Absent Tabular Coverage
- Found in the Middle: Calibrating Positional Attention Bias Improves Long-Context RAG
AI
Artificial intelligence research and applications in finance and accounting
- FinRAGBench-V: Multimodal RAG with Visual Citations in the Financial Domain
- Can LLM Agents Be CFOs? EnterpriseArena's 132-Month Simulation Reveals a Wide Gap
- WildToolBench: Why No LLM Exceeds 15% Session Accuracy in Real-World Tool Use
- LLM Confidence and Calibration: A Survey of What the Research Actually Shows
- JSONSchemaBench: Real-World Schema Complexity Breaks LLM Structured Output Guarantees
- FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under MCP
- FinTrace: Trajectory-Level Evaluation of LLM Tool Calling for Financial Tasks
- FinToolBench: Evaluating LLM Agents on Real-World Financial Tool Use
- OmniEval: Omnidirectional RAG Evaluation Benchmark for the Financial Domain
- LLM Anomaly Detection Survey (NAACL 2025): Strong Taxonomy, Absent Tabular Coverage
- Found in the Middle: Calibrating Positional Attention Bias Improves Long-Context RAG
- Uncertainty-Aware Deferral for LLM Agents: When to Escalate from Small to Large Models
Machine Learning
Machine learning techniques for financial data analysis and automation
- FinRAGBench-V: Multimodal RAG with Visual Citations in the Financial Domain
- WildToolBench: Why No LLM Exceeds 15% Session Accuracy in Real-World Tool Use
- LLM Confidence and Calibration: A Survey of What the Research Actually Shows
- JSONSchemaBench: Real-World Schema Complexity Breaks LLM Structured Output Guarantees
- FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under MCP
- FinTrace: Trajectory-Level Evaluation of LLM Tool Calling for Financial Tasks
- FinToolBench: Evaluating LLM Agents on Real-World Financial Tool Use
- OmniEval: Omnidirectional RAG Evaluation Benchmark for the Financial Domain
- LLM Anomaly Detection Survey (NAACL 2025): Strong Taxonomy, Absent Tabular Coverage
- Found in the Middle: Calibrating Positional Attention Bias Improves Long-Context RAG
- Uncertainty-Aware Deferral for LLM Agents: When to Escalate from Small to Large Models
- OpenHands: Open Platform for AI Software Agents and What It Means for Finance Automation
Beancount
Beancount ledger format, tooling, and ecosystem research
- Can your agent balance these books?
- FinRAGBench-V: Multimodal RAG with Visual Citations in the Financial Domain
- Can LLM Agents Be CFOs? EnterpriseArena's 132-Month Simulation Reveals a Wide Gap
- WildToolBench: Why No LLM Exceeds 15% Session Accuracy in Real-World Tool Use
- JSONSchemaBench: Real-World Schema Complexity Breaks LLM Structured Output Guarantees
- FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under MCP
- FinTrace: Trajectory-Level Evaluation of LLM Tool Calling for Financial Tasks
- FinToolBench: Evaluating LLM Agents on Real-World Financial Tool Use
- OmniEval: Omnidirectional RAG Evaluation Benchmark for the Financial Domain
- LLM Anomaly Detection Survey (NAACL 2025): Strong Taxonomy, Absent Tabular Coverage
- Found in the Middle: Calibrating Positional Attention Bias Improves Long-Context RAG
- Uncertainty-Aware Deferral for LLM Agents: When to Escalate from Small to Large Models
Automation
Automation techniques and tools for financial data processing workflows
- Can LLM Agents Be CFOs? EnterpriseArena's 132-Month Simulation Reveals a Wide Gap
- WildToolBench: Why No LLM Exceeds 15% Session Accuracy in Real-World Tool Use
- JSONSchemaBench: Real-World Schema Complexity Breaks LLM Structured Output Guarantees
- FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under MCP
- FinTrace: Trajectory-Level Evaluation of LLM Tool Calling for Financial Tasks
- FinToolBench: Evaluating LLM Agents on Real-World Financial Tool Use
- OmniEval: Omnidirectional RAG Evaluation Benchmark for the Financial Domain
- Found in the Middle: Calibrating Positional Attention Bias Improves Long-Context RAG
- Uncertainty-Aware Deferral for LLM Agents: When to Escalate from Small to Large Models
- OpenHands: Open Platform for AI Software Agents and What It Means for Finance Automation
- TableMaster: Adaptive Reasoning for Table Understanding with LLMs
- Zero-Shot Anomaly Detection with LLMs: How GPT-4 Performs on Tabular Data
Data Science
Data science methods applied to financial datasets and accounting workflows
- FinRAGBench-V: Multimodal RAG with Visual Citations in the Financial Domain
- WildToolBench: Why No LLM Exceeds 15% Session Accuracy in Real-World Tool Use
- LLM Confidence and Calibration: A Survey of What the Research Actually Shows
- FinToolBench: Evaluating LLM Agents on Real-World Financial Tool Use
- OmniEval: Omnidirectional RAG Evaluation Benchmark for the Financial Domain
- LLM Anomaly Detection Survey (NAACL 2025): Strong Taxonomy, Absent Tabular Coverage
- Found in the Middle: Calibrating Positional Attention Bias Improves Long-Context RAG
- Fin-RATE: How LLMs Fail at Cross-Period and Cross-Entity Financial Analysis
- FinDER: Real Analyst Queries Expose a 74% Recall Gap in Financial RAG
- Lost in the Middle: Position Bias in LLMs and Its Impact on Finance AI
- AD-LLM Benchmark: GPT-4o Hits 0.93+ AUROC Zero-Shot for Text Anomaly Detection
- CausalTAD: Causal Column Ordering for LLM Tabular Anomaly Detection
Finance
Financial research, analysis, and domain knowledge for accounting AI
- FinRAGBench-V: Multimodal RAG with Visual Citations in the Financial Domain
- LLM Confidence and Calibration: A Survey of What the Research Actually Shows
- FinTrace: Trajectory-Level Evaluation of LLM Tool Calling for Financial Tasks
- OmniEval: Omnidirectional RAG Evaluation Benchmark for the Financial Domain
- FinDER: Real Analyst Queries Expose a 74% Recall Gap in Financial RAG
- Lost in the Middle: Position Bias in LLMs and Its Impact on Finance AI
- AnoLLM: Fine-Tuning LLMs for Tabular Anomaly Detection in Financial Data
- DocFinQA: Long-Context Financial Reasoning on Full SEC Filings
- TheAgentCompany: Benchmarking LLM Agents on Real-World Enterprise Tasks
- InvestorBench: Qwen2.5-72B tops stocks at 46.15% CR
- Single-Agent: equals MAS on multi-hop at equal tokens
- M3MAD-Bench: Are Multi-Agent Debates Really Effective Across Domains and Modalities?
Plain-Text Accounting
Research grounded in plain-text accounting formats and workflows
- Can your agent balance these books?
- Uncertainty-Aware Deferral for LLM Agents: When to Escalate from Small to Large Models
- OpenHands: Open Platform for AI Software Agents and What It Means for Finance Automation
- LLMs Score 2.3% on Beancount DSL Generation: The LLMFinLiteracy Benchmark
- TableMaster: Adaptive Reasoning for Table Understanding with LLMs
- τ²-bench: Measuring the Cost of Dual-Control in Conversational AI Agents
- GAIA Benchmark: Measuring What Frontier AI Agents Can Actually Do
- WorkArena: How LLM Web Agents Perform on Real Enterprise Knowledge Work
- τ-bench: Measuring AI Agent Reliability in Real-World Tool-Use Domains
- Chain-of-Table: Evolving Tables in the LLM Reasoning Chain
- TableLlama: Can a 7B Open Model Match GPT-4 on Table Understanding?
- TAPAS: Weakly Supervised Table QA Without SQL, and What It Means for Beancount
How we approach research
Plain-text first
Start with plain-text ledgers so the inputs remain readable, portable, and open to inspection.
Reproducible by default
Make the inputs, methods, and limitations explicit so others can examine and build on the work.
Open findings
Publish research notes openly and invite the community to question, extend, and improve the findings.
Follow the research
Read the published notes and follow the next questions in the research log.