Bean Labs
Изследване на границите на автономния финансов интелект.
Изследователска инициатива на Beancount.io
Bean Labs публикува отворени изследователски бележки за автономно счетоводство — езикови модели, които разсъждават върху двойно-вписващи счетоводни книги, създават проверими одитни следи и остават основани на счетоводството в обикновен текст.
Последни изследователски бележки
Отворени експерименти и открития от Bean Labs — изследователската инициатива Finance AI Agent на Beancount.io.
Изследвания по тема
Разгледайте изследователския дневник по темите, към които се връщаме най-често.
LLM
Large language model research with applications in financial tasks
- Може ли вашият агент да изравни тези сметки?
- FinRAGBench-V: Multimodal RAG with Visual Citations in the Financial Domain
- Can LLM Agents Be CFOs? EnterpriseArena's 132-Month Simulation Reveals a Wide Gap
- WildToolBench: Why No LLM Exceeds 15% Session Accuracy in Real-World Tool Use
- LLM Confidence and Calibration: A Survey of What the Research Actually Shows
- JSONSchemaBench: Real-World Schema Complexity Breaks LLM Structured Output Guarantees
- FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under MCP
- FinTrace: Trajectory-Level Evaluation of LLM Tool Calling for Financial Tasks
- FinToolBench: Оценка на LLM агенти при реален финансов инструментариум
- OmniEval: Omnidirectional RAG Evaluation Benchmark for the Financial Domain
- LLM Anomaly Detection Survey (NAACL 2025): Strong Taxonomy, Absent Tabular Coverage
- Found in the Middle: Calibrating Positional Attention Bias Improves Long-Context RAG
AI
Artificial intelligence research and applications in finance and accounting
- FinRAGBench-V: Multimodal RAG with Visual Citations in the Financial Domain
- Can LLM Agents Be CFOs? EnterpriseArena's 132-Month Simulation Reveals a Wide Gap
- WildToolBench: Why No LLM Exceeds 15% Session Accuracy in Real-World Tool Use
- LLM Confidence and Calibration: A Survey of What the Research Actually Shows
- JSONSchemaBench: Real-World Schema Complexity Breaks LLM Structured Output Guarantees
- FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under MCP
- FinTrace: Trajectory-Level Evaluation of LLM Tool Calling for Financial Tasks
- FinToolBench: Оценка на LLM агенти при реален финансов инструментариум
- OmniEval: Omnidirectional RAG Evaluation Benchmark for the Financial Domain
- LLM Anomaly Detection Survey (NAACL 2025): Strong Taxonomy, Absent Tabular Coverage
- Found in the Middle: Calibrating Positional Attention Bias Improves Long-Context RAG
- Uncertainty-Aware Deferral for LLM Agents: When to Escalate from Small to Large Models
Machine Learning
Machine learning techniques for financial data analysis and automation
- FinRAGBench-V: Multimodal RAG with Visual Citations in the Financial Domain
- WildToolBench: Why No LLM Exceeds 15% Session Accuracy in Real-World Tool Use
- LLM Confidence and Calibration: A Survey of What the Research Actually Shows
- JSONSchemaBench: Real-World Schema Complexity Breaks LLM Structured Output Guarantees
- FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under MCP
- FinTrace: Trajectory-Level Evaluation of LLM Tool Calling for Financial Tasks
- FinToolBench: Оценка на LLM агенти при реален финансов инструментариум
- OmniEval: Omnidirectional RAG Evaluation Benchmark for the Financial Domain
- LLM Anomaly Detection Survey (NAACL 2025): Strong Taxonomy, Absent Tabular Coverage
- Found in the Middle: Calibrating Positional Attention Bias Improves Long-Context RAG
- Uncertainty-Aware Deferral for LLM Agents: When to Escalate from Small to Large Models
- OpenHands: Open Platform for AI Software Agents and What It Means for Finance Automation
Beancount
Beancount ledger format, tooling, and ecosystem research
- Може ли вашият агент да изравни тези сметки?
- FinRAGBench-V: Multimodal RAG with Visual Citations in the Financial Domain
- Can LLM Agents Be CFOs? EnterpriseArena's 132-Month Simulation Reveals a Wide Gap
- WildToolBench: Why No LLM Exceeds 15% Session Accuracy in Real-World Tool Use
- JSONSchemaBench: Real-World Schema Complexity Breaks LLM Structured Output Guarantees
- FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under MCP
- FinTrace: Trajectory-Level Evaluation of LLM Tool Calling for Financial Tasks
- FinToolBench: Оценка на LLM агенти при реален финансов инструментариум
- OmniEval: Omnidirectional RAG Evaluation Benchmark for the Financial Domain
- LLM Anomaly Detection Survey (NAACL 2025): Strong Taxonomy, Absent Tabular Coverage
- Found in the Middle: Calibrating Positional Attention Bias Improves Long-Context RAG
- Uncertainty-Aware Deferral for LLM Agents: When to Escalate from Small to Large Models
Automation
Automation techniques and tools for financial data processing workflows
- Can LLM Agents Be CFOs? EnterpriseArena's 132-Month Simulation Reveals a Wide Gap
- WildToolBench: Why No LLM Exceeds 15% Session Accuracy in Real-World Tool Use
- JSONSchemaBench: Real-World Schema Complexity Breaks LLM Structured Output Guarantees
- FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under MCP
- FinTrace: Trajectory-Level Evaluation of LLM Tool Calling for Financial Tasks
- FinToolBench: Оценка на LLM агенти при реален финансов инструментариум
- OmniEval: Omnidirectional RAG Evaluation Benchmark for the Financial Domain
- Found in the Middle: Calibrating Positional Attention Bias Improves Long-Context RAG
- Uncertainty-Aware Deferral for LLM Agents: When to Escalate from Small to Large Models
- OpenHands: Open Platform for AI Software Agents and What It Means for Finance Automation
- TableMaster: Adaptive Reasoning for Table Understanding with LLMs
- Zero-Shot Anomaly Detection with LLMs: How GPT-4 Performs on Tabular Data
Data Science
Data science methods applied to financial datasets and accounting workflows
- FinRAGBench-V: Multimodal RAG with Visual Citations in the Financial Domain
- WildToolBench: Why No LLM Exceeds 15% Session Accuracy in Real-World Tool Use
- LLM Confidence and Calibration: A Survey of What the Research Actually Shows
- FinToolBench: Оценка на LLM агенти при реален финансов инструментариум
- OmniEval: Omnidirectional RAG Evaluation Benchmark for the Financial Domain
- LLM Anomaly Detection Survey (NAACL 2025): Strong Taxonomy, Absent Tabular Coverage
- Found in the Middle: Calibrating Positional Attention Bias Improves Long-Context RAG
- Fin-RATE: How LLMs Fail at Cross-Period and Cross-Entity Financial Analysis
- FinDER: Real Analyst Queries Expose a 74% Recall Gap in Financial RAG
- Lost in the Middle: Position Bias in LLMs and Its Impact on Finance AI
- AD-LLM Benchmark: GPT-4o Hits 0.93+ AUROC Zero-Shot for Text Anomaly Detection
- CausalTAD: Causal Column Ordering for LLM Tabular Anomaly Detection
Finance
Financial research, analysis, and domain knowledge for accounting AI
- FinRAGBench-V: Multimodal RAG with Visual Citations in the Financial Domain
- LLM Confidence and Calibration: A Survey of What the Research Actually Shows
- FinTrace: Trajectory-Level Evaluation of LLM Tool Calling for Financial Tasks
- OmniEval: Omnidirectional RAG Evaluation Benchmark for the Financial Domain
- FinDER: Real Analyst Queries Expose a 74% Recall Gap in Financial RAG
- Lost in the Middle: Position Bias in LLMs and Its Impact on Finance AI
- AnoLLM: Fine-Tuning LLMs for Tabular Anomaly Detection in Financial Data
- DocFinQA: Long-Context Financial Reasoning on Full SEC Filings
- TheAgentCompany: Benchmarking LLM Agents on Real-World Enterprise Tasks
- InvestorBench: Qwen2.5-72B tops stocks at 46.15% CR
- Single-Agent: equals MAS on multi-hop at equal tokens
- M3MAD-Bench: Are Multi-Agent Debates Really Effective Across Domains and Modalities?
Plain-Text Accounting
Research grounded in plain-text accounting formats and workflows
- Може ли вашият агент да изравни тези сметки?
- Uncertainty-Aware Deferral for LLM Agents: When to Escalate from Small to Large Models
- OpenHands: Open Platform for AI Software Agents and What It Means for Finance Automation
- LLMs Score 2.3% on Beancount DSL Generation: The LLMFinLiteracy Benchmark
- TableMaster: Adaptive Reasoning for Table Understanding with LLMs
- τ²-bench: Measuring the Cost of Dual-Control in Conversational AI Agents
- GAIA Benchmark: Measuring What Frontier AI Agents Can Actually Do
- WorkArena: How LLM Web Agents Perform on Real Enterprise Knowledge Work
- τ-bench: Measuring AI Agent Reliability in Real-World Tool-Use Domains
- Chain-of-Table: Evolving Tables in the LLM Reasoning Chain
- TableLlama: Can a 7B Open Model Match GPT-4 on Table Understanding?
- TAPAS: Weakly Supervised Table QA Without SQL, and What It Means for Beancount
Нашият подход към изследванията
Обикновен текст на първо място
Започваме с книги в обикновен текст, за да останат входните данни четими, преносими и достъпни за проверка.
Възпроизводимо по подразбиране
Описваме ясно входните данни, методите и ограниченията, за да могат други да проверяват и надграждат работата.
Отворени открития
Публикуваме изследователските бележки открито и каним общността да оспорва, допълва и подобрява изводите.
Следете изследванията
Прочетете публикуваните бележки и следете следващите въпроси в изследователския дневник.