AI Financial Research Benchmarks: What the Scores Mean
A buyer's guide to Finance Agent v2, Deep FinResearch, BigFinanceBench, Fin-RATE, and FinanceBench, including samples, scoring, and limits.
Published August 28, 2026 · Updated August 30, 2026

In this article
No public finance benchmark provides a single product score a research buyer can rely on. Vals Finance Agent v2 tests precise analyst questions and tops out at 60.60% with partial credit. Deep FinResearch Bench grades complete reports and still puts professional analysts ahead, 2.84 to 2.31 on a four-point quality scale. BigFinanceBench grades a derivation and a final answer separately. The useful lesson is the gap pattern, not a league table across incompatible tests.
AllMind has not submitted a result to any benchmark discussed here. This guide uses each benchmark's paper or official leaderboard, checked August 30, 2026. It does not convert one scoring system into another or treat a foundation-model result as evidence for a commercial research platform.
Five benchmarks, five different units of work
| Benchmark | Unit tested | Sample and access | Published measure used here | Main limitation |
|---|---|---|---|---|
| Finance Agent v2 | Multi-source questions expected of a second- or third-year banking analyst | 927 questions: 27 public, 450 licensed validation, 450 held-out test; three runs per model | 60.60% partial credit and 50.88% all-pass for the leading model on August 19, 2026 | Private test set and LLM jury; tests a common harness, not a vendor platform |
| Deep FinResearch Bench | Full pre-earnings investment-research reports | 100 professional reports on 25 S&P 500 companies, compared with four deep-research agents | Professional report 2.84; leading agent 2.31 on a four-point quality scale | Two financial institutions and three sectors; automated judges support most scoring |
| BigFinanceBench | Open-ended research question plus auditable calculation and sourcing steps | 928 questions and 15,656 weighted rubric criteria; official leaderboard covered 28 models on August 3, 2026 | Leader 67.1% by rubric and 53.4% by final answer | Created by Rogo, a research vendor; live results use LLM judges |
| Fin-RATE | SEC-filing reasoning within one document, across companies, and across periods | 17 models under ground-truth-context and retrieval settings | Reported accuracy declines of 18.60% for longitudinal and 14.35% for cross-entity tasks versus single-document work | Reports degradation, not one transferable platform score |
| FinanceBench | Open-book filing question answering with an evidence string | 10,231 question-answer-evidence triplets; 150 used for expert evaluation | GPT-4-Turbo with retrieval incorrectly answered or refused 81% of the 150-case evaluation in 2023 | Old models; wrong answers and refusals share the headline measure |
The scores do not belong in one bar chart. A 60.60% question-level partial-credit result cannot be ranked against a 2.31 report-quality score or a 53.4% final-answer rate. Even the two BigFinanceBench percentages describe different objects: how much of the derivation was present and whether the submitted answer itself was correct.
What Deep FinResearch says about complete reports
Deep FinResearch Bench, posted by JPMorganChase AI Research on April 22, 2026, is the most direct public comparison between agent output and professional investment-research reports. The study used 100 pre-earnings reports from two financial institutions, covering 25 S&P 500 companies in information technology, financials, and health care. Four deep-research agents produced reports on the same companies.
The same study shows why a single report score is incomplete. Claim factuality was 86.0% for OpenAI, 75.6% for Perplexity, 69.6% for Gemini, and 53.2% for Grok. Gemini wrote the highest-rated agent report but did not lead factuality. A buyer choosing only the best-looking draft could choose more review work.
Forecast error tells a third story. The professional reports had 17.14% SMAPE across six forecast line items. The agent results ranged from 17.49% to 27.56%. One agent was close to the professional result overall, while the agents still trailed on report coherence and depth. A platform evaluation should therefore score factual support, quantitative accuracy, and decision usefulness separately.
Finance Agent v2 is useful because its failures look like analyst work
The Vals Finance Agent v2 page describes 927 expert-reviewed questions divided into nine categories. Every agent has the same two-hour limit and six tools: EDGAR search, web search, HTML parsing, retrieval over fetched pages, a calculator, and price history. Each model runs three times.
The primary metric is dealbreaker-gated partial credit. A response gets weighted credit for checks it passes, but a missed load-bearing fact can zero the question. The all-pass metric gives credit only when every check passes. On August 19, 2026, the leading model scored 60.60% under partial credit and 50.88% all-pass.
Category ceilings expose where the difficulty sits:
| Finance Agent v2 category | Best category score on August 19, 2026 |
|---|---|
| General quantitative analysis | 81.8% |
| Earnings analysis | 79.1% |
| Disclosure analysis | 71.3% |
| Adjustments | 56.3% |
| Comparables | 50.3% |
| Precedent transactions | 36.4% |
| Financial modeling | 34.5% |
These are category leaders, sometimes different models, not one model's progression. They still show a coherent pattern: retrieval and bounded analysis are stronger than work that requires normalization, transaction conventions, or a multi-step model. The benchmark paper documents the original 537-question version and its expert-authored workflow taxonomy; version 2 changes the dataset, grading, toolset, and difficulty, so v1.1 and v2 results should not be treated as a time series.
BigFinanceBench separates the trail from the answer
BigFinanceBench's paper describes 928 questions written by 52 finance subject-matter experts and audited by 12 reviewers. The questions decompose into 15,656 weighted criteria and 36,241 total rubric points. The official site reported 28 models on August 3, 2026.
For the leading model, the live leaderboard showed 67.1% rubric credit and 53.4% final-answer accuracy. That 13.7-point gap is the benchmark's most useful result for a buyer. An answer can include much of the right retrieval and calculation trail and still fail the final decision. Conversely, a correct number without its derivation may be hard to review.
Rogo created BigFinanceBench and sells a finance research product. That conflict does not invalidate the dataset, but it belongs next to the score. The live leaderboard also states that two independent model judges grade the results. A team using the leaderboard should inspect public examples and rerun a subset with human grading before treating small model differences as meaningful.
FinanceBench and Fin-RATE explain why retrieval is not enough
FinanceBench remains a useful floor. It contains 10,231 questions, answers, and evidence strings over 361 filings from 40 US public companies. Its famous 81% result came from a 150-case expert evaluation in 2023: GPT-4-Turbo with a retrieval system either answered incorrectly or refused. It was not an 81% error rate across all 10,231 cases, and it does not describe current models.
Fin-RATE moves the task from isolated filing questions to the work analysts actually struggle to automate: comparing companies and tracking one company through time. Across 17 models, its authors report accuracy declines of 18.60% for longitudinal work and 14.35% for cross-entity analysis relative to single-document reasoning. They also classify whether failures come from retrieval, generation, finance reasoning, or misunderstanding the query.
Together, these benchmarks argue for a staged test. First ask whether the system found the correct filing passage. Then test whether it aligned the right company, period, unit, definition, and restatement. Finally, grade the answer and its evidence trail.
A 20-question benchmark a buyer can actually run
Public model leaderboards cannot test a firm's licensed research, internal models, house definitions, or permission boundaries. A compact product evaluation can:
| Task block | Questions | Required evidence | Pass condition |
|---|---|---|---|
| Single-document retrieval | 4 | Exact filing or transcript passage | Correct answer, period, unit, and linked passage |
| Longitudinal change | 4 | Two or more periods and any restatement note | Correct comparable series and explanation of definition changes |
| Cross-company comparison | 4 | Same metric for at least three issuers | Entity, currency, fiscal period, and normalization all correct |
| Internal-plus-external synthesis | 4 | House model or memo joined to public material | Permission-respecting use of both source classes, each claim traceable |
| Deliverable quality | 4 | Memo, table, or workbook | Correct final conclusion, complete evidence trail, and usable format |
Run every product on the same frozen source pack. Record product version, access tier, operator, exact prompt, elapsed time, output, citations, and failure state. Score each question twice: all-pass and partial credit. Define dealbreakers before the run, such as wrong issuer, wrong period, hidden unit conversion, invented source, or access to a document the user should not see.
Do not import a public model score into the result. A product can add licensed retrieval, ontology, permissions, and verification, or it can introduce new failure points. The only comparable product score is the one earned under the same inputs and grading.
How to compare these benchmarks without inventing a league table
We used primary papers and official live leaderboards accessed August 30, 2026. Leaderboard values are a dated snapshot and can change when a new model, harness, judge, or dataset version appears. We did not reproduce the private test sets or independently audit every question.
Three limitations matter most:
- Benchmark owners have different incentives. Vals sells evaluation, while Rogo sells a research product. Ownership is disclosed instead of used as a reason to discard the work.
- Several metrics depend on LLM judges. Human spot checks and inter-rater analysis reduce, but do not eliminate, judge error.
- Public benchmarks use public filings and web data. They do not measure broker-research entitlements, internal data controls, product uptime, onboarding, or the review workflow around a finished output.
AllMind has published no result on these benchmarks as of August 30, 2026. A buyer evaluating it should use the same 20 questions and retain the failures, not accept a foundation-model leaderboard as a substitute.