ResearchPerspective

Equity Research Automation Statistics: A 2026 Evidence Check

What public surveys and finance benchmarks establish about research automation, what they do not measure, and how to calculate a team's own ROI.

Rida Malik

Published August 28, 2026 · Updated August 30, 2026

Editorial cover about evidence on equity research automation in 2026.
AllMind editorial artwork, August 2026. View article.
In this article

There is no reliable public statistic for “hours saved per equity research analyst” in 2026. Current evidence measures firm adoption, perceived benefits, onboarding friction, and model accuracy. Mercer found AI integrated into at least one investment process at 55% of 131 asset managers, but only 8% reported improved returns. Vals Finance Agent v2 shows why automation claims need task detail: the best category score was 79.1% for earnings analysis and 34.5% for financial modeling.

This article uses primary surveys, regulator guidance, and benchmark papers checked August 30, 2026. It excludes vendor time-saved claims without a published sample and method. The result is less dramatic than a “70% faster” headline and more useful for a head of research planning a controlled pilot.

What the public evidence actually measures

SourcePopulation or test setDateMeasure relevant to research automation
Mercer131 global asset managers; online survey in February and March 2026Published May 21, 202655% integrated AI into at least one investment process; 27% piloting; 69% reported operational efficiency; 8% reported improved returns
Substantive Research and Aiera35 of the largest global asset managersPublished July 16, 202677% had organization-wide general-platform deployments; 69% named licensing the top direct-feed barrier; 37% reported four-to-six-month onboarding
AIMA150 global fund managers, about $788 billion, plus 18 investorsPublished September 16, 202595% reported some generative AI use; 58% expected more use in investment processes over the following year
Vals Finance Agent v2927 expert-reviewed analyst questions; 450-question held-out test; three runs per modelLeaderboard August 19, 2026Leading model 60.60% partial credit and 50.88% all-pass; category results show task difficulty
Deep FinResearch Bench100 professional reports on 25 S&P 500 companies and reports from four agentsPosted April 22, 2026Professional report quality 2.84 versus 2.31 for the leading agent; agent factuality ranged from 53.2% to 86.0%
Bank of England and FCA118 UK-regulated financial firms across six sectorsPublished November 21, 202455% of AI use cases involved some automated decision-making; only 2% were fully autonomous; 84% of users named an accountable person
FINRA 2026 Regulatory Oversight ReportSupervisory observations, not a survey2026 reportSummarization and information extraction are the leading observed GenAI use; firms should test, log, monitor, and retain human review

These rows cannot be averaged. AIMA counts any generative AI use, Mercer counts investment-process integration, Vals tests model performance, and the Bank/FCA survey spans banking, insurance, markets, payments, and asset management.

Which equity-research tasks are closest to reliable automation?

Finance Agent v2 uses one common harness and one scoring system across nine analyst categories. Its category ceilings are therefore more comparable than productivity claims collected from different vendors and clients.

Finance Agent v2 chart showing stronger category-leading scores for earnings and disclosure analysis than for comparables, precedents, and financial modeling.
Best category scores on Vals Finance Agent v2, August 19, 2026. Different models may lead different categories, so the chart shows the current ceiling for each task, not one model's profile. Source: Vals AI.

The ordering suggests three automation tiers.

Good pilot candidates: bounded earnings and disclosure work. An analyst can specify the issuer, period, document set, and expected fields. The output can be checked against filings, earnings releases, transcripts, and consensus data. Finance Agent v2 category leaders reached 79.1% for earnings analysis and 71.3% for disclosure analysis.

Useful with structured review: adjustments and comparables. These require definition alignment, currency and period normalization, and decisions about one-time items. Category ceilings fell to 56.3% and 50.3%. The system can prepare a reviewable first pass, but a human still owns the definition and peer set.

High-risk unattended work: precedent transactions and financial modeling. Category ceilings were 36.4% and 34.5%. These tasks combine source selection, transaction or accounting conventions, several calculations, and a final judgment. They are appropriate for assistive workflows with locked assumptions and visible formulas, not a “run and publish” promise.

The full benchmark guide explains the scoring and why those percentages should not be compared with Deep FinResearch's report score.

Why adoption does not reveal analyst productivity

The current surveys leave four missing variables:

  1. Starting workflow. Automating manual copy-and-paste is different from replacing a mature data feed and model-update process.
  2. Unit of output. “One report,” “one company,” and “one earnings event” can represent radically different work.
  3. Review burden. Gross generation time ignores the time needed to verify sources, correct periods, and rebuild formulas.
  4. Failure cost. A missed source link and a wrong issuer are not equivalent. The latter should zero the task.

Mercer's 69% operational-efficiency result does not publish hours saved. It is a perception reported by respondents. Its 8% return result also uses each manager's own attribution method. Neither can be converted into a per-seat ROI without inventing assumptions.

Onboarding is part of the calculation. In the Substantive/Aiera sample, 37% said approving and onboarding models took four to six months, and another 20% reported more than six months. A 90-day pilot that begins before licensed feeds and internal permissions are ready will mostly test setup.

A worksheet for measuring automation without vendor math

Measure one recurring workflow for one quarter. An earnings-review pilot is usually cleaner than a broad “research copilot” trial because events, documents, and deadlines are observable.

FieldHow to measure it
Events in scopeCount completed earnings events, not companies licensed or seats provisioned
Baseline analyst minutesMedian active time from a sample of at least 10 pre-pilot events
AI-assisted analyst minutesSame timer boundary, including review and corrections
Completion rateEvents delivered by the agreed deadline divided by events in scope
Source coverageCheckable factual claims with a correct source passage divided by all checkable claims
Critical-error rateOutputs with any wrong issuer, period, currency, unit, or permission breach divided by outputs reviewed
Rework minutesTime spent correcting or recreating the output after first delivery
Fully loaded costTool cost plus onboarding, data, integration, and reviewer time for the quarter

Calculate net analyst hours saved by subtracting assisted time and rework time from the baseline minutes for each completed event, then convert the total to hours. Divide the fully loaded quarterly cost by those net saved hours to obtain cost per net hour saved.

If net hours are zero or negative, report that result. Do not remove failures from the denominator because a run timed out or returned an unusable deliverable.

A 12-task research-automation test

Use a frozen source pack and run the same tasks in every product:

  • three earnings reviews, including one company with non-GAAP adjustments;
  • two guidance-history updates across at least four quarters;
  • two filing-change checks across consecutive annual reports;
  • two three-company comparable tables with a written normalization policy;
  • one model update with formulas exposed;
  • one internal memo plus public-filing synthesis under a restricted role;
  • one scheduled monitor that records what changed and why it triggered.

For each run, retain the prompt, product version, operator, start and finish time, output, cited passages, corrections, and failure state. Define critical errors in advance. A wrong issuer, a period mismatch, an invented citation, or unauthorized data access should fail the task even when the prose looks plausible.

FINRA's 2026 guidance makes the governance portion explicit: formal approval, comprehensive documentation, model and output testing, prompt/output logs, ongoing monitoring, and human-in-the-loop review. Those controls are not overhead external to the automation. They are part of the production workflow being measured.

What to automate first

Start where the task has a stable input, an observable output, and a cheap verification path:

  • filing and transcript retrieval with passage-level citations;
  • earnings tables whose periods, units, and definitions are fixed;
  • guidance-history extraction with an explicit metric dictionary;
  • watchlist monitoring where every alert names the trigger and source;
  • first-draft comparison tables that expose all calculations.

Delay unattended modeling, recommendation generation, and trade-related decisions until the system has passed the bounded tasks. Mercer's 5% figure for autonomous or semi-autonomous recommendation/trade authority and the Bank/FCA's 2% fully autonomous use-case figure are different measures, but both support the same conservative sequence.

Where this evidence stops

We retained statistics only when an original publisher supplied a sample or test-set description, a date, and the measured construct. We removed claims such as “two hours per ticker,” “24 hours per week,” or “70% faster” when the source was a vendor page without a disclosed sample and method.

The surveys are voluntary and self-reported. The benchmarks test foundation models or agent harnesses over mostly public information, not commercial platforms with broker entitlements, internal models, permissions, or operational support. Live leaderboards will change. None of these sources reports a controlled, industry-wide estimate of net analyst hours saved.

The only defensible productivity statistic for a specific team is the one produced by that team's timed baseline and controlled run. Public data should set the test design and the caution level, not fill in the result.