Equity Research Automation Statistics: A 2026 Evidence Check
What public surveys and finance benchmarks establish about research automation, what they do not measure, and how to calculate a team's own ROI.
Published August 28, 2026 · Updated August 30, 2026

In this article
There is no reliable public statistic for “hours saved per equity research analyst” in 2026. Current evidence measures firm adoption, perceived benefits, onboarding friction, and model accuracy. Mercer found AI integrated into at least one investment process at 55% of 131 asset managers, but only 8% reported improved returns. Vals Finance Agent v2 shows why automation claims need task detail: the best category score was 79.1% for earnings analysis and 34.5% for financial modeling.
This article uses primary surveys, regulator guidance, and benchmark papers checked August 30, 2026. It excludes vendor time-saved claims without a published sample and method. The result is less dramatic than a “70% faster” headline and more useful for a head of research planning a controlled pilot.
What the public evidence actually measures
| Source | Population or test set | Date | Measure relevant to research automation |
|---|---|---|---|
| Mercer | 131 global asset managers; online survey in February and March 2026 | Published May 21, 2026 | 55% integrated AI into at least one investment process; 27% piloting; 69% reported operational efficiency; 8% reported improved returns |
| Substantive Research and Aiera | 35 of the largest global asset managers | Published July 16, 2026 | 77% had organization-wide general-platform deployments; 69% named licensing the top direct-feed barrier; 37% reported four-to-six-month onboarding |
| AIMA | 150 global fund managers, about $788 billion, plus 18 investors | Published September 16, 2025 | 95% reported some generative AI use; 58% expected more use in investment processes over the following year |
| Vals Finance Agent v2 | 927 expert-reviewed analyst questions; 450-question held-out test; three runs per model | Leaderboard August 19, 2026 | Leading model 60.60% partial credit and 50.88% all-pass; category results show task difficulty |
| Deep FinResearch Bench | 100 professional reports on 25 S&P 500 companies and reports from four agents | Posted April 22, 2026 | Professional report quality 2.84 versus 2.31 for the leading agent; agent factuality ranged from 53.2% to 86.0% |
| Bank of England and FCA | 118 UK-regulated financial firms across six sectors | Published November 21, 2024 | 55% of AI use cases involved some automated decision-making; only 2% were fully autonomous; 84% of users named an accountable person |
| FINRA 2026 Regulatory Oversight Report | Supervisory observations, not a survey | 2026 report | Summarization and information extraction are the leading observed GenAI use; firms should test, log, monitor, and retain human review |
These rows cannot be averaged. AIMA counts any generative AI use, Mercer counts investment-process integration, Vals tests model performance, and the Bank/FCA survey spans banking, insurance, markets, payments, and asset management.
Which equity-research tasks are closest to reliable automation?
Finance Agent v2 uses one common harness and one scoring system across nine analyst categories. Its category ceilings are therefore more comparable than productivity claims collected from different vendors and clients.
The ordering suggests three automation tiers.
Good pilot candidates: bounded earnings and disclosure work. An analyst can specify the issuer, period, document set, and expected fields. The output can be checked against filings, earnings releases, transcripts, and consensus data. Finance Agent v2 category leaders reached 79.1% for earnings analysis and 71.3% for disclosure analysis.
Useful with structured review: adjustments and comparables. These require definition alignment, currency and period normalization, and decisions about one-time items. Category ceilings fell to 56.3% and 50.3%. The system can prepare a reviewable first pass, but a human still owns the definition and peer set.
High-risk unattended work: precedent transactions and financial modeling. Category ceilings were 36.4% and 34.5%. These tasks combine source selection, transaction or accounting conventions, several calculations, and a final judgment. They are appropriate for assistive workflows with locked assumptions and visible formulas, not a “run and publish” promise.
The full benchmark guide explains the scoring and why those percentages should not be compared with Deep FinResearch's report score.
Why adoption does not reveal analyst productivity
The current surveys leave four missing variables:
- Starting workflow. Automating manual copy-and-paste is different from replacing a mature data feed and model-update process.
- Unit of output. “One report,” “one company,” and “one earnings event” can represent radically different work.
- Review burden. Gross generation time ignores the time needed to verify sources, correct periods, and rebuild formulas.
- Failure cost. A missed source link and a wrong issuer are not equivalent. The latter should zero the task.
Mercer's 69% operational-efficiency result does not publish hours saved. It is a perception reported by respondents. Its 8% return result also uses each manager's own attribution method. Neither can be converted into a per-seat ROI without inventing assumptions.
Onboarding is part of the calculation. In the Substantive/Aiera sample, 37% said approving and onboarding models took four to six months, and another 20% reported more than six months. A 90-day pilot that begins before licensed feeds and internal permissions are ready will mostly test setup.
A worksheet for measuring automation without vendor math
Measure one recurring workflow for one quarter. An earnings-review pilot is usually cleaner than a broad “research copilot” trial because events, documents, and deadlines are observable.
| Field | How to measure it |
|---|---|
| Events in scope | Count completed earnings events, not companies licensed or seats provisioned |
| Baseline analyst minutes | Median active time from a sample of at least 10 pre-pilot events |
| AI-assisted analyst minutes | Same timer boundary, including review and corrections |
| Completion rate | Events delivered by the agreed deadline divided by events in scope |
| Source coverage | Checkable factual claims with a correct source passage divided by all checkable claims |
| Critical-error rate | Outputs with any wrong issuer, period, currency, unit, or permission breach divided by outputs reviewed |
| Rework minutes | Time spent correcting or recreating the output after first delivery |
| Fully loaded cost | Tool cost plus onboarding, data, integration, and reviewer time for the quarter |
Calculate net analyst hours saved by subtracting assisted time and rework time from the baseline minutes for each completed event, then convert the total to hours. Divide the fully loaded quarterly cost by those net saved hours to obtain cost per net hour saved.
If net hours are zero or negative, report that result. Do not remove failures from the denominator because a run timed out or returned an unusable deliverable.
A 12-task research-automation test
Use a frozen source pack and run the same tasks in every product:
- three earnings reviews, including one company with non-GAAP adjustments;
- two guidance-history updates across at least four quarters;
- two filing-change checks across consecutive annual reports;
- two three-company comparable tables with a written normalization policy;
- one model update with formulas exposed;
- one internal memo plus public-filing synthesis under a restricted role;
- one scheduled monitor that records what changed and why it triggered.
For each run, retain the prompt, product version, operator, start and finish time, output, cited passages, corrections, and failure state. Define critical errors in advance. A wrong issuer, a period mismatch, an invented citation, or unauthorized data access should fail the task even when the prose looks plausible.
FINRA's 2026 guidance makes the governance portion explicit: formal approval, comprehensive documentation, model and output testing, prompt/output logs, ongoing monitoring, and human-in-the-loop review. Those controls are not overhead external to the automation. They are part of the production workflow being measured.
What to automate first
Start where the task has a stable input, an observable output, and a cheap verification path:
- filing and transcript retrieval with passage-level citations;
- earnings tables whose periods, units, and definitions are fixed;
- guidance-history extraction with an explicit metric dictionary;
- watchlist monitoring where every alert names the trigger and source;
- first-draft comparison tables that expose all calculations.
Delay unattended modeling, recommendation generation, and trade-related decisions until the system has passed the bounded tasks. Mercer's 5% figure for autonomous or semi-autonomous recommendation/trade authority and the Bank/FCA's 2% fully autonomous use-case figure are different measures, but both support the same conservative sequence.
Where this evidence stops
We retained statistics only when an original publisher supplied a sample or test-set description, a date, and the measured construct. We removed claims such as “two hours per ticker,” “24 hours per week,” or “70% faster” when the source was a vendor page without a disclosed sample and method.
The surveys are voluntary and self-reported. The benchmarks test foundation models or agent harnesses over mostly public information, not commercial platforms with broker entitlements, internal models, permissions, or operational support. Live leaderboards will change. None of these sources reports a controlled, industry-wide estimate of net analyst hours saved.
The only defensible productivity statistic for a specific team is the one produced by that team's timed baseline and controlled run. Public data should set the test design and the caution level, not fill in the result.