How Accurate Are AI Earnings Call Summaries?
There is no universal accuracy rate. This guide shows how to audit factual support, omissions, qualifiers, attribution, and source traceability.
Published August 28, 2026 · Updated August 30, 2026

In this article
There is no defensible universal accuracy percentage for AI earnings-call summaries. Accuracy changes with the transcript version, prompt, model, source packet, claim type, and scoring rule. A useful audit separates factual support, material omissions, qualifier preservation, speaker attribution, period and unit accuracy, and citation validity. The result should report errors by task, with the tested inputs and date, not compress them into a marketing number.
The earlier version of this article claimed a 50-claim test that could not be reproduced from committed inputs and outputs. That claim has been removed. This replacement provides a protocol a research team can run and publish honestly. It uses public research and vendor documentation checked August 30, 2026; no AllMind output or comparative product score is presented.
“Accurate” contains at least six different questions
| Dimension | Audit question | Example failure |
|---|---|---|
| Factual support | Does the source passage support the claim? | Summary invents a cause absent from the call |
| Numeric context | Are value, unit, period, basis, and range correct? | Basis points become percent; quarter becomes full year |
| Qualifier preservation | Are hedges and conditions retained? | “Approximately” or “subject to” disappears |
| Attribution | Is the right speaker attached to the statement? | Analyst assertion is presented as management guidance |
| Completeness | Are all decision-material items represented? | Guidance cut or discontinued KPI is omitted |
| Traceability | Can a reviewer open the exact source? | Citation points to a whole document or broken link |
A summary can be factually supported and still be poor because it omits the only item relevant to the thesis. It can also reproduce every number while losing a condition that changes the interpretation.
Grounded summarization research does not supply an earnings-call rate
Research on factual consistency shows why task definition matters. The Vectara hallucination leaderboard evaluates grounded summarization on specified datasets and model versions. Its figures should not be transferred to earnings calls, which contain speaker turns, financial periods, ranges, prepared text, and Q&A.
Academic work also distinguishes factual consistency from broad summary quality. FActScore breaks generated text into atomic facts and checks support, while TRUE evaluates factual consistency across several datasets and metrics. These methods inform an audit design. They do not establish that a particular transcript product has a given accuracy rate.
Vendor pages make narrower claims. AlphaSense Transcript Summaries documents click-through from summary items to matching transcript passages. FactSet Transcript Assistant describes transcript questions, summaries, and sentiment views. Aiera publishes a vendor-sponsored transcript benchmark focused on transcript production, including speaker identification and verbatim accuracy. None of these pages supplies a common, independent benchmark of generated earnings summaries across the products.
Build the test packet before generating a summary
Use two or more completed earnings events. Include one clean call and one difficult call with a range, multiple segments, changed definitions, and a corrected transcript.
Freeze the company and event, fiscal period, transcript provider and version, and the transcript's publication and correction timestamps. Preserve the earnings release URL, filing URL and accession number, and presentation URL. The run record should also identify the model version, the exact instructions and generation settings, the run timestamp, and the operator.
The release and filing help resolve reported values. The transcript is the source for spoken claims. Keep the exact generated output. Without those artifacts, another reviewer cannot reproduce the audit.
Create an answer key independently
Before reading the generated summary, two reviewers should build a list of decision-material source items. Include reported results, guidance, changes in guidance assumptions, newly disclosed KPIs, discontinued metrics, capital allocation, and material Q&A.
The answer key row should look like this:
| Item ID | Atomic source statement | Source and passage | Claim type | Required qualifier | Materiality reason |
|---|---|---|---|---|---|
| A01 | Management narrowed full-year revenue guidance | Reviewed transcript, management answer, exact timestamp and passage | Guidance | Retain the currency, range, period, and stated conditions | The change may affect the base-case revenue estimate |
This is an illustrative row, not a claim about a named company.
Reviewers can disagree about materiality. Record and resolve that disagreement before scoring omissions. This prevents the generated summary from defining its own test.
Keep the pre-adjudication labels as well as the final answer key. Reviewer agreement is evidence about how well defined the rubric is. A low agreement rate may mean the claim category or materiality rule needs revision; it should not be hidden by a final consensus label.
Score atomic claims, then score omissions
Split every summary sentence into atomic claims. One sentence can contain a supported figure, an unsupported cause, and a wrong period. Score each separately.
| Summary claim ID | Claim | Supported? | Period and unit correct? | Qualifier retained? | Attribution correct? | Citation valid? | Error note |
|---|---|---|---|---|---|---|---|
| S01 | Full-year revenue guidance narrowed | Yes | Yes | Yes | Yes | Yes | Direct passage supports the claim |
Then match the answer key against the summary:
| Answer-key item | Present in summary? | Complete enough for the intended reader? | Omission severity |
|---|---|---|---|
| A01 | Yes | Partial | Material |
These filled rows demonstrate the audit structure. A real test must retain every claim, disagreement, and error rather than copy the example judgments.
Do not let ten correct low-value facts cancel one omitted guidance change. Report the raw counts and severity distribution.
Publish a scorecard readers can reconstruct
A compact report should publish the following measures with both the count and its denominator:
| Measure | Reporting requirement |
|---|---|
| Atomic claims reviewed | State the total and identify the saved summary version |
| Fully, partially, and unsupported claims | Report each category separately against all reviewed claims |
| Wrong period, unit, scope, or range | Report the error count against all claims where those fields apply |
| Material qualifiers lost | Use only claims with a material qualifier as the denominator |
| Attribution errors | Use attributed claims as the denominator |
| Valid passage citations | Report valid links against all checkable claims |
| Material answer-key items omitted | Report omissions against all items judged material before scoring |
| Reviewer agreement | Report agreement before adjudication against all independently labeled decisions |
If a combined score is necessary, publish the formula and weights. Keep the task-level table next to it. A procurement team may reasonably set citation failure or a wrong guidance period as an automatic fail.
Test common failure modes deliberately
Add cases that reveal whether the system understands the source:
- a guidance range with several conditions;
- an analyst states a number and management declines to confirm it;
- one speaker corrects another later in the call;
- a metric is year-to-date while the prompt asks for the quarter;
- segment and company-wide values share a label;
- the live transcript contains an error fixed in the reviewed version;
- a topic is absent and the correct output is “not stated”;
- a current-quarter claim resembles language from the prior quarter.
Rerun the same packet after any material model, prompt, retrieval, or transcript-provider change. Version drift can break a previously acceptable workflow.
What citations solve, and what they do not
Passage links reduce verification cost and make errors discoverable. They do not prove the model retrieved every important passage, selected the right source version, or preserved a qualifier. A cited claim can still be wrong if the passage does not support it.
The product requirement should be: every checkable claim opens the exact passage, with speaker, event, and transcript version. Exported summaries should retain those links or stable source IDs.
Our own Document Search is built around passage-linked research, and our data layer includes live Aiera calls and transcripts inside an estate of 750M+ documents and 6,800+ premium data sources licensed from 100+ providers and partners. The same workflow can cross-check the call against filings, reported financials, consensus estimates, guidance history, broker research and Expert Insights instead of treating a transcript as an isolated upload. We publish no reproducible earnings-summary accuracy benchmark, and neither does any vendor named here, so the same frozen packet, answer key, and scoring rules should apply to AllMind and every alternative.
A safe operating rule
Treat an AI earnings summary as a navigation and drafting aid until a reviewer checks reported values, guidance, ranges, conditions, attribution, and material omissions. Teams can reduce review effort once repeated, versioned tests show which claim types are reliable in their own corpus. They should preserve spot checks and failure monitoring after deployment.
Any material model, retrieval, transcript-provider, or prompt change should return the workflow to a fuller audit before review effort is reduced again.
Evidence and correction note
This guide uses public grounded-summarization research, official product documentation from AlphaSense, FactSet, and Aiera, and our own AllMind pages, checked August 30, 2026. Vendor statements, ours included, are not independent comparative evidence. No claim-level product run was performed for this revision. The prior unreproducible result was removed, and the replacement protocol is intended to make any future test reproducible, including inputs, raw outputs, limitations, and reviewer disagreements.