How to Extract KPIs From Earnings Transcripts With AI
A schema-first method for extracting company KPIs from earnings transcripts with period, unit, definition, speaker, source, and review controls.
Published August 20, 2026 · Updated August 30, 2026

In this article
AI can extract KPIs from earnings transcripts, but a usable record needs more than a metric name and value. Store the period, unit, scope, definition, speaker, transcript passage, document version, and review status with every observation. Then reconcile the transcript against the filing or release when the metric appears there. The process fails when a plausible number is attached to the wrong quarter, segment, or definition.
This guide is a public-source workflow, checked August 30, 2026. It does not report a product benchmark or an AllMind product run. We build AllMind and sell extraction and grid software, so our stake and our product's limitations are stated where relevant.
Use the right source for the right field
Earnings transcripts are valuable for management-defined operating metrics and the language around them. They are not always the authoritative source for reported financial-statement values.
| Source | Use it for | Do not assume |
|---|---|---|
| 10-Q or 10-K | Reported financials, accounting context, risk and footnote detail | A company KPI has a standardized tag |
| Earnings release or supplement | Quarter highlights, guidance tables, non-GAAP reconciliations | Every number uses the same basis as consensus |
| Prepared remarks | Management emphasis, operational drivers, selected KPIs | The statement is independent confirmation |
| Q&A | Clarifications, caveats, newly disclosed detail | A number in an analyst's question was accepted by management |
| Slide deck | Charts, segment bridges, management-defined views | The chart's axis and period are obvious after extraction |
The SEC explains that Inline XBRL combines human-readable filings with tagged facts and their contexts. Its EDGAR APIs expose standardized and company-specific filing data. Use that structure where it exists. Use transcript extraction for narrative KPIs and for values whose context is spoken.
Define the KPI before asking the model to find it
A metric dictionary prevents each run from interpreting the label differently.
| Field | Required value |
|---|---|
metric_id | Stable internal identifier |
canonical_name | Name used in the research model |
company_labels | Exact labels and known variants |
definition | Numerator, denominator, population, exclusions |
unit | USD, percent, users, locations, days, or other |
scope | Company, segment, geography, product |
period_type | Point-in-time, quarter, year-to-date, trailing period |
source_priority | Filing, release, transcript, presentation |
model_destination | Workbook, sheet, row, or database field |
review_owner | Person accountable for definition changes |
For “active customers,” record whether the company counts paying accounts, organizations, seats, or users. For “bookings,” record duration and cancellation policy if disclosed. A value is not comparable merely because two quarters use the same label.
The extraction record
Ask the model to return one row per observation. Do not ask for a paragraph first. Each row should preserve four groups of information:
| Part of the record | Fields to retain | Review question |
|---|---|---|
| Metric identity | Stable metric ID, company-reported label, value, and unit | Is this the requested KPI, rather than a nearby metric with a similar name? |
| Period and scope | Period start and end, period type, company or segment scope, and definition | Does the value cover the right quarter, entity, geography, and population? |
| Evidence | Speaker, source document and version, exact passage, URL, and material qualifiers | Can a reviewer open the source and see the statement in context? |
| Workflow status | Exact, derived, ambiguous, or absent; calculation inputs when derived; reviewer note | Should the row load automatically, wait for review, or remain explicitly missing? |
The status field is deliberately narrow. “Exact” means the source directly states the value on the requested basis. “Derived” requires visible inputs and a formula. “Ambiguous” means the value, period, scope, or definition is unclear. “Absent” prevents the system from filling a gap with a nearby metric or model memory.
Run extraction in four passes
Pass 1: locate candidate passages
Use exact labels, label variants, and semantic retrieval. Return more passages than the final answer needs. Include surrounding sentences and the speaker. This is a retrieval pass, so do not deduplicate values yet.
Pass 2: normalize context
For each candidate, identify period, unit, scope, and definition. Reject numbers stated by an analyst unless management confirms them. Keep ranges as ranges. Keep “approximately,” “more than,” “constant currency,” and similar qualifiers in a separate field and in the quoted passage.
Pass 3: reconcile authoritative documents
If the release or filing states the same KPI, compare values and definitions. Do not silently replace the transcript value. Record both sources and explain any difference, such as rounding, later correction, or a different reporting scope.
Pass 4: validate and load
Run deterministic checks before a row enters a model or database:
- period end falls within the requested fiscal quarter;
- unit matches the metric dictionary;
- percentages and basis points are not interchanged;
- segment values are not labeled company-wide;
- range lower bound is less than upper bound;
- a derived value exposes inputs and formula;
- source URL and passage are present;
- definition changes create a new version;
- one reviewer owns each ambiguous or changed row.
A small test reveals most extraction defects
Build an answer key with 30 observations from three companies. Include ten easy metrics, ten metrics with changing labels or definitions, and ten traps. Useful traps include a number in an analyst question, a range, a year-to-date value, a currency conversion, a corrected transcript, a segment-only figure, and a metric discontinued this quarter.
Report results by failure type:
| Measure | Denominator | What counts as failure |
|---|---|---|
| Retrieval recall | Answer-key observations | Required passage was never returned |
| Field accuracy | Fields in retrieved observations | Value, unit, period, or scope is wrong |
| Qualifier preservation | Qualified observations | Material qualifier is missing |
| Definition integrity | Metrics with definition changes | Old and new versions are merged |
| Citation validity | Output rows | Link does not open the supporting passage |
| Abstention | Deliberately absent metrics | System supplies a value anyway |
Do not publish one blended accuracy percentage unless the weighting and answer key are visible. A system can score well on common values and fail every changed definition.
How product surfaces differ
FactSet Transcript Assistant supports transcript questions, summaries, Q&A review, and sentiment views within the Workstation. AlphaSense Transcript Summaries links summary items to transcript passages. Quartr AI Chat searches first-party IR material, which can help reconcile transcript, filing, report, and slides. These pages establish product surfaces, not extraction accuracy.
Our Grids page makes AllMind a candidate for running a fixed metric schema across companies. The meaningful limitation sits on our side: we do not publish recall, definition-handling, or source-coverage results, so the interface description alone verifies none of them. The 30-observation answer-key test should be run on the buyer's actual tickers before adoption.
A general assistant is workable for one transcript if policy permits the upload and the analyst can inspect the full source. It becomes difficult to govern across a coverage list because the team must own document ingestion, versioning, permissions, schemas, repeat runs, and audit logs.
Management commentary belongs beside the KPI, not inside it
Store management's explanation separately from the value. A useful commentary record includes driver, direction, time horizon, stated evidence, caveat, and source passage. This enables questions such as “Which companies attributed lower gross margin to mix this quarter?” without changing the underlying metric table.
Do not convert commentary into a numeric forecast unless the method and assumptions are explicit. “We expect improvement in the back half” has analytical value, but it is not a basis for selecting a point estimate on its own.
What remains an analyst decision
An analyst approves the canonical definition, decides whether two periods are comparable, chooses which source controls a model row, interprets management commentary, and determines materiality. AI can make the evidence packet faster to build and easier to rerun. It should leave a visible exception when the evidence is incomplete.
The reviewer also owns every override and definition version, so the next quarter can reproduce the decision instead of rediscovering it.
Evidence behind the extraction method
This workflow uses SEC documentation for filing and Inline XBRL structure plus official product pages from FactSet, AlphaSense, and Quartr, and our own AllMind pages, checked August 30, 2026. Competitor product statements are vendor-reported. AllMind statements are our own first-party claims. We did not execute the 30-observation extraction test or compare products. Buyers should verify document coverage, transcript versions and corrections, field-level passage links, entitlements, retention, schema export, and abstention behavior in their own environment.