Skip to main content
AI Fin Hub

Calculator

Earnings-Call Summarization Cost Calculator

Compute LLM cost per stock per quarter to summarize earnings transcripts across Sonnet, Opus, GPT-5.5, Gemini 2.5 Pro/Flash. Cache-hit-rate aware.

Rates from aidevhub.io model catalog (each price verified on the vendor page), Anthropic pricing, OpenAI pricing and Google AI pricing, checked on October 3, 2026.

Runs in your browser. Nothing you enter is uploaded, and no account or API key is needed.

Education, not investment advice. Past performance does not predict future results. How we check our numbers.

Workload

Tickers per quarter

One earnings call per ticker per quarter.

Input tokens per call

Transcript plus prompt. A one-hour call is roughly 10,000 to 15,000 tokens.

Summary tokens (output)
Input tokens served from cache

A fresh transcript is a cache miss; only a stable prompt prefix and re-reads hit. 10% ≈ a 1,500-token cached prompt on a 15,000-token call.

Attempts per transcript

Re-runs or chained passes over the same transcript.

Pricing snapshot as of 2026-10-03. List prices, no batch discount.

Cost per stock per quarter

$0.04

Claude Sonnet 5.5: $0.14 per stock per year, $7.06 a year for 50 tickers.

Ranked by annual cost (full universe)

Model$/stock/qtr$/stock/yrUniverse/yr
OpenAI, GPT-6 Luna$0.002$0.01$0.35
Google, Gemini 3.5 Flash-Lite$0.006$0.02$1.22
Google, Gemini 3.8 Flash$0.013$0.05$2.65
OpenAI, GPT-5.4 Mini$0.014$0.06$2.77
Anthropic, Claude Haiku 4.5$0.018$0.07$3.53
OpenAI, GPT-6.1 Sol$0.035$0.14$7.03
Anthropic, Claude Sonnet 5.5$0.035$0.14$7.06
Google, Gemini 3.1 Pro Preview$0.037$0.15$7.38
Anthropic, Claude Opus 5.5$0.070$0.28$14.06
OpenAI, GPT-5.5$0.092$0.37$18.45
Anthropic, Claude Fable 5.1$0.175$0.70$35.08
OpenAI, GPT-6 Astra$0.177$0.71$35.30

Caveats

Estimates assume average tokens; real transcripts vary several-fold in length. Cache reads only apply to a stable prompt prefix (system prompt, schema, few-shot examples) or to a transcript read more than once; a transcript's first read is cold input. Cached tokens are priced at each provider's cache-read rate; Anthropic's 1.25× cache-write surcharge and Gemini cache storage fees are not included.

How to use it

  1. Pick the model for the headline number and enter tickers per quarter and input tokens per call (transcript plus prompt; a one-hour call is roughly 10,000 to 15,000 tokens). The Financial Document Token Estimator can count a pasted transcript.
  2. Set the summary length in output tokens (500 to 1,500 for a structured summary) and the number of passes per transcript.
  3. Set the share of input tokens served from cache: only a repeated prompt prefix or re-read transcripts qualify, so keep it low (around 10%) for single-pass summaries.
  4. Read the cost per stock per quarter, per stock per year and for the whole universe per year.
  5. Compare models in the ranked table, which uses the same inputs for every model. Pricing is list price as of the snapshot date shown under the inputs.

Questions people ask

What's the per-call cost range?

It depends on transcript length and model. At list prices with no caching, a 12,000-token call summarized into 1,000 tokens costs about $0.051 on Claude Sonnet 4.6 ($3/$15 per 1M tokens), $0.085 on Claude Opus 4.8, $0.090 on GPT-5.5 and $0.0016 on Gemini 2.5 Flash-Lite. For 500 companies a quarter that is about $25.50 on Sonnet, $42.50 on Opus and $0.80 on Flash-Lite per quarter. The ranked table shows every model for your own inputs.

Why does long-context model selection matter?

For single-call summaries it rarely does: even an 80-minute call with Q&A is a few tens of thousands of tokens, well inside the 200K to 1M-token windows of every model in the table. The tool assumes one call per transcript and does not model chunking. Context size matters again if you batch several transcripts or a full quarter of filings into one prompt.

How does prompt caching help here?

Caching only discounts tokens that repeat across calls: your system prompt, extraction schema and examples, or a transcript you query more than once. The cache-hit input is the share of all input tokens that does repeat, priced at each provider's cache-read rate (0.1× input on Anthropic). For 500 calls a quarter, caching a 2,000-token prompt on a 14,000-token call saves about $2.70 a quarter on Sonnet 4.6 and $4.50 on Opus 4.8: worth doing, but small next to the transcript itself, which is a cache miss on first read.

Can I summarize calls in real-time?

The tool prices tokens, not latency or transcription. A summary call itself returns quickly; in practice the wait is for the transcript to be published. Summarizing live audio adds a speech-to-text step whose cost is not included here.

What if I need to extract structured data, not summary?

Use the same calculator with a larger output-token setting: structured extraction of guidance changes, KPIs and risk mentions usually needs longer outputs and sometimes a second validation pass (set attempts to 2). There is no separate extraction mode; the cost scales linearly with output tokens and attempts.

  • Calculators Financial Document Token Estimator

    Paste a 10-K, 10-Q, 8-K or earnings transcript and see token count + one-pass extraction cost across ten frontier LLMs, with cache-hit toggle.

  • Calculators Token-Cost Optimizer

    Compute the dollar cost of a trading research loop across Claude, GPT, and Gemini. Prompt length × model × retry × call volume → cost per idea and per.

  • Comparators Model Selector for Finance

    Input task, latency budget, cost budget, context size, and quality sensitivity; get ranked model recommendations with rationale — grounded in published.

  • 12 min read Prompt Patterns for Earnings Calls

    Five copy-paste patterns — speaker attribution, hedged-guidance confidence, multi-quarter delta, risk aggregator, forward-outlook separator.

  • 9 min read Earnings Call Summarisation

    Cost and architecture guide for earnings-call summarisation across eight production LLMs: verifiable 2026 vendor pricing and qualitative failure modes.

  • 9 min read Earnings Call Summarization: 250 Tickers, Nine Models

    Engine returns $0.91/year (Gemini 2.5 Flash-Lite) to $42.80/year (Opus 4.8) for 250-ticker quarterly coverage. Operator review time dominates API spend by 200×.

All articles

Use it from code

The same calculation as a JavaScript module you can import. It runs where you import it, with no request, key or rate limit.

import { compute } from "https://aifinhub.io/engines/earnings-call-summarization-cost.js";

Input and output contract and the guide for agents.