Calculator
Financial Document Token Estimator
Financial document token estimator: price 10-K, 10-Q, 8-K and earnings call runs across 10 frontier LLMs. Context-fit + one-pass + peer synthesis cost.
Rates from aidevhub.io model catalog (each price verified on the vendor page), Anthropic pricing, OpenAI pricing and Google AI / Gemini pricing, checked on October 3, 2026.
Runs in your browser. Nothing you enter is uploaded, and no account or API key is needed.
1. Configure the document
Business + MD&A + risk factors body, excluding exhibits.
Length of the extraction or summary.
Same-size documents read together in one synthesis call.
0% for a document read once; caching pays on repeated prompts or re-reads.
One-pass run
18.0K
input tokens (GPT-6 Luna’s tokenizer), cheapest at $0.0026 on GPT-6 Luna, 110.1× spread to priciest
Priciest: $0.281 on Claude Fable 5.1. Output: 1.5K tok
2. Cost per model
| Model | Input tokens | Output tokens | Cache-read | One-pass cost | Synthesis cost | Context | Fits |
|---|---|---|---|---|---|---|---|
| GPT-6 Lunaopenai | 18.0K | 1.5K | 0 | $0.0026 | — | 1.05M | ✓ |
| Gemini 3.5 Flash-Litegoogle | 18.0K | 1.5K | 0 | $0.0091 | — | 1.048576M | ✓ |
| Gemini 3.8 Flashgoogle | 18.0K | 1.5K | 0 | $0.019 | — | 1.048576M | ✓ |
| GPT-5.4 Miniopenai | 18.0K | 1.5K | 0 | $0.020 | — | 400K | ✓ |
| Claude Haiku 4.5anthropic | 20.6K | 1.5K | 0 | $0.028 | — | 200K | ✓ |
| GPT-6.1 Solopenai | 18.0K | 1.5K | 0 | $0.051 | — | 1.05M | ✓ |
| Gemini 3.1 Pro Previewgoogle | 18.0K | 1.5K | 0 | $0.054 | — | 1.048576M | ✓ |
| Claude Sonnet 5.5anthropic | 20.6K | 1.5K | 0 | $0.056 | — | 1M | ✓ |
| Claude Opus 5.5anthropic | 20.6K | 1.5K | 0 | $0.112 | — | 1M | ✓ |
| GPT-5.5openai | 18.0K | 1.5K | 0 | $0.135 | — | 1.05M | ✓ |
| GPT-6 Astraopenai | 18.0K | 1.5K | 0 | $0.255 | — | 1.05M | ✓ |
| Claude Fable 5.1anthropic | 20.6K | 1.5K | 0 | $0.281 | — | 1M | ✓ |
Sorted cheapest first on one-pass cost. Input-token count differs slightly per provider because each tokenizer has a different char-per-token ratio.
Approximation notes
Tokenization varies per model. Estimates use published char-per-token ratios from vendor docs (Anthropic ~3.5, OpenAI ~4.0, Gemini ~4.0). For precise counts, use tiktoken (OpenAI) or Anthropic’s count_tokens endpoint. List prices last verified 2026-10-03. Cache reads use each provider’s discount; cache-write surcharges (1.25× on Anthropic) are not added.
How to use it
- Pick an archetype (10-K, 10-Q, 8-K, or earnings call) for a representative estimate, or paste the actual document text (up to about 50,000 characters).
- Read the token estimate. Pasted text uses its exact character count divided by each model's characters-per-token ratio (3.5 for Anthropic, 4.0 for the others), so it is an approximation; presets use a representative size (the 10-K body is about 18,000 to 20,600 tokens depending on the tokenizer).
- Check the Fits column — it flags, per model, whether the document overflows that model's context window (windows range from about 200K to 1M tokens). Documents that don't fit need chunking; see the SEC Filing Chunk Optimizer.
- Read the one-pass extraction cost across all ten models. Raise the cache share only for prompts or documents you send repeatedly; a filing read once is a cache miss.
- Multiply per-document cost by your monthly document volume to size the bill. For batch optimization across models, pair with the Token Cost Optimizer.
Questions people ask
What documents does it support?
Four archetype presets — 10-K, 10-Q, 8-K, and earnings call transcript — or you can paste raw text (up to about 50,000 characters) to estimate directly. There is no file upload and no page-count input: pick an archetype for a representative estimate, or paste the actual text for a precise character-based count.
How accurate is the estimate?
The paste-text path counts your characters exactly and divides by a fixed characters-per-token ratio per provider (3.5 Anthropic, 4.0 OpenAI and Google), so the token count is an approximation; tables and long numbers usually take more tokens per character than prose. For an exact count use the provider's tokenizer or token-counting endpoint. The archetype presets are representative mid-range samples, and real filings of the same type span an order of magnitude (a large-cap 10-K with full exhibits can exceed 40,000 tokens), so treat a preset number as a starting estimate, not a guarantee.
Why does token count matter for finance use?
Two reasons: (1) cost — input tokens drive LLM API cost, so you estimate before deciding to process the doc; (2) feasibility — context windows range from about 200K to 1M tokens across the models shown, and the Fits column flags, per model, whether a document overflows. Documents that don't fit need chunking.
How do I estimate without pasting the full text?
Pick the closest archetype (10-K, 10-Q, 8-K, or earnings call). Each preset carries a representative token count baked in from typical filings — the 10-K preset is around 18,000 tokens. There is no tokens-per-page or page-count path; for a precise number, paste the actual text and the tool counts characters divided by the model's tokenizer ratio.
Does it work for non-English filings?
The character-to-token ratios assume English prose (about 3.5–4 characters per token). Other languages tokenize differently — CJK text is denser, agglutinative languages sparser — so estimates for non-English text are rougher. There is no language selector: the ratio is fixed per model.
Related tools
- Generators SEC Filing Chunk Optimizer
Pick a filing archetype, tune chunk size and overlap, and see chunk count, embedding cost, and structural-boundary warnings across three chunking strategies.
- Calculators Token-Cost Optimizer
Compute the dollar cost of a trading research loop across Claude, GPT, and Gemini. Prompt length × model × retry × call volume → cost per idea and per.
- Calculators Batch vs Real-Time Cost Calculator
Jobs per day, tokens per job, model, deadline — get real-time vs batch cost side-by-side with savings estimate and batch-eligibility flag. Based.
Articles
- 14 min read Reading Financial Filings With LLMs: 2026 Playbook
A map of eight filing tasks — extraction, summarization, peer comparison, Q&A, classification, sentiment, forecasting input, compliance — with model.
- 12 min read Fine-Tuning vs RAG vs Long-Context for Filings
Decision matrix for finance LLMs: when RAG wins, when long-context wins, and when fine-tuning makes sense. Cost math from published 2026-04 vendor rates.
- 11 min read Prompt Caching Economics for Finance
How Anthropic, OpenAI, and Gemini prompt caching works on finance workloads — 5-minute TTL, hit-rate patterns, and 50-90% input savings at the right design.
Use it from code
The same calculation as a JavaScript module you can import. It runs where you import it, with no request, key or rate limit.
import { compute } from "https://aifinhub.io/engines/financial-document-token-estimator.js";