Skip to main content
AI Fin Hub

Generator

SEC Filing Chunk Optimizer

SEC filing chunk sizing + 10-K chunking cost calculator. Pick archetype, chunk size, overlap, strategy, and embedding model.

Runs in your browser. Nothing you enter is uploaded, and no account or API key is needed.

Education, not investment advice. Past performance does not predict future results. How we check our numbers.

1. Configure chunk strategy

Chunking strategy
Chunk size (tokens)
Overlap between chunks (%)

Share of each chunk repeated at the start of the next one.

Queries to embed

Search queries embedded against the index, about 40 tokens each.

Estimated chunks for one 10-K (full body)

138 chunks

Average 1,021 tokens each at a 1,024-token target, 15% overlap, structural split; 140,898 tokens embedded.

Ingest $0.002818, 100 queries $0.000080, text-embedding-3-small

Strategy note

Respects Items / section headers / speaker turns. Preserves table blocks by keeping heading+table together. Chunk sizes are uneven but semantically clean.

No structural warnings at these settings. Still run a retrieval eval before production — heuristics can't replace ground truth.

Archetype reference: Form 10-K business + risk + MD&A + financials. ~12 Items. Dense tables in Item 7 / 8.

Detail

Total chunks

138

Avg tokens/chunk

1,021

min 1,024, max 1,024

Ingest cost (once)

$0.002818

text-embedding-3-small

Query cost (100 re-embeds)

$0.000080

Tokens embedded

140,898

3. Compare strategies (same archetype + chunk size)

StrategyChunksAvg tokMin / MaxIngest costTradeoff
structuralselected1381,0211,024 / 1,024$0.002818Highest fidelity; uneven chunk sizes.
recursive1381,021614 / 1,024$0.002818Cheap + deterministic; blind to tables.
semantic1321,061409 / 1,433$0.002801Coherent prose groups; variable sizes, higher compute.

How the estimate works

stride        = chunk_size × (1 − overlap_pct)
base_count    = ceil(total_tokens / stride)
structural    → max(boundary_count, base_count)
recursive     → base_count
semantic      → ceil(base_count × 0.95), wider size variance
ingest_cost   = tokens_embedded × $/M_tokens
query_cost    = (40 × n_queries) × $/M_tokens

Pricing verified 2026-04-23.

How to use it

  1. Pick a filing archetype (10-K body, MD&A, notes to the financial statements, earnings-call transcript) and an embedding model.
  2. Choose a chunking strategy (structural, recursive or semantic), the chunk size in tokens and the overlap.
  3. Read the estimated chunk count for one filing, the average chunk size, the tokens embedded and the one-time embedding cost.
  4. Read the warnings: small chunks on table-heavy filings, high overlap, semantic chunking under 1K tokens and others.
  5. Compare the three strategies at the same settings in the table, then confirm the choice with a retrieval eval on your own filings.

Questions people ask

What chunks does the tool produce?

None: it is an estimator. From a representative token count for each filing type it estimates how many chunks each strategy produces at your chunk size and overlap, their size range, the tokens embedded and the embedding cost. It does not download, parse or split an actual filing.

Why not use a generic text splitter?

A generic recursive splitter ignores document structure, so it can cut a table mid-row or split an Item section; the tool warns about this for table-heavy archetypes. Structural chunking that follows Items, headings and speaker turns keeps those units intact, at the cost of uneven chunk sizes.

How big are the chunks?

You set the target, from 256 to 8,192 tokens, and an overlap from 0 to 30%; the page starts at 1,024 tokens with 15% overlap. The estimate shows the average and the expected minimum and maximum chunk size for the chosen strategy. Larger chunks mean fewer, more diluted embeddings.

Does it handle XBRL?

No. The estimate works from token counts of the narrative filing and does not read XBRL. For exact reported figures, pull the XBRL facts from SEC EDGAR directly rather than from chunked text.

What's the optimal chunk strategy for filing Q&A?

There is no universal answer; it depends on your questions and your retriever. Structural chunks of roughly 512 to 1,024 tokens with modest overlap are a common starting point for 10-K question answering. Measure recall on your own question set before settling; this tool shows the cost side of the choice.

  • Calculators Financial Document Token Estimator

    Paste a 10-K, 10-Q, 8-K or earnings transcript and see token count + one-pass extraction cost across ten frontier LLMs, with cache-hit toggle.

  • Playgrounds Structured Schema Validator for Finance

    Paste LLM JSON output and validate against four pre-built finance schemas — research output, trade decision, risk snapshot, peer comparison — with sanity.

  • Playgrounds Hallucination Detector

    Paste a source document + an LLM's extraction. Every numeric claim in the output is checked against the source. Client-side. Catches silent fabrication.

All articles

Use it from code

The same calculation as a JavaScript module you can import. It runs where you import it, with no request, key or rate limit.

import { compute } from "https://aifinhub.io/engines/sec-filing-chunk-optimizer.js";

Input and output contract and the guide for agents.