Calculator
Token-Cost Optimizer
Estimate LLM token cost for a trading research loop: prompt length, model, retries, and call volume map to dollars per idea and per validated trade.
Rates from aidevhub.io model catalog (each price verified on the vendor page), Anthropic pricing, OpenAI pricing and Google AI / Gemini pricing, checked on October 3, 2026.
Runs in your browser. Nothing you enter is uploaded, and no account or API key is needed.
1. Configure your research loop
Prompt + context sent on every call.
Model calls in one research loop.
Extra calls as a share of calls (15% = 1.15× the calls).
Share of ideas that become a validated trade.
Monthly cost
$41
Claude Sonnet 5.5, cache hit 50%, 15.4× cheapest (GPT-6 Luna)
Annual run-rate: $500
Unit cost breakdown
Per call
$0.024
effective, after cache
Per idea
$0.137
5 calls × retries
Per validated trade
$0.456
30% pass rate
2. Model comparison at these inputs
| Model | Per call | Per idea | Per month | Per validated trade |
|---|---|---|---|---|
| GPT-6 Lunaopenai | $0.00155 | $0.00891 | $3 | $0.030 |
| Gemini 3.5 Flash-Litegoogle | $0.00615 | $0.035 | $11 | $0.118 |
| Gemini 3.8 Flashgoogle | $0.012 | $0.067 | $20 | $0.223 |
| Claude Haiku 4.5anthropic | $0.012 | $0.068 | $21 | $0.228 |
| GPT-5.4 Miniopenai | $0.013 | $0.073 | $22 | $0.244 |
| Claude Sonnet 5.5anthropicprimary | $0.024 | $0.137 | $41 | $0.456 |
| GPT-6.1 Solopenai | $0.031 | $0.178 | $53 | $0.594 |
| Gemini 3.1 Pro Previewgoogle | $0.034 | $0.196 | $59 | $0.652 |
| Claude Opus 5.5anthropic | $0.047 | $0.269 | $81 | $0.897 |
| GPT-5.5openai | $0.085 | $0.489 | $147 | $1.63 |
| Claude Fable 5.1anthropic | $0.116 | $0.667 | $200 | $2.22 |
| GPT-6 Astraopenai | $0.155 | $0.891 | $267 | $2.97 |
How the cost flows
effective_call = input × price_in + output × price_out
(Anthropic: cache-hit fraction priced at cache_read)
calls_per_idea = calls × (1 + retry_rate)
cost_per_idea = effective_call × calls_per_idea
cost_per_day = cost_per_idea × ideas_per_day
cost_per_val = cost_per_idea / validation_rate
cost_per_year = cost_per_day × 365Pricing last verified 2026-05-25. Cache writes and batch discounts are not modelled.
How to use it
- Pick the primary model and enter input and output tokens per call and calls per research loop (idea).
- Set the retry rate, ideas per day and the share of ideas that become validated trades.
- For Anthropic models, set the share of input tokens served from the prompt cache. It is ignored for OpenAI and Google models.
- Read the monthly cost (30 days), the annual run-rate, and the cost per call, per idea and per validated trade.
- Use the comparison table to see every model at the same inputs, sorted by monthly cost.
Questions people ask
How are model prices kept current?
From a price table in the tool, checked by hand against the Anthropic, OpenAI and Google pricing pages; the date of the last check is shown on the page. There is no automatic price feed and no history of past rates.
What's the difference between input and output token cost?
Input tokens are what you send to the model (prompt, system message, tool definitions); output tokens are what it generates. For the models in this table output costs 4 to 8 times as much as input. In research loops with long context and short answers, input cost dominates; in long-form generation it flips.
Does the tool account for prompt caching?
For Anthropic models only: the cached share of input tokens is billed at the published cache-read rate, one tenth of the base input price. The one-time cache-write premium is not added. OpenAI and Google cached-input discounts are not modelled, so the cache field changes nothing for those models.
How do I estimate token count from text?
The tool takes token counts, not text. As a rough guide, English runs about 4 characters per token; for exact counts use the provider's tokenizer or token-counting endpoint, or the Financial Document Token Estimator for filings.
What's a 'research loop' cost in practice?
The tool multiplies the cost of one call by calls per idea plus retries, then by ideas per day; a month is 30 days. Cost per validated trade divides the cost per idea by the share of ideas that become trades, so a low validation rate makes each trade expensive even when calls are cheap.
Related tools
- Playgrounds Prompt Regression Tester
Run the same prompt against multiple models (Claude 4.5/4.6/4.7, GPT-5, Gemini 2.5) with your own keys. Diff outputs, score drift, catch regressions.
- Comparators Market Data API Cost Calculator
Compute annual cost of market data across Databento, Polygon, Alpaca, Tiingo, FMP, and Alpha Vantage for your exact universe, bar resolution, and real-time needs.
- Playgrounds Agent Skill Tester for Markets
Paste a SKILL.md definition + sample input + your Anthropic API key. See structured extraction, token cost, and latency — all in your browser. No signup.
Articles
- 10 min read Financial QA LLM Benchmarks 2026: FinanceBench & Fin-RATE
Financial QA LLM benchmarks 2026: FinQA, FinanceBench, DocFinQA, and Fin-RATE leaderboard scores, plus whole-filing read costs verified 2026-06-17.
- 7 min read DeepSeek V4 for Finance 2026: SEC Filing Extraction Cost
DeepSeek V4 for finance 2026: V4.1-Flash reads a full 10-K for about $0.041 at $0.30/$1.20 per Mtok, half that off-peak, with a 1M-token window.
- 10 min read The 2026 Engineer's Guide to AI in Markets
An engineer's map of where LLMs, MCP servers, and market-data APIs fit into a 2026 trading stack — and where they still break. Direct, no hype, no grift.
Workflows that use this tool
- Workflow Plan your agent stack
Estimate first-year cost for an LLM agent — token budget, vendor selection, MCP servers.
Use it from code
The same calculation as a JavaScript module you can import. It runs where you import it, with no request, key or rate limit.
import { compute } from "https://aifinhub.io/engines/token-cost-optimizer.js";