Comparator
Model Selector for Finance
Model selector finance: pick the right LLM for extract, summarize, forecast, compare, rank, synthesize — cost, latency, context, quality axes.
Rates from aidevhub.io model catalog (each price verified on the vendor page), Anthropic pricing, OpenAI pricing and Google AI / Gemini pricing, checked on October 3, 2026.
Runs in your browser. Nothing you enter is uploaded, and no account or API key is needed.
Recommended model
GPT-6 Luna
$4/mo
openai, haiku tier, 1.1M ctx, thinking
Reference workload: 6,000 in / 1,200 out × 3,000 calls/mo.
1. Configure your task profile
Reference workload for cost fit: 6,000 in / 1,200 out × 3,000 calls/mo.
2. Top 3 recommendations
GPT-6 Luna
haiku tier, 1.1M ctx, thinking
OpenAI fast-tier model: 1.05M context at $0.1 in / $0.5 out per 1M tokens (prompts up to 272K input tokens). Reference monthly spend at this tool's default workload is ~$4, within the $50/mo budget. Published context window 1.1M covers the 32K–200K requirement. Vendor positions the Haiku tier for summarize workloads.
Vendor pricingGemini 3.8 Flash
sonnet tier, 1.0M ctx, thinking
Google mid-tier model: 1.05M context at $0.75 in / $3.75 out per 1M tokens (promotional price through December 31, 2026). Reference monthly spend at this tool's default workload is ~$27, within the $50/mo budget. Published context window 1.0M covers the 32K–200K requirement. Vendor positions the Sonnet tier for summarize workloads.
Vendor pricingGPT-5.4 Mini
sonnet tier, 400K ctx, thinking
OpenAI mid-tier model: 400K context at $0.75 in / $4.5 out per 1M tokens. Reference monthly spend at this tool's default workload is ~$30, within the $50/mo budget. Published context window 400K covers the 32K–200K requirement. Vendor positions the Sonnet tier for summarize workloads.
Vendor pricingPublished-rate-based; verify with your own eval harness (see D1 — Eval harness for finance LLMs).
3. Full ranked list with why-not notes
Passes all gates; simply outranked by a model with better combined fit.
Passes all gates; simply outranked by a model with better combined fit.
Passes all gates; simply outranked by a model with better combined fit.
Passes all gates; simply outranked by a model with better combined fit.
Passes all gates; simply outranked by a model with better combined fit.
Over the chosen cost budget at default workload.
Over the chosen cost budget at default workload.
Over the chosen cost budget at default workload.
Over the chosen cost budget at default workload.
Over the chosen cost budget at default workload.
Over the chosen cost budget at default workload.
Over the chosen cost budget at default workload.
4. Per-axis comparison (all models)
| Model | Input $/1M | Output $/1M | Context | Thinking | Ref $/mo | Cost | Latency | Ctx | Capability |
|---|---|---|---|---|---|---|---|---|---|
| GPT-6 Luna | $0.10 | $0.50 | 1.1M | yes | $4 | pass | pass | pass | pass |
| Gemini 3.8 Flash | $0.75 | $3.75 | 1.0M | yes | $27 | pass | pass | pass | pass |
| GPT-5.4 Mini | $0.75 | $4.50 | 400K | yes | $30 | pass | pass | pass | pass |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 1.0M | yes | $14 | pass | pass | pass | pass |
| Claude Haiku 4.5 | $1.00 | $5.00 | 200K | yes | $36 | pass | pass | pass | pass |
| Claude Sonnet 5.5 | $2.00 | $10.00 | 1M | yes | $72 | fail | pass | pass | pass |
| GPT-6.1 Sol | $2.00 | $10.00 | 1.1M | yes | $72 | fail | pass | pass | pass |
| Claude Fable 5.1 | $10.00 | $50.00 | 1M | yes | $360 | fail | fail | pass | fail |
| Claude Opus 5.5 | $4.00 | $20.00 | 1M | yes | $144 | fail | fail | pass | fail |
| GPT-6 Astra | $10.00 | $50.00 | 1.1M | yes | $360 | fail | fail | pass | fail |
| GPT-5.5 | $5.00 | $30.00 | 1.1M | yes | $198 | fail | fail | pass | fail |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 | 1.0M | yes | $79 | fail | fail | pass | fail |
Hover or focus a cell for the axis note. Rates and context windows from vendor pricing pages, checked 2026-05-25.
Scoring framework
score = cost_match + latency_match + context_match
+ capability_bonus + quality_boost
cost_match : 0 if monthly estimate > budget ceiling
latency_match : 0 if tier slower than latency budget
context_match : 0 if context window < required
capability : bonus if task ∈ model.best_for
quality : boost flagship tiers when quality = highDeliberately no accuracy numbers. See the framework article for deeper rationale.
How to use it
- Pick the task type, latency budget, monthly cost budget, context-size need and quality sensitivity from the five dropdowns.
- Read the recommended model and its estimated monthly spend at the reference workload (6,000 input and 1,200 output tokens per call, 3,000 calls a month).
- Check the top three cards and the full ranked list. A model that fails a hard gate (cost, latency or context) ranks below every model that passes.
- Use the per-axis table to compare published input and output prices per 1M tokens, context windows and which gates each model passes.
- Treat the ranking as a shortlist: it uses published prices and vendor positioning only, so confirm quality on your own eval set before switching.
Questions people ask
How does the selector recommend a model?
It scores every model in its table on five axes: whether the monthly cost at a fixed reference workload fits your budget, whether the model's tier fits your latency budget, whether its context window covers your need, whether the vendor positions it for your task type, and how it fits your quality sensitivity. Cost, latency and context are hard gates: a model that fails one ranks after every model that passes. No accuracy benchmark enters the score.
Why is Claude Sonnet recommended for most workloads?
It is not recommended by default. At the page's starting inputs (summarize, under 5 seconds, $50 a month, 32K to 200K context, medium quality) the cheapest models that pass every gate rank first. Sonnet-tier models gain points when quality sensitivity is medium or high and when the vendor positions that tier for your task; change the inputs and the ranking changes with them.
When does the selector recommend a GPT model over Claude?
When its published price, context window, latency tier and task positioning score higher for your inputs. The selector does not model tool-use maturity, rate limits or tail latency; it uses only the published per-token prices, context windows and vendor tier positioning shown in the per-axis table.
Are open-source models considered?
No. The table covers hosted Anthropic, OpenAI and Google models with published per-token API prices. Self-hosted open-weight models have no comparable published per-token price, so they are not scored.
How often is the selector updated?
Prices and context windows are checked by hand against the vendor pricing pages, and the date of the last check is shown on the page. No scheduled eval suite sits behind the ranking.
Related tools
- Calculators Token-Cost Optimizer
Compute the dollar cost of a trading research loop across Claude, GPT, and Gemini. Prompt length × model × retry × call volume → cost per idea and per.
- Calculators Financial Document Token Estimator
Paste a 10-K, 10-Q, 8-K or earnings transcript and see token count + one-pass extraction cost across ten frontier LLMs, with cache-hit toggle.
- Calculators Batch vs Real-Time Cost Calculator
Jobs per day, tokens per job, model, deadline — get real-time vs batch cost side-by-side with savings estimate and batch-eligibility flag. Based.
Articles
- 10 min read Financial QA LLM Benchmarks 2026: FinanceBench & Fin-RATE
Financial QA LLM benchmarks 2026: FinQA, FinanceBench, DocFinQA, and Fin-RATE leaderboard scores, plus whole-filing read costs verified 2026-06-17.
- 12 min read Model Selection Framework for Finance Tasks
A task × latency × cost × context decision tree for finance LLM workloads. Ten concrete scenarios mapped to tier bands. Grounded in published pricing, not.
- 11 min read Thinking Tokens for Finance Tasks
When extended-thinking and reasoning-effort modes earn their 3-10x cost tax on finance workloads — and when they are a silent drain on the budget.
Use it from code
The same calculation as a JavaScript module you can import. It runs where you import it, with no request, key or rate limit.
import { compute } from "https://aifinhub.io/engines/model-selector-finance.js";