Skip to main content
AI Fin Hub

Comparator

Model Selector for Finance

Model selector finance: pick the right LLM for extract, summarize, forecast, compare, rank, synthesize — cost, latency, context, quality axes.

Rates from aidevhub.io model catalog (each price verified on the vendor page), Anthropic pricing, OpenAI pricing and Google AI / Gemini pricing, checked on October 3, 2026.

Runs in your browser. Nothing you enter is uploaded, and no account or API key is needed.

Education, not investment advice. Past performance does not predict future results. How we check our numbers.

Recommended model

GPT-6 Luna

$4/mo

openai, haiku tier, 1.1M ctx, thinking

Reference workload: 6,000 in / 1,200 out × 3,000 calls/mo.

1. Configure your task profile

Reference workload for cost fit: 6,000 in / 1,200 out × 3,000 calls/mo.

2. Top 3 recommendations

#1openai

GPT-6 Luna

haiku tier, 1.1M ctx, thinking

OpenAI fast-tier model: 1.05M context at $0.1 in / $0.5 out per 1M tokens (prompts up to 272K input tokens). Reference monthly spend at this tool's default workload is ~$4, within the $50/mo budget. Published context window 1.1M covers the 32K–200K requirement. Vendor positions the Haiku tier for summarize workloads.

Vendor pricing
#2google

Gemini 3.8 Flash

sonnet tier, 1.0M ctx, thinking

Google mid-tier model: 1.05M context at $0.75 in / $3.75 out per 1M tokens (promotional price through December 31, 2026). Reference monthly spend at this tool's default workload is ~$27, within the $50/mo budget. Published context window 1.0M covers the 32K–200K requirement. Vendor positions the Sonnet tier for summarize workloads.

Vendor pricing
#3openai

GPT-5.4 Mini

sonnet tier, 400K ctx, thinking

OpenAI mid-tier model: 400K context at $0.75 in / $4.5 out per 1M tokens. Reference monthly spend at this tool's default workload is ~$30, within the $50/mo budget. Published context window 400K covers the 32K–200K requirement. Vendor positions the Sonnet tier for summarize workloads.

Vendor pricing

Published-rate-based; verify with your own eval harness (see D1 — Eval harness for finance LLMs).

3. Full ranked list with why-not notes

#1GPT-6 Lunaopenai, haiku
score 86

Passes all gates; simply outranked by a model with better combined fit.

#2Gemini 3.8 Flashgoogle, sonnet
score 85

Passes all gates; simply outranked by a model with better combined fit.

#3GPT-5.4 Miniopenai, sonnet
score 84

Passes all gates; simply outranked by a model with better combined fit.

#4Gemini 3.5 Flash-Litegoogle, haiku
score 83

Passes all gates; simply outranked by a model with better combined fit.

#5Claude Haiku 4.5anthropic, haiku
score 76

Passes all gates; simply outranked by a model with better combined fit.

#6Claude Sonnet 5.5anthropic, sonnetfails a gate
score 13

Over the chosen cost budget at default workload.

#7GPT-6.1 Solopenai, sonnetfails a gate
score 13

Over the chosen cost budget at default workload.

#8Claude Fable 5.1anthropic, opusfails a gate
score 0

Over the chosen cost budget at default workload.

#9Claude Opus 5.5anthropic, opusfails a gate
score 0

Over the chosen cost budget at default workload.

#10GPT-6 Astraopenai, opusfails a gate
score 0

Over the chosen cost budget at default workload.

#11GPT-5.5openai, opusfails a gate
score 0

Over the chosen cost budget at default workload.

#12Gemini 3.1 Pro Previewgoogle, opusfails a gate
score 0

Over the chosen cost budget at default workload.

4. Per-axis comparison (all models)

ModelInput $/1MOutput $/1MContextThinkingRef $/moCostLatencyCtxCapability
GPT-6 Luna$0.10$0.501.1Myes$4passpasspasspass
Gemini 3.8 Flash$0.75$3.751.0Myes$27passpasspasspass
GPT-5.4 Mini$0.75$4.50400Kyes$30passpasspasspass
Gemini 3.5 Flash-Lite$0.30$2.501.0Myes$14passpasspasspass
Claude Haiku 4.5$1.00$5.00200Kyes$36passpasspasspass
Claude Sonnet 5.5$2.00$10.001Myes$72failpasspasspass
GPT-6.1 Sol$2.00$10.001.1Myes$72failpasspasspass
Claude Fable 5.1$10.00$50.001Myes$360failfailpassfail
Claude Opus 5.5$4.00$20.001Myes$144failfailpassfail
GPT-6 Astra$10.00$50.001.1Myes$360failfailpassfail
GPT-5.5$5.00$30.001.1Myes$198failfailpassfail
Gemini 3.1 Pro Preview$2.00$12.001.0Myes$79failfailpassfail

Hover or focus a cell for the axis note. Rates and context windows from vendor pricing pages, checked 2026-05-25.

Scoring framework

score = cost_match + latency_match + context_match
      + capability_bonus + quality_boost
cost_match    : 0 if monthly estimate > budget ceiling
latency_match : 0 if tier slower than latency budget
context_match : 0 if context window < required
capability    : bonus if task ∈ model.best_for
quality       : boost flagship tiers when quality = high

Deliberately no accuracy numbers. See the framework article for deeper rationale.

How to use it

  1. Pick the task type, latency budget, monthly cost budget, context-size need and quality sensitivity from the five dropdowns.
  2. Read the recommended model and its estimated monthly spend at the reference workload (6,000 input and 1,200 output tokens per call, 3,000 calls a month).
  3. Check the top three cards and the full ranked list. A model that fails a hard gate (cost, latency or context) ranks below every model that passes.
  4. Use the per-axis table to compare published input and output prices per 1M tokens, context windows and which gates each model passes.
  5. Treat the ranking as a shortlist: it uses published prices and vendor positioning only, so confirm quality on your own eval set before switching.

Questions people ask

How does the selector recommend a model?

It scores every model in its table on five axes: whether the monthly cost at a fixed reference workload fits your budget, whether the model's tier fits your latency budget, whether its context window covers your need, whether the vendor positions it for your task type, and how it fits your quality sensitivity. Cost, latency and context are hard gates: a model that fails one ranks after every model that passes. No accuracy benchmark enters the score.

Why is Claude Sonnet recommended for most workloads?

It is not recommended by default. At the page's starting inputs (summarize, under 5 seconds, $50 a month, 32K to 200K context, medium quality) the cheapest models that pass every gate rank first. Sonnet-tier models gain points when quality sensitivity is medium or high and when the vendor positions that tier for your task; change the inputs and the ranking changes with them.

When does the selector recommend a GPT model over Claude?

When its published price, context window, latency tier and task positioning score higher for your inputs. The selector does not model tool-use maturity, rate limits or tail latency; it uses only the published per-token prices, context windows and vendor tier positioning shown in the per-axis table.

Are open-source models considered?

No. The table covers hosted Anthropic, OpenAI and Google models with published per-token API prices. Self-hosted open-weight models have no comparable published per-token price, so they are not scored.

How often is the selector updated?

Prices and context windows are checked by hand against the vendor pricing pages, and the date of the last check is shown on the page. No scheduled eval suite sits behind the ranking.

  • Calculators Token-Cost Optimizer

    Compute the dollar cost of a trading research loop across Claude, GPT, and Gemini. Prompt length × model × retry × call volume → cost per idea and per.

  • Calculators Financial Document Token Estimator

    Paste a 10-K, 10-Q, 8-K or earnings transcript and see token count + one-pass extraction cost across ten frontier LLMs, with cache-hit toggle.

  • Calculators Batch vs Real-Time Cost Calculator

    Jobs per day, tokens per job, model, deadline — get real-time vs batch cost side-by-side with savings estimate and batch-eligibility flag. Based.

All articles

Use it from code

The same calculation as a JavaScript module you can import. It runs where you import it, with no request, key or rate limit.

import { compute } from "https://aifinhub.io/engines/model-selector-finance.js";

Input and output contract and the guide for agents.