Skip to main content
AI Fin Hub

Calculator

Agent Cost Envelope Calculator

Agent cost envelope for LLM research loops — model, tokens per step, tool-use steps, convergence check, markets per day → per-loop, daily, monthly cost.

Rates from aidevhub.io model catalog (each price verified on the vendor page), Anthropic pricing, OpenAI pricing and Google AI / Gemini pricing, checked on October 3, 2026.

Runs in your browser. Nothing you enter is uploaded, and no account or API key is needed.

Education, not investment advice. Past performance does not predict future results. How we check our numbers.

1. Configure your agent loop

tokens

Prompt, context and tool results sent to the model each step.

tokens

Reasoning, tool-call arguments and the answer.

Model calls (tool-use rounds) per market analyzed.

Convergence-check cost

Final self-check as a share of one normal step. 0% = no check.

One loop per market, ticker or idea.

$

Drives budget used and the token cap.

Calendar

Monthly cost

$35/mo

Claude Sonnet 5.5, 10 markets/day, 22 days, inside $200 target

Per loop: $0.158, Per day: $1.58, Budget used: 17%

Cap recommendation

Per-loop budget

$0.909

$200 / month ÷ 220 loops

Max tokens per loop

278.6K

Current: 48.5K

Suggested cap per step

54.6K

You have headroom

2. Per-loop breakdown

StepInput costOutput costTotal
Step 1 — tool call + reasoning$0.016$0.015$0.031
Step 2 — tool call + reasoning$0.016$0.015$0.031
Step 3 — tool call + reasoning$0.016$0.015$0.031
Step 4 — tool call + reasoning$0.016$0.015$0.031
Step 5 — tool call + reasoning$0.016$0.015$0.031
Convergence check — final analysis$0.00160$0.00150$0.00310
Per-loop total$0.158

3. Sensitivity — swap model tier

Same inputs, every model in the current lineup (12), cheapest first. Shows what the envelope becomes if you change tier and which models fit your target budget.

ModelTierPer loopPer dayPer monthvs target
GPT-6 Lunaopenaieconomy$0.00791$0.079$2inside
Gemini 3.5 Flash-Litegoogleeconomy$0.031$0.314$7inside
Gemini 3.8 Flashgooglemid$0.059$0.593$13inside
GPT-5.4 Miniopenaimid$0.065$0.650$14inside
Claude Haiku 4.5anthropiceconomy$0.079$0.790$17inside
Claude Sonnet 5.5anthropic(selected)mid$0.158$1.58$35inside
GPT-6.1 Solopenaimid$0.158$1.58$35inside
Gemini 3.1 Pro Previewgooglefrontier$0.173$1.73$38inside
Claude Opus 5.5anthropicfrontier$0.316$3.16$70inside
GPT-5.5openaifrontier$0.433$4.33$95inside
Claude Fable 5.1anthropicfrontier$0.790$7.90$174inside
GPT-6 Astraopenaifrontier$0.790$7.90$174inside

How the envelope is priced

step_cost          = input_tokens × in_rate + output_tokens × out_rate
convergence_cost   = step_cost × convergence_pct
loop_cost          = steps × step_cost + convergence_cost
daily_cost         = loop_cost × markets_per_day
monthly_cost       = daily_cost × (22 biz | 30 crypto)
budget_per_loop    = target_monthly / (markets_per_day × days_per_month)
max_tokens_per_loop = budget_per_loop / blended_$_per_token

List prices per 1M tokens, verified 2026-10-03 against each provider's pricing page. Prompt caching and batch discounts are not applied: the envelope is a ceiling.

How to use it

  1. Pick the primary model and enter what one agent step consumes: input tokens (prompt, context and tool results sent to the model) and output tokens (reasoning, tool-call arguments and the answer).
  2. Set steps per loop (model calls per market analyzed) and the convergence-check cost as a percentage of one step. Set it to 0% if the loop has no final self-check.
  3. Enter markets analyzed per day, your monthly budget, and the calendar: 22 business days for equities or 30 days for 24/7 crypto markets.
  4. Read the monthly cost, then the per-loop cost, per-day cost and share of budget used. The cap recommendation back-solves the maximum tokens per loop and per step that keep the loop inside your budget.
  5. Use the sensitivity table to price the same loop on every model in the rate table, cheapest first, and see which ones fit the budget.

Questions people ask

What's an agent cost envelope?

The cost ceiling of a bounded research loop: (steps × cost of one step + convergence check) × markets per day × days per month. Each step is priced at the model's list input and output rates per 1M tokens. Caching and batch discounts are deliberately left out, so the result is an upper bound to budget against, and the cap recommendation tells you how many tokens per loop and per step fit inside your monthly target.

How do retry assumptions affect cost?

The calculator has no retry input: every step is billed once. Each retry repeats the step's input cost and adds a fresh output cost, so if roughly 10% of steps are retried, raise steps per loop by about 10% (or add one step per loop) to keep the envelope honest.

Should I budget for the median or the 95th percentile?

The tool is deterministic and reports a single envelope, not a distribution, so there is no median or 95th percentile to choose between. Treat the envelope as the cost of a loop that runs exactly as configured, and add headroom for loops that take extra steps. The budget-used figure and the per-loop cap show how much room you have.

Does the tool include human-review cost?

No. It models LLM call cost only. Adding human-in-the-loop review can dominate cost for high-stakes agents — but the time/cost varies so much by team that the calculator stays focused on the LLM bill.

Why split prompt-caching savings into a separate field?

The envelope does not apply caching at all: every input token is priced at the full list rate, which keeps it a ceiling. Cache hit rate depends on how the loop is deployed, not on the model, so savings belong in a separate pass. Anthropic bills cache reads at about 0.1× the input rate; the Token-Cost Optimizer models a cache-hit rate on the same price table.

  • Calculators Token-Cost Optimizer

    Compute the dollar cost of a trading research loop across Claude, GPT, and Gemini. Prompt length × model × retry × call volume → cost per idea and per.

  • Calculators Batch vs Real-Time Cost Calculator

    Jobs per day, tokens per job, model, deadline — get real-time vs batch cost side-by-side with savings estimate and batch-eligibility flag. Based.

  • Playgrounds Fallback Chain Simulator

    Define a provider fallback chain, simulate rate-limit and latency failures, and see p50/p95/p99 latency, success rate, total cost, and the degradation-event distribution.

  • 11 min read Observability Patterns for LLM Trading Agents

    Three patterns that stop silent failure: trace-ID propagation, structured log schema with per-step cost and confidence, and a deterministic replay harness.

  • 10 min read Bounded-Cost Agentic Research

    Three gates stop runaway agent loops: hard token budget, step-count cap, and a cost-convergence check that halts when belief stops moving.

  • 11 min read Agent Memory Patterns for Finance Research

    Three memory tiers for finance agents — working, episodic, long-term lesson library — with retention policies and runnable Python for each.

All articles
  • Workflow Plan your agent stack

    Estimate first-year cost for an LLM agent — token budget, vendor selection, MCP servers.

Use it from code

The same calculation as a JavaScript module you can import. It runs where you import it, with no request, key or rate limit.

import { compute } from "https://aifinhub.io/engines/agent-cost-envelope-calculator.js";

Input and output contract and the guide for agents.