Calculator
Agent Cost Envelope Calculator
Agent cost envelope for LLM research loops — model, tokens per step, tool-use steps, convergence check, markets per day → per-loop, daily, monthly cost.
Rates from aidevhub.io model catalog (each price verified on the vendor page), Anthropic pricing, OpenAI pricing and Google AI / Gemini pricing, checked on October 3, 2026.
Runs in your browser. Nothing you enter is uploaded, and no account or API key is needed.
1. Configure your agent loop
Prompt, context and tool results sent to the model each step.
Reasoning, tool-call arguments and the answer.
Model calls (tool-use rounds) per market analyzed.
Final self-check as a share of one normal step. 0% = no check.
One loop per market, ticker or idea.
Drives budget used and the token cap.
Monthly cost
$35/mo
Claude Sonnet 5.5, 10 markets/day, 22 days, inside $200 target
Per loop: $0.158, Per day: $1.58, Budget used: 17%
Cap recommendation
Per-loop budget
$0.909
$200 / month ÷ 220 loops
Max tokens per loop
278.6K
Current: 48.5K
Suggested cap per step
54.6K
You have headroom
2. Per-loop breakdown
| Step | Input cost | Output cost | Total |
|---|---|---|---|
| Step 1 — tool call + reasoning | $0.016 | $0.015 | $0.031 |
| Step 2 — tool call + reasoning | $0.016 | $0.015 | $0.031 |
| Step 3 — tool call + reasoning | $0.016 | $0.015 | $0.031 |
| Step 4 — tool call + reasoning | $0.016 | $0.015 | $0.031 |
| Step 5 — tool call + reasoning | $0.016 | $0.015 | $0.031 |
| Convergence check — final analysis | $0.00160 | $0.00150 | $0.00310 |
| Per-loop total | $0.158 |
3. Sensitivity — swap model tier
Same inputs, every model in the current lineup (12), cheapest first. Shows what the envelope becomes if you change tier and which models fit your target budget.
| Model | Tier | Per loop | Per day | Per month | vs target |
|---|---|---|---|---|---|
| GPT-6 Lunaopenai | economy | $0.00791 | $0.079 | $2 | inside |
| Gemini 3.5 Flash-Litegoogle | economy | $0.031 | $0.314 | $7 | inside |
| Gemini 3.8 Flashgoogle | mid | $0.059 | $0.593 | $13 | inside |
| GPT-5.4 Miniopenai | mid | $0.065 | $0.650 | $14 | inside |
| Claude Haiku 4.5anthropic | economy | $0.079 | $0.790 | $17 | inside |
| Claude Sonnet 5.5anthropic(selected) | mid | $0.158 | $1.58 | $35 | inside |
| GPT-6.1 Solopenai | mid | $0.158 | $1.58 | $35 | inside |
| Gemini 3.1 Pro Previewgoogle | frontier | $0.173 | $1.73 | $38 | inside |
| Claude Opus 5.5anthropic | frontier | $0.316 | $3.16 | $70 | inside |
| GPT-5.5openai | frontier | $0.433 | $4.33 | $95 | inside |
| Claude Fable 5.1anthropic | frontier | $0.790 | $7.90 | $174 | inside |
| GPT-6 Astraopenai | frontier | $0.790 | $7.90 | $174 | inside |
How the envelope is priced
step_cost = input_tokens × in_rate + output_tokens × out_rate convergence_cost = step_cost × convergence_pct loop_cost = steps × step_cost + convergence_cost daily_cost = loop_cost × markets_per_day monthly_cost = daily_cost × (22 biz | 30 crypto) budget_per_loop = target_monthly / (markets_per_day × days_per_month) max_tokens_per_loop = budget_per_loop / blended_$_per_token
List prices per 1M tokens, verified 2026-10-03 against each provider's pricing page. Prompt caching and batch discounts are not applied: the envelope is a ceiling.
How to use it
- Pick the primary model and enter what one agent step consumes: input tokens (prompt, context and tool results sent to the model) and output tokens (reasoning, tool-call arguments and the answer).
- Set steps per loop (model calls per market analyzed) and the convergence-check cost as a percentage of one step. Set it to 0% if the loop has no final self-check.
- Enter markets analyzed per day, your monthly budget, and the calendar: 22 business days for equities or 30 days for 24/7 crypto markets.
- Read the monthly cost, then the per-loop cost, per-day cost and share of budget used. The cap recommendation back-solves the maximum tokens per loop and per step that keep the loop inside your budget.
- Use the sensitivity table to price the same loop on every model in the rate table, cheapest first, and see which ones fit the budget.
Questions people ask
What's an agent cost envelope?
The cost ceiling of a bounded research loop: (steps × cost of one step + convergence check) × markets per day × days per month. Each step is priced at the model's list input and output rates per 1M tokens. Caching and batch discounts are deliberately left out, so the result is an upper bound to budget against, and the cap recommendation tells you how many tokens per loop and per step fit inside your monthly target.
How do retry assumptions affect cost?
The calculator has no retry input: every step is billed once. Each retry repeats the step's input cost and adds a fresh output cost, so if roughly 10% of steps are retried, raise steps per loop by about 10% (or add one step per loop) to keep the envelope honest.
Should I budget for the median or the 95th percentile?
The tool is deterministic and reports a single envelope, not a distribution, so there is no median or 95th percentile to choose between. Treat the envelope as the cost of a loop that runs exactly as configured, and add headroom for loops that take extra steps. The budget-used figure and the per-loop cap show how much room you have.
Does the tool include human-review cost?
No. It models LLM call cost only. Adding human-in-the-loop review can dominate cost for high-stakes agents — but the time/cost varies so much by team that the calculator stays focused on the LLM bill.
Why split prompt-caching savings into a separate field?
The envelope does not apply caching at all: every input token is priced at the full list rate, which keeps it a ceiling. Cache hit rate depends on how the loop is deployed, not on the model, so savings belong in a separate pass. Anthropic bills cache reads at about 0.1× the input rate; the Token-Cost Optimizer models a cache-hit rate on the same price table.
Related tools
- Calculators Token-Cost Optimizer
Compute the dollar cost of a trading research loop across Claude, GPT, and Gemini. Prompt length × model × retry × call volume → cost per idea and per.
- Calculators Batch vs Real-Time Cost Calculator
Jobs per day, tokens per job, model, deadline — get real-time vs batch cost side-by-side with savings estimate and batch-eligibility flag. Based.
- Playgrounds Fallback Chain Simulator
Define a provider fallback chain, simulate rate-limit and latency failures, and see p50/p95/p99 latency, success rate, total cost, and the degradation-event distribution.
Articles
- 11 min read Observability Patterns for LLM Trading Agents
Three patterns that stop silent failure: trace-ID propagation, structured log schema with per-step cost and confidence, and a deterministic replay harness.
- 10 min read Bounded-Cost Agentic Research
Three gates stop runaway agent loops: hard token budget, step-count cap, and a cost-convergence check that halts when belief stops moving.
- 11 min read Agent Memory Patterns for Finance Research
Three memory tiers for finance agents — working, episodic, long-term lesson library — with retention policies and runnable Python for each.
Workflows that use this tool
- Workflow Plan your agent stack
Estimate first-year cost for an LLM agent — token budget, vendor selection, MCP servers.
Use it from code
The same calculation as a JavaScript module you can import. It runs where you import it, with no request, key or rate limit.
import { compute } from "https://aifinhub.io/engines/agent-cost-envelope-calculator.js";