Skip to main content
AI Fin Hub

Playground

Fallback Chain Simulator

LLM fallback chain simulator: Monte Carlo a primary + two fallbacks across Anthropic, OpenAI, Google. Success rate, p50/p95/p99 latency, cost, degradations.

Rates from aidevhub.io model catalog (each price verified on the vendor page), Anthropic pricing, OpenAI pricing and Google AI / Gemini pricing, checked on October 3, 2026.

Runs in your browser. Nothing you enter is uploaded, and no account or API key is needed.

Education, not investment advice. Past performance does not predict future results. How we check our numbers.

1. Configure the chain

Primary

Reference median latency 800 ms

429 rate

Share of calls throttled.

p99 latency

Values below the median are treated as the median.

Fallback 1

Reference median latency 800 ms

429 rate

Share of calls throttled.

p99 latency

Values below the median are treated as the median.

Fallback 2

Reference median latency 800 ms

429 rate

Share of calls throttled.

p99 latency

Values below the median are treated as the median.

ms

One budget for the whole chain, fallbacks included.

Success rate

97.7%

977 / 1000 requests answered inside the 3.00s deadline, p95 2.19s, p99 3.00s

p50 811ms (failed requests count at the deadline). Cost $0.012 per request on average, $12.309 across all 1,000. Answered by fallback 1: 85, by fallback 2: 3.

2. Provider utilization (share of successful trials)

Claude Sonnet 5.5primary91.0%, 889 trials
GPT-5.4 Minifallback 18.7%, 85 trials
Gemini 3.8 Flashfallback 20.3%, 3 trials
Failed (every leg throttled, or deadline passed)2.3%, 23 trials

Recommendation

On cost per successful call alone, Gemini 3.8 Flash beats the current primary ($0.0144 → $0.0050). Swap only if its output quality is acceptable for the task; quality is not modeled.

How the trial is simulated

for each request:
  elapsed = 0
  for leg in [primary, fallback1, fallback2?]:
    if uniform() < rate_429:        # throttled: fast, no tokens billed
      elapsed += 50ms; continue
    latency ~ LogNormal(median = model p50, 99th pct = your p99)
    elapsed += latency; bill tokens
    if elapsed <= deadline: return success
    return failure                  # deadline passed, no time for a fallback
  return failure                    # every leg throttled

The model p50 values are rough reference medians, not provider SLAs; set each leg's p99 from your own telemetry. Failures are independent per leg (rate limits and latency tail). Real outages are correlated and bursty, so treat the success rate as an upper bound during an incident.

How to use it

  1. Add up to three legs in priority order: primary, then fallback 1 and fallback 2.
  2. For each leg, pick the provider and model, then set its 429 (throttle) rate and p99 latency. The median latency is the model's reference value shown under the model.
  3. Set the request-level inputs: input and output tokens per call and the end-to-end deadline for the whole chain. Cost is derived from tokens × each model's price; you don't enter a per-call cost.
  4. Read the success rate, p50/p95/p99 latency, average cost per request, and how many requests each fallback answered. Results update as you type and are reproducible (fixed seed).
  5. Compare a 2-leg chain against a 3-leg chain; the marginal success-rate gain usually shrinks after the second fallback. Each leg is tried once — there is no retry toggle.

Questions people ask

What's a fallback chain?

A sequence of LLM endpoints to try in order when an upstream call fails: e.g., primary Anthropic, fallback OpenAI, last-resort cache. Production agent systems use chains to maintain availability when individual providers go down or rate-limit. The simulator models cost, latency, and success rate across the chain.

How are failure rates set?

You set them. Each leg has a 429 (throttle) rate from 0 to 50% that you set to your observed provider failure rate, plus a p99 latency; the defaults are 8%, 5% and 3% for the three legs. The simulator does not auto-calibrate from status pages or published reliability targets: the numbers are yours to supply, and each leg's median latency is a rough reference value shown under the model.

What's the cost of a fallback path?

Cumulative cost of every leg that actually runs before success. In the simulator a 429 is fast (about 50ms) and bills no tokens, so if the primary fails 0.5% of the time and the fallback is 3× more expensive, your effective cost is 0.995 × primary + 0.005 × 3·primary ≈ 1.010× primary. A slow answer that misses the end-to-end deadline is still billed (the tokens were generated) and ends the request, because there is no time left for a fallback.

Should the fallback always be a different provider?

Multi-provider fallback gives correlation-of-failure protection (one provider's outage doesn't take down both). Same-provider fallback (different model tier) gives capacity-overflow protection (rate-limit on Opus, retry on Sonnet). The simulator lets you mix both strategies.

Does the simulator model retries?

No. Each leg is tried exactly once, and on failure the request falls through to the next leg — there is no retry-on-timeout or retry-on-error toggle. To approximate a retry against the same provider, add another leg pointing at the same provider and model.

  • Calculators Agent Cost Envelope Calculator

    Model an LLM research loop end-to-end — steps, tool calls, convergence checks, markets per day — and see per-loop, daily, and monthly cost with cost-cap.

  • Generators Trading System Blueprinter

    Pick your data source, LLM, broker, storage, risk engine, and logger. Get a Mermaid architecture diagram and a copyable starter file tree — the full stack before you write code.

  • Calculators Token-Cost Optimizer

    Compute the dollar cost of a trading research loop across Claude, GPT, and Gemini. Prompt length × model × retry × call volume → cost per idea and per.

All articles

Use it from code

The same calculation as a JavaScript module you can import. It runs where you import it, with no request, key or rate limit.

import { compute } from "https://aifinhub.io/engines/fallback-chain-simulator.js";

Input and output contract and the guide for agents.