Skip to main content
AI Fin Hub

Playground

Prompt Regression Tester

Run one prompt across Claude, GPT-5, and Gemini with your own API keys and diff the outputs to catch prompt drift before it ships. Browser-only,

Runs in your browser with your own Anthropic, OpenAI or Google API key, which is sent only to that provider.

Education, not investment advice. Past performance does not predict future results. How we check our numbers.

Your own API keys, used only in your browser

You provide your own API keys for each target. Keys stay in your browser's React state for the session; nothing is persisted or sent anywhere except the respective provider's API.

1. Prompt

2. Targets

How to use it

  1. Write or paste the prompt to compare. A balance-sheet summary prompt is preloaded.
  2. Set up targets, each a provider, model and your own API key, and enable the ones to run.
  3. Run all enabled targets in parallel at temperature 0 and read each output with its latency and token counts.
  4. Read the mean pairwise drift (1 − Jaccard similarity of 3-character sequences) and the drift matrix between every pair of successful runs.
  5. After a model upgrade, re-run the same prompt and read high-drift outputs by hand. The score measures surface text, not meaning.

Questions people ask

What's prompt regression?

Detecting when a prompt that previously produced good output starts producing degraded output, usually because the model behind the API changed. New model versions, even minor ones, can shift outputs subtly. A regression tester catches this before users do.

How does the tester decide outputs are equivalent?

It does not judge equivalence. It scores surface drift between two outputs as 1 − the Jaccard similarity of their 3-character sequences (0 = identical text, 1 = no overlap) and labels the mean below 0.2 tight agreement, 0.2 to 0.5 moderate divergence and above 0.5 high divergence. Schema checks and semantic similarity are not part of the score.

How do I baseline?

Run the prompt on your current model and keep that output. The page does not store runs between sessions; keep the current model as one of the targets, or send your saved output as the baseline through the agent API and compare a new output against it.

How often should I re-baseline?

After every confirmed model upgrade in production, plus quarterly if you're tracking ambient drift on a 'static' model version. Do not re-baseline reactively after a regression alert — investigate first.

What if my prompts produce non-deterministic outputs?

Every run uses temperature 0, but outputs can still differ between calls. Add the same model twice as two targets to see the noise floor before reading drift between different models.

  • Playgrounds Agent Skill Tester for Markets

    Paste a SKILL.md definition + sample input + your Anthropic API key. See structured extraction, token cost, and latency — all in your browser. No signup.

  • Playgrounds Prompt Injection Tester

    Red-team a finance agent against 23 documented prompt-injection attacks — direct override, role confusion, indirect injection via retrieved content.

  • Calculators Token-Cost Optimizer

    Compute the dollar cost of a trading research loop across Claude, GPT, and Gemini. Prompt length × model × retry × call volume → cost per idea and per.

All articles
  • Workflow Audit your pipeline

    Catch hallucinations, prompt injections, and regression drift before they ship.

Use it from code

The same calculation as a JavaScript module you can import. It runs where you import it, with no request, key or rate limit.

import { compute } from "https://aifinhub.io/engines/prompt-regression-tester.js";

Input and output contract and the guide for agents.