Playground
Prompt Regression Tester
Run one prompt across Claude, GPT-5, and Gemini with your own API keys and diff the outputs to catch prompt drift before it ships. Browser-only,
Runs in your browser with your own Anthropic, OpenAI or Google API key, which is sent only to that provider.
Your own API keys, used only in your browser
You provide your own API keys for each target. Keys stay in your browser's React state for the session; nothing is persisted or sent anywhere except the respective provider's API.
1. Prompt
2. Targets
How to use it
- Write or paste the prompt to compare. A balance-sheet summary prompt is preloaded.
- Set up targets, each a provider, model and your own API key, and enable the ones to run.
- Run all enabled targets in parallel at temperature 0 and read each output with its latency and token counts.
- Read the mean pairwise drift (1 − Jaccard similarity of 3-character sequences) and the drift matrix between every pair of successful runs.
- After a model upgrade, re-run the same prompt and read high-drift outputs by hand. The score measures surface text, not meaning.
Questions people ask
What's prompt regression?
Detecting when a prompt that previously produced good output starts producing degraded output, usually because the model behind the API changed. New model versions, even minor ones, can shift outputs subtly. A regression tester catches this before users do.
How does the tester decide outputs are equivalent?
It does not judge equivalence. It scores surface drift between two outputs as 1 − the Jaccard similarity of their 3-character sequences (0 = identical text, 1 = no overlap) and labels the mean below 0.2 tight agreement, 0.2 to 0.5 moderate divergence and above 0.5 high divergence. Schema checks and semantic similarity are not part of the score.
How do I baseline?
Run the prompt on your current model and keep that output. The page does not store runs between sessions; keep the current model as one of the targets, or send your saved output as the baseline through the agent API and compare a new output against it.
How often should I re-baseline?
After every confirmed model upgrade in production, plus quarterly if you're tracking ambient drift on a 'static' model version. Do not re-baseline reactively after a regression alert — investigate first.
What if my prompts produce non-deterministic outputs?
Every run uses temperature 0, but outputs can still differ between calls. Add the same model twice as two targets to see the noise floor before reading drift between different models.
Related tools
- Playgrounds Agent Skill Tester for Markets
Paste a SKILL.md definition + sample input + your Anthropic API key. See structured extraction, token cost, and latency — all in your browser. No signup.
- Playgrounds Prompt Injection Tester
Red-team a finance agent against 23 documented prompt-injection attacks — direct override, role confusion, indirect injection via retrieved content.
- Calculators Token-Cost Optimizer
Compute the dollar cost of a trading research loop across Claude, GPT, and Gemini. Prompt length × model × retry × call volume → cost per idea and per.
Articles
- 8 min read The Price-Blind LLM Research Harness
Price-blind LLM research — most harnesses leak the current price and the model confabulates. The architectural fix and a 30-line Python scaffold.
- 9 min read The 8-Step LLM Research Prompt Template
Free-form prompts yield uncalibrated LLM output. An 8-step template makes research reproducible and better-calibrated across model versions.
- 8 min read The Token-Cost Reality of LLM Trading Research
What LLM trading research costs per idea and per validated trade across Claude, GPT-5, and Gemini 2.5. Pricing, caching, model-mix under $200/month.
Workflows that use this tool
- Workflow Audit your pipeline
Catch hallucinations, prompt injections, and regression drift before they ship.
Use it from code
The same calculation as a JavaScript module you can import. It runs where you import it, with no request, key or rate limit.
import { compute } from "https://aifinhub.io/engines/prompt-regression-tester.js";