Playground
Agent Skill Tester for Markets
Test Anthropic Agent Skills for market extraction. Paste SKILL.md + sample 10-K excerpt + your key. See output, token cost, latency.
Runs in your browser with your own Anthropic, OpenAI or Google API key, which is sent only to that provider.
Where your key and data go
- Anthropic call goes direct from your browser to
api.anthropic.com. The API key never reaches aifinhub. - Scoring imports the
/engines/agent-skill-tester.jsmodule and scores the prompt + response in your browser. No key, no model id, no network call. Deterministic checks, no LLM.
1. Config
2. Skill definition (SKILL.md)
3. Sample input document
4. Pass/fail rubric (one criterion per line)
Recognised: valid_json, has_field:<path>, field_type:<path>:<type>, contains:<text>, regex:<pattern>, no_apology, numbers_grounded_in_prompt, min_length:<n>, max_length:<n>, not_contains:<text>, non_empty, refuses, cites_source. Lines starting with # are ignored.
How to use it
- Paste your SKILL.md instructions into the skill definition box. They are sent as the system prompt of a single model call.
- Paste a sample input document that looks like what the skill will see in production. Use a realistic excerpt, not a minimal one.
- Enter your Anthropic API key and pick a model. The call goes straight from your browser to api.anthropic.com. The key stays in browser memory unless you tick 'Remember in this browser', which saves it to localStorage.
- Write the pass/fail rubric, one criterion per line (for example valid_json, has_field:company_name, field_type:revenue_usd:number, numbers_grounded_in_prompt), then click Run skill + score.
- Read how many criteria passed, the verdict and detail for each criterion, and the latency, token counts and cost of the run. Re-run a few times: verdicts that flip between runs mean the skill is under-constrained.
Questions people ask
What's a SKILL.md?
A markdown instruction file that packages one repeatable agent capability, for example 'extract the headline figures from a 10-K excerpt as JSON'. The tester sends your SKILL.md as the system prompt and the sample document as the user message in one Anthropic API call, then scores the response against your rubric.
Why does the tester need my own API key?
The call goes from your browser directly to api.anthropic.com, so the cost lands on your account and no proxy ever sees the key or your documents. The key stays in browser memory unless you tick 'Remember in this browser', which stores it in this browser's localStorage until you clear it.
What does the tester measure?
Rubric compliance first: each criterion you list (valid JSON, required fields, field types, length limits, substrings, regexes, refusals, citations, numbers grounded in the input) gets a pass, fail or error verdict, and the result shows how many passed. It also reports latency, input and output tokens, and cost per run at list prices. The scoring is deterministic string and JSON checking, not an LLM judge.
Why does my SKILL.md sometimes fail validation?
There is no JSON Schema validation step: the checks are the rubric criteria you write. The usual failures are code fences or commentary around the JSON (valid_json strips a single ```json fence), a missing or null field (has_field), a number returned as a string (field_type), and figures the model computed or reformatted (numbers_grounded_in_prompt, which accepts restated scales such as $2,847 million = 2847000000 but not rounding).
Can I test multi-turn skills?
No. Each run is a single call: system prompt (your SKILL.md) plus one user message (the sample document), scored on the one response. Skills that ask clarifying questions or need tool calls have to be exercised against the API in your own harness; the same rubric vocabulary is available headlessly through the agent-skill-tester engine module for scoring those runs.
Related tools
- Playgrounds Prompt Regression Tester
Run the same prompt against multiple models (Claude 4.5/4.6/4.7, GPT-5, Gemini 2.5) with your own keys. Diff outputs, score drift, catch regressions.
- Playgrounds Hallucination Detector
Paste a source document + an LLM's extraction. Every numeric claim in the output is checked against the source. Client-side. Catches silent fabrication.
- Playgrounds Prompt Injection Tester
Red-team a finance agent against 23 documented prompt-injection attacks — direct override, role confusion, indirect injection via retrieved content.
Articles
- 9 min read The 8-Step LLM Research Prompt Template
Free-form prompts yield uncalibrated LLM output. An 8-step template makes research reproducible and better-calibrated across model versions.
- 9 min read LLM Prompt Patterns for 10-K and 8-K Extraction
Three structured patterns for auditable 10-K extractions: field-by-field JSON, citation-required verbatim quotes, and contradiction-triangle cross-check.
- 10 min read Options Greeks for LLM-Driven Trading
Options Greeks for LLM-driven trading: delta, gamma, theta, vega, rho — what each costs, three rules, plus a prompt template for multi-leg positions.
Workflows that use this tool
- Workflow Audit your pipeline
Catch hallucinations, prompt injections, and regression drift before they ship.
Use it from code
The same calculation as a JavaScript module you can import. It runs where you import it, with no request, key or rate limit.
import { compute } from "https://aifinhub.io/engines/agent-skill-tester.js";