Playground
Hallucination Detector — Finance
Paste source + LLM extraction. Every numeric claim is cross-checked against the source; ungrounded claims are flagged.
Runs in your browser. Nothing you enter is uploaded, and no account or API key is needed.
Grounding score
57%
4 / 7 numeric claims found in the source. 3 ungrounded, 4 grounded.
Numeric claims only: currencies, percents, plain numbers of 1,000 or more, and dates. Prose, names and derived figures (growth rates, ratios) are not verified; a computed number shows as not grounded.
3. Claims in the output
shadedfound in the sourcewavynot found: check it
Claims not found in the source
- percent
18.1% - currency
$780 millionnearest in source:615000000 - percent
35%
How grounding is checked
- Numbers in the output are extracted (currencies, plain numbers ≥ 1000, percents, dates).
- Currencies and plain numbers are grounded when the source states the same value within 1%, with scale words applied on both sides ($2,850 million matches $2.85 billion).
- Percents are grounded when they match a source percent up to the output’s rounding (12% matches 12.3%).
- Dates are grounded when their year appears in the source.
- Prose-level fabrication is not detected. This pass catches the numeric class only.
How to use it
- Paste the source document (the filing excerpt or retrieval result the model was supposed to ground its answer in).
- Paste the model's output to be checked. The check runs as you type.
- Read the grounding score: the share of numeric claims (amounts, percents, numbers of 1,000 or more, dates) whose value appears in the source.
- Read the highlighted output and the list of claims not found in the source, with the nearest source number for each. Computed figures such as growth rates are flagged by design, so confirm their derivation.
- For many outputs, call the hallucination-detector engine module with each (source, output) pair and track the grounding rate across your pipeline.
Questions people ask
How does the detector decide a model output is hallucinated?
It runs one check: numeric grounding. Every currency amount, percent, number of 1,000 or more and date in the output is looked up in the source. Amounts match within 1% after applying scale words ($2,850 million = $2.85 billion), percents match up to the output's rounding, and dates match on the year. Anything not found is flagged, with the nearest source number shown. There is no entity check, no LLM judge and no multi-sample consistency test.
What's the false-positive rate?
It has not been benchmarked, so there is no measured rate. The known false-positive pattern is by design: figures the model computed correctly (a growth rate, a margin, a sum) are flagged because they do not appear in the source. Approximations such as 'about $4 billion' for $4.2 billion are flagged too. Review flagged claims rather than auto-rejecting them.
Does it work for non-English text?
Partly. The number patterns are language-neutral for digits, '%', $ € £ and the scale words million, billion, thousand, bn, mn and k, but they assume a period as the decimal separator and a comma for thousands. Text that writes 1.234,5 or uses words such as Milliarden will be misread, so normalize number formats first.
Can it detect 'plausible but wrong' hallucinations?
Only when the wrong part is a number the source contradicts, such as an EPS of $0.81 when the source says $0.78. It cannot catch an invented company name, a correct number attached to the wrong line item (net income reported as revenue), or a claim the source simply does not address. Pair it with field-level checks against a structured extraction for those.
How many samples does self-consistency need?
The detector does not sample the model: it checks one output against one source, deterministically. Self-consistency checks (asking the model several times and comparing answers) are a separate technique; you can run each sample through the detector and compare grounding rates, but the tool does not do that for you.
Related tools
- Playgrounds Agent Skill Tester for Markets
Paste a SKILL.md definition + sample input + your Anthropic API key. See structured extraction, token cost, and latency — all in your browser. No signup.
- Playgrounds Price-Blind Research Auditor
Paste a research prompt or agent context bundle. The auditor flags price numbers, directional words, and outcome-leaking phrases that cause LLMs.
- Playgrounds Prompt Injection Tester
Red-team a finance agent against 23 documented prompt-injection attacks — direct override, role confusion, indirect injection via retrieved content.
Articles
- 7 min read DeepSeek V4 for Finance 2026: SEC Filing Extraction Cost
DeepSeek V4 for finance 2026: V4.1-Flash reads a full 10-K for about $0.041 at $0.30/$1.20 per Mtok, half that off-peak, with a 1M-token window.
- 8 min read The Price-Blind LLM Research Harness
Price-blind LLM research — most harnesses leak the current price and the model confabulates. The architectural fix and a 30-line Python scaffold.
- 9 min read LLM Prompt Patterns for 10-K and 8-K Extraction
Three structured patterns for auditable 10-K extractions: field-by-field JSON, citation-required verbatim quotes, and contradiction-triangle cross-check.
Workflows that use this tool
- Workflow Audit your pipeline
Catch hallucinations, prompt injections, and regression drift before they ship.
Use it from code
The same calculation as a JavaScript module you can import. It runs where you import it, with no request, key or rate limit.
import { compute } from "https://aifinhub.io/engines/hallucination-detector.js";