Skip to main content
AI Fin Hub

Playground

Hallucination Detector — Finance

Paste source + LLM extraction. Every numeric claim is cross-checked against the source; ungrounded claims are flagged.

Runs in your browser. Nothing you enter is uploaded, and no account or API key is needed.

Education, not investment advice. Past performance does not predict future results. How we check our numbers.

Grounding score

57%

4 / 7 numeric claims found in the source. 3 ungrounded, 4 grounded.

Numeric claims only: currencies, percents, plain numbers of 1,000 or more, and dates. Prose, names and derived figures (growth rates, ratios) are not verified; a computed number shows as not grounded.

3. Claims in the output

Company: SYNTHETIC_A Corp. Period: fiscal year ended 2025-12-31. Revenue: $2,847 million (18.1% YoY growth). Net income: $412 million. Operating cash flow: $780 million. R&D intensity: 12.3% of revenue. Notable: primary customer concentration exceeds 35% of revenue.

shadedfound in the sourcewavynot found: check it

Claims not found in the source

  • percent18.1%
  • currency$780 millionnearest in source: 615000000
  • percent35%

How grounding is checked

  • Numbers in the output are extracted (currencies, plain numbers ≥ 1000, percents, dates).
  • Currencies and plain numbers are grounded when the source states the same value within 1%, with scale words applied on both sides ($2,850 million matches $2.85 billion).
  • Percents are grounded when they match a source percent up to the output’s rounding (12% matches 12.3%).
  • Dates are grounded when their year appears in the source.
  • Prose-level fabrication is not detected. This pass catches the numeric class only.

How to use it

  1. Paste the source document (the filing excerpt or retrieval result the model was supposed to ground its answer in).
  2. Paste the model's output to be checked. The check runs as you type.
  3. Read the grounding score: the share of numeric claims (amounts, percents, numbers of 1,000 or more, dates) whose value appears in the source.
  4. Read the highlighted output and the list of claims not found in the source, with the nearest source number for each. Computed figures such as growth rates are flagged by design, so confirm their derivation.
  5. For many outputs, call the hallucination-detector engine module with each (source, output) pair and track the grounding rate across your pipeline.

Questions people ask

How does the detector decide a model output is hallucinated?

It runs one check: numeric grounding. Every currency amount, percent, number of 1,000 or more and date in the output is looked up in the source. Amounts match within 1% after applying scale words ($2,850 million = $2.85 billion), percents match up to the output's rounding, and dates match on the year. Anything not found is flagged, with the nearest source number shown. There is no entity check, no LLM judge and no multi-sample consistency test.

What's the false-positive rate?

It has not been benchmarked, so there is no measured rate. The known false-positive pattern is by design: figures the model computed correctly (a growth rate, a margin, a sum) are flagged because they do not appear in the source. Approximations such as 'about $4 billion' for $4.2 billion are flagged too. Review flagged claims rather than auto-rejecting them.

Does it work for non-English text?

Partly. The number patterns are language-neutral for digits, '%', $ € £ and the scale words million, billion, thousand, bn, mn and k, but they assume a period as the decimal separator and a comma for thousands. Text that writes 1.234,5 or uses words such as Milliarden will be misread, so normalize number formats first.

Can it detect 'plausible but wrong' hallucinations?

Only when the wrong part is a number the source contradicts, such as an EPS of $0.81 when the source says $0.78. It cannot catch an invented company name, a correct number attached to the wrong line item (net income reported as revenue), or a claim the source simply does not address. Pair it with field-level checks against a structured extraction for those.

How many samples does self-consistency need?

The detector does not sample the model: it checks one output against one source, deterministically. Self-consistency checks (asking the model several times and comparing answers) are a separate technique; you can run each sample through the detector and compare grounding rates, but the tool does not do that for you.

  • Playgrounds Agent Skill Tester for Markets

    Paste a SKILL.md definition + sample input + your Anthropic API key. See structured extraction, token cost, and latency — all in your browser. No signup.

  • Playgrounds Price-Blind Research Auditor

    Paste a research prompt or agent context bundle. The auditor flags price numbers, directional words, and outcome-leaking phrases that cause LLMs.

  • Playgrounds Prompt Injection Tester

    Red-team a finance agent against 23 documented prompt-injection attacks — direct override, role confusion, indirect injection via retrieved content.

All articles
  • Workflow Audit your pipeline

    Catch hallucinations, prompt injections, and regression drift before they ship.

Use it from code

The same calculation as a JavaScript module you can import. It runs where you import it, with no request, key or rate limit.

import { compute } from "https://aifinhub.io/engines/hallucination-detector.js";

Input and output contract and the guide for agents.