Skip to main content
AI Fin Hub

Playground

Calibration Dojo

Train probabilistic intuition. Binary forecasting questions at any confidence level; track Brier score + reliability curve over time.

Runs in your browser. Nothing you enter is uploaded, and no account or API key is needed.

Education, not investment advice. Past performance does not predict future results. How we check our numbers.

Brier score

No answers yet

Lower is better. Perfect calibration scores 0.000; always answering 50% scores 0.250

0 answered. Answers are stored in this browser only.

1. Read the statement

Loading…

Reliability curve

Dot size ∝ number of answers in that decile. Diagonal = perfect calibration.

History persists in your browser's localStorage only.

How to use it

  1. Read the statement and set the probability that it is true (0% = certainly false, 100% = certainly true), then click Commit probability.
  2. The reveal shows the answer, your Brier score for that question, and the rationale and source behind the answer. Click Next question to continue.
  3. Read the reliability curve: dots on the diagonal mean your 70% answers come true about 70% of the time. Dots below the diagonal on the right mean you are overconfident.
  4. Compare your Brier score with the base-rate benchmark shown under it. A skill score above 0 means you beat always answering the base rate; reliability and resolution show whether calibration or discrimination is holding you back.
  5. Answer at least 50 questions before drawing conclusions; progress is saved in this browser and Clear local history resets it.

Questions people ask

What's calibration in this context?

How well a model's stated confidence matches its empirical hit rate. A 70%-confident forecast that's right 70% of the time is well-calibrated; a 90%-confident forecast that's right 60% of the time is overconfident. The dojo trains this with immediate feedback: you commit a probability to a true-or-false statement, see the answer and its source, and watch your reliability curve fill in.

How is the Brier score calculated?

Brier score is the mean squared error between probability forecast and outcome. For a single forecast: (probability − outcome)². For a series: average across forecasts. Lower is better: 0 is perfect, and answering 50% on everything scores exactly 0.25. The dojo also shows the score of always answering the base rate (the share of statements that were true), base rate × (1 − base rate), and your skill score against it.

What's the difference between calibration and resolution?

Calibration is whether your stated probabilities match observed frequencies. Resolution is whether you give different probabilities to different outcomes (vs. 50% on everything). A forecaster who always says the base rate is perfectly calibrated but has zero resolution, and is useless. The dojo reports both terms of the Murphy decomposition from decile bins: reliability (the calibration gap, lower is better) and resolution (higher is better). Brier score combines them.

How many forecasts before I get a stable calibration estimate?

At least 50, ideally 100 or more. The reliability curve plots, for each 10%-wide probability bin, how often the statements you put in that bin were true; with a handful of answers per bin one surprise moves a dot a long way, and the dojo shows no confidence interval, so read early curves loosely. The bank has 24 questions; after that the dojo re-asks them at random.

Why does a low Brier score not always mean I'm right?

A Brier score is meaningful only relative to a benchmark. Forecasting 'AAPL up tomorrow' at 50% gives Brier 0.25 with zero skill. The dojo always reports your score against the benchmark of forecasting the base rate, so you can see whether your work is adding signal or just hitting the average.

Where can I see forecast calibration scored in public on real markets?

Probability-scored forecasts are easiest to find outside markets: Good Judgment Open scores every forecast with the Brier score once its question resolves, the same squared error the dojo uses, and ranks forecasters against the crowd (https://www.gjopen.com/faq). Public records of stock calls usually grade direction rather than probability. FrontierPicks, a sister site, grades each AI stock pick against the S&P 500 with a kill line written down before the pick opens (https://frontierpicks.com/track-record/): a record of commitment you can check, but not a calibration score, because no probability is stated.

  • Calculators Returns Distribution Analyzer

    Paste a returns CSV. Histogram, normal QQ plot, skewness, excess kurtosis, Jarque-Bera test, tail-weight index. See why Sharpe alone misleads.

  • Calculators Backtest Overfitting Score

    Upload a backtest trade log and compute Probability of Backtest Overfitting (PBO), Deflated Sharpe Ratio, and the odds your edge survives live trading.

  • Calculators Kelly Criterion Calculator

    Size positions with full, half, quarter, or eighth Kelly and stress-test the choice with a drawdown Monte Carlo simulator. Client-side. Private by default.

  • 8 min read The Price-Blind LLM Research Harness

    Price-blind LLM research — most harnesses leak the current price and the model confabulates. The architectural fix and a 30-line Python scaffold.

  • 8 min read Conviction-Scaled Kelly Bet Sizing

    Full Kelly is brutally unforgiving of over-estimation. Quarter-Kelly with a conviction-tier mapping and a per-trade cap is the defensible default.

  • 10 min read Calibrating LLM Forecasts with Isotonic Regression

    LLM probabilities are systematically miscalibrated. Isotonic regression via PAV is the cheapest robust fix: 40 lines of Python, no distributional priors.

All articles
  • Workflow Validate your strategy

    Pressure-test a quant or LLM-augmented strategy before paper-trading or production.

Use it from code

The same calculation as a JavaScript module you can import. It runs where you import it, with no request, key or rate limit.

import { compute } from "https://aifinhub.io/engines/calibration-dojo.js";

Input and output contract and the guide for agents.