Skip to main content
AI Fin Hub

Calculator

Backtest Overfitting Score

Free PBO calculator: upload a wide-format returns CSV for the probability of backtest overfitting and deflated Sharpe ratio via CSCV, in your browser.

Runs in your browser. Nothing you enter is uploaded, and no account or API key is needed.

Education, not investment advice. Past performance does not predict future results. How we check our numbers.

1. Upload your backtest returns

Wide-format CSV: one column per candidate strategy, one row per observation. Optional date column as the first column. Returns are interpreted as simple (non-log) daily returns. All computation runs in your browser — nothing uploaded.

What this tool answers

If you backtested many candidate strategies and picked the best one, how likely is it that the winner is real versus the winner of a lucky lottery? Two complementary signals:

  • PBO (Probability of Backtest Overfitting, via Combinatorially-Symmetric Cross-Validation): fraction of splits where the in-sample winner ranks below median out-of-sample. High PBO = likely overfit.
  • DSR (Deflated Sharpe Ratio): probability that the Sharpe is statistically real, adjusted for how many strategies you tested and how non-normal the returns are. Low DSR = Sharpe probably a coincidence.

Load the synthetic demo for a working example, or upload your own CSV.

How to use it

  1. Upload a wide-format CSV: one row per period (an optional date column first), one column per strategy variant, simple returns as decimals. Annualization assumes daily rows (×√252). You need at least 2 variants and 60 rows; more variants give a finer rank distribution.
  2. Pick the number of time blocks S (default 16). The tool tests every split of the blocks into equal in-sample and out-of-sample halves, sampling 500 with a fixed seed when there are more. Each block needs at least 4 rows.
  3. Read PBO (probability of backtest overfitting) — values above 0.5 mean the in-sample winner is likely to underperform out-of-sample.
  4. Read Deflated Sharpe Ratio alongside. PBO measures relative overfitting; DSR measures absolute statistical significance after multiple-testing penalty.
  5. If PBO > 0.5 or DSR < 0.95, treat the backtest as curve-fit. DSR here is a probability in [0,1], so 0.95 is the 95%-confidence bar. Reduce variant count, lengthen sample, or test on truly fresh data before live deployment.

Questions people ask

What is Probability of Backtest Overfitting (PBO)?

PBO is the probability that the strategy with the best in-sample Sharpe ranks below median out-of-sample. Bailey, Borwein, Lopez de Prado, and Zhu (2017) introduced the metric. A PBO above 0.5 means the in-sample winner is more likely to underperform than outperform in production — i.e., the backtest is more curve-fit than predictive.

How is PBO computed from a returns matrix?

Combinatorially Symmetric Cross-Validation (CSCV): cut the returns matrix into S equal time blocks and treat every choice of S/2 blocks as in-sample, the rest as out-of-sample. For each split the tool picks the variant with the best in-sample Sharpe and checks where it ranks out of sample. PBO is the share of splits where it ranks below the out-of-sample median. With S = 16 there are 12,870 splits, so the tool evaluates a fixed-seed sample of 500, which makes the result reproducible.

What's the Deflated Sharpe Ratio?

Bailey and López de Prado's (2014) correction of the Sharpe ratio for skewness, kurtosis and the number of variants tried. Here it is reported as a probability: the chance that the true Sharpe beats the expected maximum Sharpe of N zero-skill variants, using the null sampling variance 1/(T−1) for the Sharpe estimate. Above 95% is the usual bar for 'unlikely to be a lucky pick'. The table shows the annualized Sharpe, that expected-maximum benchmark in the same units, and the DSR for every variant.

How many strategy variants do I need to compute PBO?

The tool needs at least 2, but with N variants the out-of-sample rank can only take N values, so with 2 or 3 variants PBO moves in coarse steps and the tool shows a warning. Ten or more variants give a much smoother estimate. Equally important is the sample: each time block needs at least 4 rows, and longer blocks make each split's Sharpe less noisy.

Does a low PBO mean my strategy will work live?

No. PBO measures relative overfitting across the variants you tested, not absolute predictive power. A strategy can have low PBO (your best variant is genuinely better than the average variant) and still lose money live if all variants are weak. Combine PBO with deflated Sharpe and out-of-sample equity-curve inspection.

  • Playgrounds Walk-Forward Validator

    Upload a returns CSV. Rolling or expanding IS/OOS windows, per-window Sharpe, walk-forward efficiency, and a concatenated OOS equity curve. Catches regime.

  • Calculators Risk-Adjusted Returns Calculator

    Paste a returns CSV. Sharpe, Sortino, Calmar, Omega, alpha, beta, tracking error, information ratio, max drawdown, and tail moments — plus.

  • Calculators Returns Distribution Analyzer

    Paste a returns CSV. Histogram, normal QQ plot, skewness, excess kurtosis, Jarque-Bera test, tail-weight index. See why Sharpe alone misleads.

All articles
  • Workflow Validate your strategy

    Pressure-test a quant or LLM-augmented strategy before paper-trading or production.

Use it from code

The same calculation as a JavaScript module you can import. It runs where you import it, with no request, key or rate limit.

import { compute } from "https://aifinhub.io/engines/backtest-overfitting-score.js";

Input and output contract and the guide for agents.