Apollo 8
← Blog

Why Backtests Lie: The Five Biases That Make Strategies Look Better Than They Are

Jul 17, 2026 · 9 min read

Why Backtests Lie: The Five Biases That Make Strategies Look Better Than They Are

Key Takeaways

  • Definition: a backtest simulates an investment strategy on historical data to estimate what it would have earned. Indispensable in quantitative management – and dangerously easy to embellish, most often with no intent to deceive.
  • Survivorship bias: testing a strategy on today's universe ignores the securities and funds that disappeared along the way – bankruptcies, delistings, mergers. Academic research has documented since the 1990s that this omission systematically inflates simulated performance.
  • Look-ahead bias: using a data point before its actual publication date. Our own data illustrates it: across the Swiss real estate fund reports in our database with a recorded publication date, release comes a median of 73 days after the fiscal-year close – a backtest using figures "as of the closing date" therefore trades two and a half months ahead of reality.
  • Overfitting: try enough variants and a spectacular backtest always turns up – by pure chance. Academic research proposes markedly higher significance hurdles for strategies discovered through multiple trials, and provides tools to estimate the probability that a backtest is overfit.
  • Forgotten costs: transaction fees, bid-ask spreads, market impact, taxes – a cost-free backtest overstates performance in proportion to how fast the strategy trades.
  • The single regime: a strategy calibrated on one era – say, Switzerland's negative-rate years – learns the rules of a world that can stop existing, as it did in 2022–2023.
  • Key reflex: in front of a backtest, the useful questions never change: survivorship-free universe? data dated at publication? realistic costs? how many variants tried? out-of-sample validation? The answers separate a simulation from a sales argument.

Introduction

Every systematic investment strategy arrives with a simulated performance curve – smooth, ascending, convincing. That is the backtest's power: it turns an idea into numbers. It is also its danger: those numbers are manufactured from the past, through a process where every step can – often unintentionally – flatter the result.

In quantitative management, fighting these biases takes at least as much work as building the strategies themselves. This article describes the five mechanisms that make backtests lie – survivorship bias, look-ahead bias, overfitting, forgotten costs, the single regime – what academic research says about them, and the safeguards that resist them. Not to disqualify the tool: to learn how to read it.

What a Backtest Is – and Why It Seduces

A backtest applies precise investment rules to history: "each month, buy the ten cheapest securities by this criterion, sell by that one," then measures what this fictional portfolio would have earned over ten or twenty years. The exercise is legitimate and irreplaceable: it forces rules to be made explicit, reveals how an idea behaved through past crises, and lets approaches be compared on equal terms.

Its seduction comes from its apparent precision: annualized return, volatility, worst drawdown – everything quantified to the decimal. But that precision concerns a reconstructed past, and the reconstruction is a perilous exercise. The five biases that follow are its main traps – from the easiest to detect to the most insidious.

Bias No. 1: Survivorship – Testing on the Living

The most natural way to test a strategy is to start from today's investment universe: the members of an index, the funds of a category. It is also the surest way to be wrong: that universe contains, by construction, only the survivors. Bankrupt companies, delisted securities, funds liquidated or absorbed after bad years have vanished from the lists – and with them, the losses the strategy would have suffered holding them.

Academic research documented this bias as early as the 1990s, particularly on US mutual funds: measuring a category's average performance while ignoring dead funds systematically overstates the result. The magnitude varies across universes and periods; the direction of the error, never.

The remedy is known but demanding: databases that include vanished securities ("survivorship-bias free"), reconstructing the universe as it existed at each date. Our own field of study offers a concrete illustration: the universe of listed Swiss real estate funds we track counts 46 vehicles today – but the sector's recent mergers and absorptions are a reminder that this list has not always carried the same names, and that a backtest run on the current list would mechanically inherit the bias.

Bias No. 2: Look-Ahead – Knowing Before Everyone Else

An honest backtest can only use, at each simulated date, the information actually available on that date. The rule sounds obvious; it is violated constantly, for a simple reason: databases file figures under their accounting date, not their publication date.

Our data offers a direct measure of the problem. Across the Swiss real estate fund reports in our database whose publication date is recorded – just over a quarter of the base to date – the gap between the fiscal-year close and the report's release reaches 73 days at the median; a quarter of them appear more than three months after the close, some up to nine months. A backtest that uses the NAV or the annual result "as of December 31" starting January 1 therefore trades two and a half months ahead of real investors – a fictional informational edge, invisible in the curve, impossible to replicate in practice.

The same mechanism operates everywhere: corporate earnings restated after the fact, economic indicators revised, index compositions known retroactively. The remedy has a name: point-in-time data, dated at the moment the market learned it – which is precisely why our extraction pipeline records, for every report, its publication date and not only its accounting date.

Bias No. 3: Overfitting – Searching Until You Find

The most insidious bias falsifies no data: it is born of repetition. Try a strategy; if the backtest disappoints, adjust a parameter, move a threshold, test a variant – and repeat. With enough attempts, a spectacular result will eventually appear. The problem is that it no longer measures a market regularity: it measures your perseverance in interrogating randomness.

Academic research has formalized the mechanism. Work published in the Notices of the American Mathematical Society shows that the best simulated performance among a large number of trials grows mechanically with the number of trials – even when every strategy tested is worthless; its authors describe presenting a backtest without disclosing the number of trials as a practice bordering on "financial charlatanism." In the same spirit, a reference study on the hundreds of published equity factors recommends holding strategies discovered through multiple testing to a markedly higher statistical bar – a t-statistic of roughly 3 rather than 2 – and methods now exist to estimate the probability that a given backtest is overfit.

The remedies are a matter of discipline more than technique: set aside a validation period never touched during research ("out-of-sample"), advance through rolling windows ("walk-forward"), count the variants tried honestly – all of them, including the abandoned ones – and distrust a result in proportion to how long it took to find. Above all: start from an economic hypothesis – why would this premium exist, who pays it, why would it persist? – rather than letting the optimizer mine the data blindly.

Bias No. 4: Forgotten Costs

A gross backtest simulates free, instantaneous transactions. Reality bills every line: brokerage commissions, bid-ask spreads, market impact – moving the price by buying – the federal stamp duty on transactions for Swiss securities, and, depending on the strategy, securities-borrowing costs. Each item is small; their sum, multiplied by portfolio turnover, is not.

The reading rule is simple: the faster a strategy trades, the wider the gap between its gross curve and its net reality – and the more a backtest must detail its cost assumptions to be taken seriously. A strategy whose simulated edge disappears at 20 basis points of costs per trade does not have an edge: it has a cost assumption.

Added to this is capacity: a strategy tested on small caps or thinly traded securities – Swiss real estate funds know something about this – can be real at 10 million and illusory at 500 million: backtests ignore that the buyer eventually becomes the market.

Bias No. 5: The Single Regime

The last bias comes neither from the data nor from the method, but from history itself. A strategy calibrated on a given period learns the rules of that period's world – and those rules change. The Swiss example speaks for itself: between 2015 and 2022, the SNB policy rate sat at −0.75%; it climbed to 1.75% in 2023, before returning to 0% since June 2025. A yield strategy built and tested only on the negative-rate era – when every source of income commanded a premium – met a world it had never seen in 2022–2023; the correction in listed real estate funds illustrated it in its own way.

The remedy is not to test further into the past – data quality degrades quickly – but to test across regimes: check the strategy's behavior in different rate, inflation, and liquidity environments, and above all identify explicitly which regime its economic engine depends on. An honest strategy can say in which world it stops working.

The Safeguards: What a Serious Backtest Documents

Let us summarize what the five biases demand. A seriously presented backtest documents:

  • its universe – reconstructed at each date, vanished securities included;
  • its dates – every data point used at its actual publication date ("point-in-time"), not its accounting date;
  • its costs – commissions, spreads, impact, taxes, with the assumptions quantified;
  • its research process – number of variants tried, existence of an untouched validation period, parameter-selection method;
  • its regimes – behavior across different environments, and the economic conditions its signal depends on;
  • its capacity – up to what volume the execution assumptions remain realistic.

No backtest satisfies all six criteria perfectly – ours included. But the gap between what a simulation documents and what it stays silent about is, in practice, the best indicator of the confidence it deserves.

The Limits of the Critique

Three closing nuances. First, the backtest remains irreplaceable: the alternative – investing on convictions never confronted with history – is worse than all the biases combined. The critique targets badly built or badly presented backtests, not the tool.

Second, an honest backtest is still an estimate: even stripped of every bias, a simulated performance describes past markets. It bounds the expectation; it promises nothing – which is precisely why a simulation, however careful, never constitutes a forecast.

Third, sophistication does not protect against bias – it disguises it. The more complex a method, the more parameters it offers to tune, and the more likely overfitting becomes for the same amount of searching. In backtesting, documented simplicity beats silent complexity.

Conclusion

Backtests rarely lie by fraud; they lie by construction – whenever a universe forgets its dead, a data point travels through time, a researcher forgets to count the attempts, a curve ignores its costs, or an era mistakes itself for a law. These biases all push in the same direction – embellishment – which is exactly why simulated curves are so consistently more beautiful than the live performance that follows them.

The good news is that each of these biases has a known remedy, and the list of questions to ask fits on a page. It is the lens we apply to our own work – from the database built report by report, publication dates included, to the strategies we study – and the one this series of articles will keep applying, numbers in hand, to the markets it covers.

Sources

  1. Bailey, Borwein, López de Prado, Zhu – "Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance", Notices of the American Mathematical Society, 61(5), May 2014
  2. Bailey, Borwein, López de Prado, Zhu – "The Probability of Backtest Overfitting", Journal of Computational Finance, 20(4), 2017 (SSRN 2326253)
  3. Harvey, Liu, Zhu – "…and the Cross-Section of Expected Returns", Review of Financial Studies, 29(1), 2016 (SSRN 2249314)
  4. Elton, Gruber, Blake – "Survivorship Bias and Mutual Fund Performance", Review of Financial Studies, 9(4), 1996
  5. López de Prado – Advances in Financial Machine Learning, Wiley, 2018
  6. SNB – Current interest rates and exchange rates (policy rate history)