DashboardCopilotBrokerage
Get the app ↗
← Back to terminalVolaren School
Risk and returnFactor investingPortfolio optimizationBacktesting and its traps

Backtesting and its traps

A backtest is a simulation of a strategy on the past, and it is the only laboratory markets offer. It is also the most efficient self-deception machine ever built for investors, because the past holds still while you search it, and enough searching always finds something. The traps below are not exotic; every one of them ships in the average published track record.

Trap 1: overfitting, the master trap

Test enough rules on one history and some fit it perfectly by chance. The 200-day average worked; try 187 days, better; add a volume filter, better still. Each refinement fits the NOISE of this one sample, and noise does not repeat. The tell is fragility: a real edge survives its parameters being wiggled, its start date moving, its universe changing. An overfit one shatters.

The multiple-testing arithmetic
test 100 random strategies -> ~5 will look "significant" at p < 0.05
                              BY CONSTRUCTION, with zero real edge

the honest question is never "did something pass?"
it is "how many things did I try before this passed?"
(and the industry's answer is: thousands, silently)

Traps 2-4: the data lying quietly

  • Survivorship bias. Testing on today's index members means testing on companies that, by definition, survived. The delisted, bankrupt and acquired are missing, and they are exactly the stocks a strategy would have bought on the way down. Point-in-time universes are the fix, and most free data is not point-in-time.
  • Look-ahead bias. Using information before it existed: trading on Q4 earnings dated December 31 that were not filed until late February, or on a "final" economic number later revised. The subtlest version hides in restated financials and in signals built on today's cleaned dataset.
  • Costs and capacity ignored. Paper trades execute free at the close; real ones pay spreads, impact, borrow fees and taxes. High-turnover strategies are routinely profitable before costs and worthless after, and a strategy's capacity (how much money it absorbs before its own trading moves the price) is invisible in a backtest.

Trap 5: the regime assumption

Every backtest assumes tomorrow is drawn from the same distribution as the sample. A strategy tested on 2010-2021 was tested on one regime: falling rates, quiet inflation, buy-the-dip equities. 2022 was drawn from a different urn, and portfolios "proven" on the sample met it unprotected. The discipline is testing ACROSS regimes by name (the inflation seventies, the 2008 crash, the 2020 whipsaw) and reporting performance per regime, not one blended number that lets a decade of calm bury two months of ruin.

The disciplines that separate research from curve-fitting

DisciplineWhat it does
Hold-out sampleFit on one period, judge ONLY on the untouched one; touching it twice makes it training data
Walk-forwardRe-fit on rolling windows, trade the next window: closest simulation of live research
Economic rationale firstWrite down WHY the edge should exist (who pays it, and why they keep paying) before touching data
Parameter plateausDemand the edge survive neighboring parameters; cliffs are noise
Cost stressDouble the assumed costs; a real edge survives, a microstructure artifact dies
Deflated metricsDiscount the Sharpe for how many trials produced it (the number nobody volunteers)
The paper-to-live haircut

The empirical rule of thumb across the industry: expect live performance around HALF the backtest, and treat anything better as a gift. The gap is every trap above plus one more that cannot be simulated: the researcher who found the strategy was themselves selected by it working. When a live strategy tracks its backtest closely, that is the anomaly worth investigating.

Reading someone else's backtest

  • Ask for the trade count: statistical claims on 30 trades are stories.
  • Ask what was tried and discarded: the denominator of the search decides the meaning of the survivor.
  • Ask for net-of-cost, per-regime results, and the worst drawdown IN the sample against the worst plausible one out of it.
  • Ask when the sample ends: a strategy backtested to a convenient end date chose its ending.
Glossary for this guide
Backtest
A simulation of a strategy on historical data: the only laboratory markets offer, and an efficient self-deception machine, because the past holds still while you search it.
Overfitting
Fitting the noise of one sample rather than a repeatable signal. The tell is fragility: real edges survive parameter wiggles and date shifts; overfit ones shatter.
Multiple testing
Test a hundred random ideas and about five pass at conventional significance by construction. The meaning of a result depends on how many trials produced it, the number nobody volunteers.
Survivorship bias
Testing on today's members means testing only on survivors; the delisted and bankrupt are missing, and they are what a strategy would have bought on the way down.
Look-ahead bias
Using information before it existed: trading on year-end earnings not filed until February, or on later-revised data. The subtlest versions hide in restatements and cleaned datasets.
Point-in-time data
Data as it stood on each historical date, including companies later delisted and numbers later revised. The antidote to survivorship and look-ahead, and rarely what free datasets are.
Transaction costs
Spreads, market impact, borrow fees and taxes: what paper trading ignores. High-turnover strategies are routinely profitable before costs and worthless after.
Capacity
How much money a strategy absorbs before its own trading moves prices against it. Invisible in backtests, and where good live records go to die.
Regime
A market era with its own rules: the inflation seventies, the QE 2010s. Backtests assume tomorrow is drawn from the sample's regime; report performance per named regime instead of one blended number.
Walk-forward testing
Fit on a rolling window, trade the next, repeat: the closest simulation of how research meets the future.
Hold-out sample
Data set aside untouched until final judgment. Consult it twice and it silently becomes training data; discipline about this separates research from curve-fitting.
Parameter plateau
An edge that persists across neighboring parameter values. Plateaus suggest signal; a sharp peak at one magic number is noise wearing a crown.
Volaren

Hedge fund infrastructure, built for retail investors.

Asset classes
Stocks
Options
Commodities
Crypto
ETFs
Rates & bonds
Indices
Currencies
Tools
Dashboard
Volaren Copilot
Connect brokerage
Settings
About us
Who we are
Our story
Methodology
Volaren School
Reference
Contact us
Pricing
Security
Terms
Privacy
Backed by
Y CombinatorAfore Capital
Partners
OpenAI
Not investment advice. Investing involves risk, including loss of principal. Volaren Inc. is not a registered investment adviser.© 2026 Volaren Inc.