Look-ahead bias and survivorship bias
These are the two most common ways a backtest lies to you — both make results look better than reality, and both are easy to introduce accidentally.
Look-ahead bias
Using information in a signal/decision that wouldn't actually have been available at that point in time.
Classic examples:
- Using a day's closing price to decide a trade you then simulate entering at that same day's close — in reality you'd only know the close *after* it happened, too late to act on it that same bar.
- Computing an indicator (e.g. a pivot point, a day's high/low range) using the full day's data, then applying the signal to trades executed earlier in that same day.
- Using restated/corrected historical data (e.g. earnings figures later revised) as if it were known at the time.
- The most common one: shifting a signal calculation incorrectly, so today's decision effectively uses tomorrow's price.
# WRONG: signal computed from today's close, "entered" at today's close — impossible live
signal = df["close"] > df["close"].rolling(20).mean()
returns = df["close"].pct_change()
strategy_returns = signal * returns # BUG: no shift — trades on same-bar information
# CORRECT: signal decided using data up to and including today,
# but the resulting trade only executes on tomorrow's bar
strategy_returns = signal.shift(1) * returns
That single .shift(1) is the difference between a legitimate backtest and a fantasy. Audit every signal calculation for this class of bug before trusting any result.
Survivorship bias
Testing only on instruments that *still exist today* — silently excluding companies that were delisted, went bankrupt, or were merged out of existence during your test period.
Why this inflates results: the worst-performing stocks are exactly the ones most likely to disappear from a current index constituent list or a "top NIFTY 500 stocks today" data pull. A backtest built on today's NIFTY 500 list, applied retroactively to 2015, silently excludes every stock that failed or was removed since — making the historical "universe" look better than it actually was to someone trading it live in 2015 without hindsight of who'd survive.
# WRONG: today's index constituents, applied to a historical backtest
current_nifty500 = get_current_nifty500_list()
backtest(current_nifty500, start="2015-01-01") # survivorship bias baked in
# CORRECT: point-in-time constituent lists for each period tested
def get_historical_constituents(as_of_date) -> list[str]:
# requires point-in-time index membership data — a real data engineering
# task, not available from a simple "current constituents" API call
...
Practical mitigation
- Always shift signals by at least one bar relative to any execution price derived from the same bar (look-ahead).
- Source point-in-time universe membership data where available, or explicitly caveat results as "current-universe survivorship-biased" and discount confidence accordingly if point-in-time data isn't accessible (survivorship).
- When in doubt about whether a calculation is look-ahead-safe, ask: "if I were standing at this exact bar, live, with no knowledge of what comes next, could I have computed this exact signal value?"