Walk-forward validation

A single train/test split (optimize parameters on 2018-2021, test on 2022-2024) tells you how the strategy did on *one* out-of-sample period. Walk-forward validation repeats this process across multiple rolling windows, which is both more data-efficient and directly tests whether a strategy's edge holds up as market conditions change, not just once.

The pattern

|--train--|--test--|
      |--train--|--test--|
            |--train--|--test--|
                  ...

Each window: optimize parameters on the "train" slice, then evaluate (without re-optimizing) on the immediately following "test" slice. Move the window forward and repeat.

def walk_forward_windows(df: pd.DataFrame, train_days: int, test_days: int, step_days: int):
    windows = []
    start = 0
    while start + train_days + test_days <= len(df):
        train = df.iloc[start : start + train_days]
        test = df.iloc[start + train_days : start + train_days + test_days]
        windows.append((train, test))
        start += step_days
    return windows

def run_walk_forward(df, strategy_fn, optimize_fn, train_days=252, test_days=63, step_days=63):
    results = []
    for train, test in walk_forward_windows(df, train_days, test_days, step_days):
        best_params = optimize_fn(train)                 # optimize ONLY on train
        test_result = strategy_fn(test, **best_params)     # evaluate on unseen test, no re-fitting
        results.append({"test_start": test.index[0], "params": best_params, **test_result})
    return pd.DataFrame(results)

What a good result looks like

  • Consistency across windows, not just a good average — a strategy profitable in 8 of 10 windows tells a different story than one profitable overall only because 2 windows were spectacular and the rest roughly broke even or lost.
  • Parameter stability — if the "best" parameters swing wildly between adjacent windows (e.g. optimal moving-average length jumping from 10 to 80 and back), that's a sign you're fitting noise, not a persistent pattern (directly connects to chapter 85's overfitting discussion).

What a bad result looks like, and what it means

If test-period performance is consistently much worse than train-period performance across most windows, the strategy is overfitting to whatever specific data it was tuned on each time, rather than capturing something that generalizes. This is exactly the failure mode a single train/test split can hide if you got lucky with which period you tested on.

Practical note for Indian markets specifically

Make sure your walk-forward windows span genuinely different regimes — Indian markets have had distinct trending bull runs, sharp corrections (2020 COVID crash), extended choppy/range-bound stretches, and high-IV periods around major events (elections, budget days). A walk-forward validation that happens to only sample similar regime types across all windows understates how much regime-dependence risk you're actually carrying.

Next: 085 — Overfitting red flags