Overfitting red flags

This is the direct, expanded answer to "what should I scrutinize when someone shows me a strategy" — the checklist that should run through your head whether you're evaluating your own backtest or someone else's pitch.

1. Too many free parameters relative to trade count

A strategy with 8 tunable parameters (entry threshold, exit threshold, two moving average lengths, an RSI level, a volume filter, a time-of-day filter, a volatility filter) tested on 200 historical trades has enormous room to fit noise — with enough knobs, you can always find *some* combination that looks great on past data. Rough rule of thumb: be skeptical of any strategy with more free parameters than roughly sqrt(number_of_trades), and increasingly skeptical beyond that.

2. Suspiciously smooth equity curve

Real trading strategies have losing streaks, choppy stretches, and periods of underperformance — that's the nature of an edge playing out against noise. An equity curve that goes up and to the right with almost no drawdown is either extremely rare or, far more likely, is overfit, uses look-ahead bias (chapter 83), or the reported numbers exclude costs/slippage that would introduce realistic bumpiness.

3. Performance concentrated in a short period or a few trades

def check_concentration(trade_pnls: list[float], top_n: int = 5) -> float:
    sorted_pnls = sorted(trade_pnls, reverse=True)
    total = sum(trade_pnls)
    top_contribution = sum(sorted_pnls[:top_n])
    return top_contribution / total if total else 0

If the top 5 trades out of 200 account for 60%+ of total profit, the strategy's apparent edge is really a handful of lucky outlier events, not a repeatable process — ask what would happen to the overall result if you removed those top few trades entirely.

4. Parameters that only work in a narrow, specific range

If a moving-average crossover strategy is profitable with a 19-day fast MA but unprofitable at 18 or 20 days, that's not a real edge — it's curve-fitting to the specific noise pattern in your specific dataset. A genuine edge is usually robust across a reasonable neighborhood of parameter values, not a knife-edge optimum.

5. No causal story for why the edge should exist

"This combination of 6 indicators produced 68% win rate historically" with no explanation of *why* that combination should predict anything is a red flag. Compare to: "index rebalancing creates predictable forced buying/selling in the days around reconstitution" — a mechanism you can reason about independently of the backtest number.

6. Backtest period doesn't match walk-forward robustness (chapter 84)

A single glowing backtest number with no walk-forward or out-of-sample validation shown should be treated as unverified, regardless of how compelling the headline stat is.

7. Reported numbers exclude realistic costs

Chapter 86 covers this in depth — a backtest with zero slippage assumption and unrealistically low transaction costs will look better than any live version of the same strategy ever will.

The single most useful question to ask a strategy pitch

*"Show me the walk-forward, out-of-sample equity curve, with realistic costs included, and tell me how many free parameters you tuned to get here."* Most strategy pitches that can't answer this cleanly don't survive the question.

Next: 086 — Modeling slippage and transaction costs