Statistical significance and robustness testing

Closes Part 14 and the extended course. Chapter 130 introduced bootstrap confidence intervals for expectancy; this chapter builds the full significance-testing toolkit — the final, formal answer to "is this edge real, or could it plausibly be noise?"

t-test: is mean return significantly different from zero?

from scipy import stats as scipy_stats

def t_test_returns(returns: pd.Series) -> dict:
    t_stat, p_value = scipy_stats.ttest_1samp(returns.dropna(), 0)
    return {"t_statistic": t_stat, "p_value": p_value, "significant_at_5pct": p_value < 0.05}

Important caveat: the standard t-test assumes returns are independent and normally distributed — chapter 127 already established financial returns are typically fat-tailed and often autocorrelated (today's return isn't fully independent of yesterday's, especially for a trend-following strategy). A t-test p-value here is a useful signal, not a rigorous proof — treat it as one input, not the final word, and prefer the bootstrap approach below when the independence/normality assumptions are shaky, which for trading strategies is most of the time.

Bootstrap resampling — the more robust, assumption-light approach

def bootstrap_test(pnl_series: pd.Series, num_iterations: int = 10000) -> dict:
    """Resamples trade P&Ls with replacement to build a distribution of possible mean outcomes,
    without assuming normality."""
    observed_mean = pnl_series.mean()
    bootstrap_means = [
        pnl_series.sample(len(pnl_series), replace=True).mean() for _ in range(num_iterations)
    ]
    ci_low, ci_high = np.percentile(bootstrap_means, [2.5, 97.5])
    p_value_approx = np.mean([m <= 0 for m in bootstrap_means]) if observed_mean > 0 else np.mean([m >= 0 for m in bootstrap_means])
    return {"observed_mean": observed_mean, "ci_95_low": ci_low, "ci_95_high": ci_high, "p_value_approx": p_value_approx}

If the 95% confidence interval includes zero, you cannot yet statistically distinguish this strategy's edge from noise at your current sample size — a direct, quantitative gate before committing meaningful capital, more honest than eyeballing an equity curve.

Monte Carlo permutation test — is the *sequence* of trades meaningful, or just the set of outcomes?

def permutation_test(pnl_series: pd.Series, num_permutations: int = 10000) -> dict:
    """Shuffles trade ORDER (not values) to test whether your strategy's specific sequencing
    (e.g. a compounding position-sizing scheme) adds value beyond just having this set of trade outcomes."""
    observed_final_equity = (1 + pnl_series / 100000).cumprod().iloc[-1]   # example: compounding on 100k base
    permuted_finals = []
    for _ in range(num_permutations):
        shuffled = pnl_series.sample(frac=1, replace=False).reset_index(drop=True)
        permuted_finals.append((1 + shuffled / 100000).cumprod().iloc[-1])
    percentile_rank = (np.array(permuted_finals) < observed_final_equity).mean() * 100
    return {"observed_final_equity": observed_final_equity, "percentile_rank_vs_random_order": percentile_rank}

This isolates a subtle question distinct from "is the edge real": given this exact set of trade outcomes, did the *order* they happened in (and any sizing scheme reacting to that order, like a Kelly-style scaling) add or destroy value versus a random shuffle of the same trades? Useful for validating dynamic position-sizing logic specifically, separate from validating the underlying signal.

Multiple-testing correction — the honest reckoning with chapter 85's warning

If you tested 50 different indicator/parameter combinations (chapter 117's grid search) and one came back with p < 0.05, that is not surprising by chance alone — with 50 independent tests at a 5% threshold, you'd expect roughly 2-3 "significant" results purely from noise.

def bonferroni_correction(p_values: list[float], alpha: float = 0.05) -> float:
    """Adjusted significance threshold accounting for the number of tests run."""
    return alpha / len(p_values)

def significant_after_correction(p_value: float, num_tests_run: int, alpha: float = 0.05) -> bool:
    return p_value < bonferroni_correction([None] * num_tests_run, alpha)

Report and use the *corrected* threshold, not the raw p-value, whenever a result emerged from searching across multiple parameter combinations or indicator sets (chapters 112-117) rather than testing one pre-registered hypothesis (chapter 79) — this is the rigorous, formal version of chapter 85's "too many free parameters" red flag.

The closing standard this entire extended course has been building toward

Before trusting any strategy — your own or someone else's pitch — you should now be able to answer, with actual numbers rather than impressions: What's the win rate, expectancy, and profit factor (ch 130)? What's the Sharpe, Sortino, Calmar, and max drawdown duration (ch 131)? Does the equity curve survive walk-forward validation with realistic, measured slippage (ch 84, 86, 129)? Is the result still significant after correcting for how many things were tried (this chapter)? A strategy that clears all of these honestly is a rare, valuable thing — and now you have the complete toolkit to tell the difference.

One more thing before you deploy anything for real: the regulatory ground under all of this shifted in 2026.

Next: 134 — SEBI's 2026 retail algo trading framework