Distribution shape: skew, kurtosis, percentiles
Chapter 126's mean/median/std describe the *center* and *spread* of a distribution. This chapter describes its *shape* — which matters enormously for trading returns, since strategy return distributions are routinely far from the symmetric bell curve that mean/std alone would imply.
from scipy import stats as scipy_stats
def distribution_shape(data: pd.Series) -> dict:
return {
"skewness": scipy_stats.skew(data.dropna()),
"kurtosis": scipy_stats.kurtosis(data.dropna()), # excess kurtosis (0 = normal distribution)
"percentile_5": data.quantile(0.05),
"percentile_25": data.quantile(0.25),
"percentile_75": data.quantile(0.75),
"percentile_95": data.quantile(0.95),
"iqr": data.quantile(0.75) - data.quantile(0.25),
}
Skewness — asymmetry, and why it matters for strategy comparison
- Positive skew: many small losses, occasional large wins (typical of trend-following/breakout strategies, options buying).
- Negative skew: many small wins, occasional large losses (typical of mean-reversion strategies, options selling/premium collection).
def compare_strategies_by_skew(strategy_returns: dict[str, pd.Series]) -> pd.DataFrame:
return pd.DataFrame({name: distribution_shape(returns) for name, returns in strategy_returns.items()}).T
Two strategies with identical Sharpe ratios (chapter 131) can have completely different risk characters if their skew differs — a negatively skewed strategy (steady small wins, rare large losses) can look deceptively smooth and low-risk on a standard metrics table while carrying tail risk that only shows up in an extreme, rare event.
Kurtosis — how fat the tails are (how often extreme outcomes occur)
Excess kurtosis > 0 ("leptokurtic") means more extreme outcomes than a normal distribution would predict — financial return series almost universally exhibit this ("fat tails"). This is a direct, quantified reason why risk models assuming normally-distributed returns systematically underestimate real tail risk.
def tail_risk_flag(data: pd.Series, kurtosis_threshold: float = 3.0) -> bool:
return scipy_stats.kurtosis(data.dropna()) > kurtosis_threshold
Percentiles — a more robust way to communicate "typical" outcomes than mean alone
def trade_outcome_percentiles(pnl_series: pd.Series) -> dict:
return {
"worst_5pct_avg": pnl_series[pnl_series <= pnl_series.quantile(0.05)].mean(),
"median": pnl_series.median(),
"best_5pct_avg": pnl_series[pnl_series >= pnl_series.quantile(0.95)].mean(),
}
"What's my median trade outcome, and what does a bad-5%-of-the-time trade look like?" is a more actionable question for setting expectations than a single mean number, especially for a skewed distribution where mean and median diverge meaningfully (chapter 126's describe_series already exposes both — checking whether they're close is itself a quick skew diagnostic).
Visualizing shape, not just computing numbers
def plot_return_distribution(data: pd.Series):
import matplotlib.pyplot as plt
fig, axes = plt.subplots(1, 2, figsize=(12, 4))
axes[0].hist(data.dropna(), bins=50)
axes[0].set_title("Return distribution")
scipy_stats.probplot(data.dropna(), dist="norm", plot=axes[1])
axes[1].set_title("Q-Q plot vs normal distribution")
plt.show()
The Q-Q plot is a quick, honest visual check for exactly the fat-tail and skew properties this chapter quantifies — points curving away from the diagonal line at the extremes is the fat-tails signature, directly visible before you've computed a single number.