Correlation and covariance across instruments
Chapter 74 referenced correlation-aware sizing without building the computation; chapter 115 referenced measuring how correlated two indicator conditions are before treating their agreement as meaningful confirmation. This chapter builds the actual tools both referenced.
def correlation_matrix(returns_df: pd.DataFrame) -> pd.DataFrame:
"""returns_df: columns are instruments/strategies, rows are periods, values are returns."""
return returns_df.corr()
def covariance_matrix(returns_df: pd.DataFrame) -> pd.DataFrame:
return returns_df.cov()
Applying this to chapter 74's portfolio concentration problem
def find_highly_correlated_pairs(returns_df: pd.DataFrame, threshold: float = 0.7) -> list[tuple]:
corr = correlation_matrix(returns_df)
pairs = []
for i in range(len(corr.columns)):
for j in range(i + 1, len(corr.columns)):
if abs(corr.iloc[i, j]) > threshold:
pairs.append((corr.columns[i], corr.columns[j], corr.iloc[i, j]))
return sorted(pairs, key=lambda x: abs(x[2]), reverse=True)
If two "different" positions in your book show a correlation above, say, 0.7, chapter 74's exposure limits should treat them as much closer to one combined bet than two independent ones — this function turns that qualitative concern into a specific, checkable number per pair.
Applying this to chapter 115's indicator redundancy problem
def indicator_condition_correlation(conditions: dict[str, pd.Series]) -> pd.DataFrame:
"""conditions: {name: signal_series (-1/0/1)} — measures how redundant your confluence rules actually are."""
df = pd.DataFrame(conditions)
return df.corr()
If rsi_bias and stochastic_bias (chapter 105's note about RSI/ Stochastic similarity) show 0.85+ correlation, requiring both to agree in an AND condition (chapter 115) provides almost no additional confirmation over requiring just one — this is the concrete measurement behind that chapter's qualitative warning.
Rolling correlation — because correlation is not stable over time
def rolling_correlation(series_a: pd.Series, series_b: pd.Series, window: int = 60) -> pd.Series:
return series_a.rolling(window).corr(series_b)
A pair of stocks correlated at 0.3 historically can spike to 0.9+ during a market-wide stress event (correlations tend to rise sharply in selloffs — "correlations go to 1 in a crash") — a static, long-history correlation number can understate exactly the risk that matters most: diversification failing precisely when you need it most. Check rolling correlation, not just a single full-period number, especially for positions you're relying on for diversification.
Correlation between strategies, not just instruments — directly extends chapter 100's scaling note
def strategy_correlation(strategy_returns: dict[str, pd.Series]) -> pd.DataFrame:
return correlation_matrix(pd.DataFrame(strategy_returns))
Two strategies trading completely different instruments (say, a NIFTY options strategy and an equity swing strategy) can still be highly correlated if they both, structurally, tend to lose in the same kind of market condition (e.g. both suffer in a sharp, fast selloff) — this is exactly what chapter 100 flagged as "not real diversification," now with a concrete way to check it using actual return series rather than intuition about how "different" the strategies feel.
Correlation is not causation — and financial correlations are often spurious or unstable
A high correlation discovered in a specific historical window is not guaranteed to persist — always re-verify correlation assumptions periodically rather than computing them once and hardcoding portfolio construction decisions on a stale matrix indefinitely.