Frequency distributions and return histograms
A frequency distribution counts how often values fall into each range — the raw material behind chapter 126's binned mode and chapter 127's histograms, formalized here as its own reusable tool, plus the specific trading applications that make it useful beyond just "plot a histogram."
def frequency_table(data: pd.Series, bins: int = 20) -> pd.DataFrame:
counts, bin_edges = np.histogram(data.dropna(), bins=bins)
return pd.DataFrame({
"bin_start": bin_edges[:-1],
"bin_end": bin_edges[1:],
"count": counts,
"frequency_pct": counts / counts.sum() * 100,
})
Frequency of trade outcomes — a concrete example
def trade_outcome_frequency(pnl_series: pd.Series) -> dict:
total = len(pnl_series)
return {
"win_frequency_pct": (pnl_series > 0).sum() / total * 100,
"loss_frequency_pct": (pnl_series < 0).sum() / total * 100,
"breakeven_frequency_pct": (pnl_series == 0).sum() / total * 100,
"big_win_frequency_pct": (pnl_series > pnl_series.quantile(0.9)).sum() / total * 100,
"big_loss_frequency_pct": (pnl_series < pnl_series.quantile(0.1)).sum() / total * 100,
}
Frequency of signal occurrence — how often does a strategy actually trade?
def signal_frequency(signal_series: pd.Series, timeframe_bars_per_day: int) -> dict:
total_bars = len(signal_series)
signal_count = (signal_series != 0).sum()
return {
"signals_per_day": signal_count / (total_bars / timeframe_bars_per_day),
"pct_bars_with_signal": signal_count / total_bars * 100,
}
This directly informs practical decisions: a strategy generating 40 signals/day on 1-minute NIFTY futures needs a genuinely low-latency, well-tested execution path (chapters 43-68) and realistic cost modeling (chapter 86) that a strategy generating 2 signals/week does not — trade frequency changes which parts of this entire course matter most for a given strategy.
Frequency of a market condition — validating a hypothesis's applicability
def condition_frequency(df: pd.DataFrame, condition: pd.Series) -> float:
"""e.g., how often is ADX > 20 (trending) vs not, over the test period?"""
return condition.mean() * 100
Recall chapter 79's hypothesis needs a stated universe and time period — if your opening-range-breakout strategy's entry condition only fires in a trending regime (chapter 106's ADX gate), and the trending regime only occurred 15% of the days in your backtest period, that 15% frequency number is itself important context for interpreting the backtest's overall statistics; a strategy validated on a regime that occurs 15% of the time tells you little about the other 85%.
Comparing frequency distributions across time periods — a walk-forward diagnostic
def compare_frequency_across_windows(pnl_by_window: dict[str, pd.Series]) -> pd.DataFrame:
return pd.DataFrame({
window: trade_outcome_frequency(pnl) for window, pnl in pnl_by_window.items()
}).T
If win frequency is stable at ~40% across every walk-forward window (chapter 84) but suddenly drops to 15% in the most recent window, that's a concrete, quantified version of the "process failure vs bad luck" question chapter 98 raised — a frequency shift this large is a specific, checkable signal, not just a feeling that something's off.