Trade performance stats: win rate, expectancy, profit factor
The standard metrics for judging whether a strategy's trade-level outcomes actually constitute an edge — built on chapters 126-129's statistical primitives, applied specifically to P&L per trade.
def trade_performance_stats(pnl_series: pd.Series) -> dict:
wins = pnl_series[pnl_series > 0]
losses = pnl_series[pnl_series < 0]
win_rate = len(wins) / len(pnl_series) if len(pnl_series) else 0
avg_win = wins.mean() if len(wins) else 0
avg_loss = losses.mean() if len(losses) else 0 # negative number
expectancy = win_rate * avg_win + (1 - win_rate) * avg_loss
profit_factor = wins.sum() / abs(losses.sum()) if losses.sum() != 0 else float("inf")
return {
"total_trades": len(pnl_series),
"win_rate_pct": win_rate * 100,
"avg_win": avg_win,
"avg_loss": avg_loss,
"reward_risk_ratio": abs(avg_win / avg_loss) if avg_loss != 0 else float("inf"),
"expectancy_per_trade": expectancy,
"profit_factor": profit_factor,
}
Why expectancy matters more than win rate alone
A 35% win rate with a 3:1 reward:risk ratio has positive expectancy (0.35 × 3R - 0.65 × 1R = +0.4R per trade) — genuinely profitable despite losing most individual trades. A 70% win rate with a 1:3 reward:risk ratio (small wins, occasional large losses) can have *negative* expectancy (0.70 × 1R - 0.30 × 3R = -0.2R) despite feeling "successful" most of the time. Win rate in isolation, without reward: risk context, is close to meaningless — this is exactly why chapter 85's scrutiny checklist doesn't lead with win rate.
Profit factor — a robustness cross-check on expectancy
def profit_factor_interpretation(pf: float) -> str:
if pf < 1.0:
return "losing strategy — total losses exceed total wins"
elif pf < 1.5:
return "marginal — likely doesn't survive realistic costs (ch 86)"
elif pf < 2.0:
return "reasonable, worth walk-forward validating further (ch 84)"
else:
return "strong on paper — verify this isn't concentration risk (ch 85) before trusting it"
A very high profit factor (>3) is itself worth suspicion, not celebration — cross-check against chapter 126's outlier-contribution flag and chapter 85's overfitting checklist before assuming it reflects a durable edge rather than a few lucky outlier trades or a look-ahead-biased backtest (chapter 83).
Breaking expectancy down by regime/condition — connects to chapter 128's frequency work
def expectancy_by_condition(df: pd.DataFrame, condition_col: str) -> pd.DataFrame:
"""df: trade journal with a column tagging each trade's market condition (e.g. 'trending'/'ranging')."""
return df.groupby(condition_col)["pnl"].apply(lambda pnl: pd.Series(trade_performance_stats(pnl))).unstack()
If a strategy's overall positive expectancy is entirely driven by trending-regime trades, with negative expectancy in ranging conditions, that's actionable — it directly suggests adding a regime filter (chapter 106's ADX gate) rather than trading unconditionally, and it's a concrete, testable refinement rather than a vague intuition.
Running expectancy against a null hypothesis — is this even distinguishable from noise?
def expectancy_confidence_interval(pnl_series: pd.Series, num_bootstrap: int = 10000) -> tuple[float, float]:
"""Bootstrap resampling to get a confidence interval on expectancy — chapter 133 covers this
and significance testing in full depth; this is the specific application to expectancy."""
means = [pnl_series.sample(len(pnl_series), replace=True).mean() for _ in range(num_bootstrap)]
return np.percentile(means, 2.5), np.percentile(means, 97.5)
A confidence interval on expectancy that includes zero means you cannot yet statistically distinguish this strategy's results from a coin-flip with the observed trade count — chapter 133 builds this out properly, but the seed of the idea belongs here: a positive expectancy number computed from few trades is a much weaker claim than the same number computed from hundreds.