Slippage statistics: measuring and aggregating

Chapter 86 modeled slippage theoretically for backtesting. This chapter builds the measurement side — computing real slippage statistics from your own order log (chapter 58's tag-based idempotent placement, chapter 19's trade book) using every descriptive tool from chapters 126-128 applied specifically to this one number, since it's the single biggest source of backtest-vs-live divergence flagged back at the very start of this course.

Defining slippage precisely, per order

def compute_order_slippage(intended_price: float, filled_price: float, side: str) -> float:
    """Positive slippage = cost (worse than intended); negative = favorable (better than intended)."""
    if side == "BUY":
        return filled_price - intended_price   # paid more than intended = positive slippage = bad
    else:
        return intended_price - filled_price   # received less than intended = positive slippage = bad

intended_price for a market order is the LTP/best-quote at the moment you decided to place it (chapter 24's ltp() snapshot, taken immediately before placement); for a limit order, it's the limit price itself (any fill at that price has zero slippage by definition — the "slippage" for limit orders that don't fill is a different metric, covered below as fill-rate, not price slippage).

Building the slippage dataset from your order log

def build_slippage_dataset(order_log: list[dict]) -> pd.DataFrame:
    """order_log entries: {order_id, side, intended_price, filled_price, order_type, symbol, timestamp}"""
    df = pd.DataFrame(order_log)
    df["slippage_rupees"] = df.apply(
        lambda row: compute_order_slippage(row["intended_price"], row["filled_price"], row["side"]), axis=1
    )
    df["slippage_pct"] = df["slippage_rupees"] / df["intended_price"] * 100
    return df

The full descriptive statistics pass — directly answering the original ask

def slippage_statistics(df: pd.DataFrame) -> dict:
    slippage = df["slippage_rupees"]
    return {
        "min": slippage.min(),
        "max": slippage.max(),
        "mean": slippage.mean(),
        "median": slippage.median(),
        "mode": robust_mode(slippage),                       # ch 126
        "std": slippage.std(),
        "percentile_95": slippage.quantile(0.95),             # worst-case-ish, excluding true tail outliers
        "frequency_favorable_pct": (slippage < 0).mean() * 100,   # how often you did BETTER than intended
        "frequency_adverse_pct": (slippage > 0).mean() * 100,      # how often you did WORSE than intended
        "frequency_zero_pct": (slippage == 0).mean() * 100,         # exact fills — mostly limit orders
    }

Breaking this down by order type — market vs limit slippage differs fundamentally

def slippage_by_order_type(df: pd.DataFrame) -> pd.DataFrame:
    return df.groupby("order_type").apply(lambda g: pd.Series(slippage_statistics(g)))

Market orders should show consistently positive mean slippage (you're paying the spread plus any market impact, by construction) — a market -order dataset showing negative mean slippage on average would be surprising and worth investigating (possibly a sign of favorable conditions, or a bug in how intended_price was captured). Limit orders should show slippage clustered tightly at/near zero, with a separate fill-rate metric mattering more than price slippage:

def limit_order_fill_rate(df: pd.DataFrame) -> dict:
    limit_orders = df[df["order_type"] == "LIMIT"]
    return {
        "fill_rate_pct": (limit_orders["status"] == "COMPLETE").mean() * 100,
        "partial_fill_rate_pct": ((limit_orders["filled_qty"] > 0) & (limit_orders["filled_qty"] < limit_orders["intended_qty"])).mean() * 100,
    }

Breaking down by time-of-day and instrument — slippage is not uniform

def slippage_by_time_bucket(df: pd.DataFrame) -> pd.DataFrame:
    df["hour"] = pd.to_datetime(df["timestamp"]).dt.hour
    return df.groupby("hour").apply(lambda g: pd.Series(slippage_statistics(g)))

Slippage is typically worse in the first and last few minutes of the session (wider spreads, higher volatility, thinner initial depth) and around major news/data releases — this breakdown often reveals that average slippage is a misleading single number hiding a bimodal reality (good most of the day, bad in specific windows) — directly connects back to chapter 128's frequency-distribution framing: don't just look at the mean, look at *when* the bad outcomes cluster.

Feeding this back into your backtest — closing the loop from chapter 86

def calibrated_slippage_model(slippage_stats: dict) -> float:
    """Use the empirically observed mean (or a conservative upper percentile) as your
    backtest slippage assumption, instead of a guessed constant."""
    return slippage_stats["percentile_95"]   # conservative choice — biases the backtest toward caution

This is the concrete mechanism chapter 86 referenced but deferred: "calibrate slippage from real data" now has an exact, reusable pipeline — run your strategy in paper mode (chapter 87) or small real size, log every order, run slippage_statistics(), and feed a conservative percentile back into your backtest assumptions before trusting any performance projection for larger size.

Next: 130 — Trade performance stats: win rate, expectancy, profit factor