Heartbeat and API downtime handling

A heartbeat is a periodic self-check that answers "is everything actually working?" — distinct from your bot merely still running as a process (which tells you nothing about whether it's doing its job correctly).

Components a heartbeat should check

class HealthCheck:
    def __init__(self, kite, price_cache, kill_switch):
        self.kite = kite
        self.price_cache = price_cache
        self.kill_switch = kill_switch
        self.last_heartbeat_alert = None

    def run(self) -> dict:
        checks = {
            "api_reachable": self._check_api(),
            "token_valid": self._check_token(),
            "feed_alive": self._check_feed(),
            "positions_reconciled": self._check_reconciliation(),
        }
        if not all(checks.values()):
            self._handle_failure(checks)
        return checks

    def _check_api(self) -> bool:
        try:
            self.kite.margins(segment="equity")
            return True
        except Exception:
            return False

    def _check_token(self) -> bool:
        try:
            self.kite.profile()
            return True
        except TokenException:
            return False

    def _check_feed(self) -> bool:
        return not self.price_cache.is_stale(CRITICAL_TOKEN, max_age=30)

    def _check_reconciliation(self) -> bool:
        diffs = reconcile_positions(self.kite, get_local_positions())
        return len(diffs) == 0

    def _handle_failure(self, checks: dict):
        failed = [k for k, v in checks.items() if not v]
        send_alert(f"Health check failed: {failed}", urgent=True)
        if "token_valid" in failed:
            self.kill_switch.trigger(KillLevel.HALT_NEW_ENTRIES, "Access token invalid")
        if "feed_alive" in failed:
            self.kill_switch.trigger(KillLevel.HALT_NEW_ENTRIES, "Market data feed stale")

Broker API downtime — a real, not-hypothetical scenario

Brokers do have outages, especially during high-volume periods (results season, budget day, extreme volatility days) — exactly when you might most want your bot working. Your bot's response should be conservative by default:

def handle_api_downtime(consecutive_failures: int, kill_switch: KillSwitch):
    if consecutive_failures >= 3:
        kill_switch.trigger(KillLevel.HALT_NEW_ENTRIES, "Broker API appears down — halting new entries")
    if consecutive_failures >= 10:
        send_alert("Broker API down for extended period — manual intervention needed", urgent=True)

Note: during an actual outage, you likely cannot flatten positions either (the same API is down) — this is a real limit of what any bot can do, and part of why position sizing (chapters 71-74) should never assume you can always exit on demand.

Run the heartbeat on its own independent schedule

def heartbeat_loop(health_check: HealthCheck, interval_s=30):
    while True:
        health_check.run()
        time.sleep(interval_s)

Run this on a separate thread/process from your main strategy loop, so a hang or crash in strategy logic doesn't also silently kill your ability to detect that something's wrong.

Next: 093 — Running the bot as a service