Historical data limits and pagination
Kite's historical API doesn't paginate via a cursor — it enforces a hard date-range limit per request, per interval, and you handle "pagination" yourself by chunking date ranges (chapter 28's fetch_intraday_chunked is the pattern).
Approximate limits (verify current values in the docs before relying on them)
| Interval | Max range per request |
|---|---|
minute | ~60 days |
3minute–15minute | ~100 days |
30minute–60minute | ~180-200 days |
day | ~2000 days |
Rate limits stack with range limits
Fetching 2 years of 1-minute data for one instrument means ~12+ chunked requests. At ~3 req/sec, that's a few seconds — fine for one instrument, but multiply by 50 instruments in a universe and you're looking at minutes, and you'll want backoff + progress tracking, not a naive loop that dies halfway through and restarts from scratch.
import time
def fetch_all_chunks(kite, token, interval, start, end, chunk_days, sleep_s=0.4):
results = []
cursor = start
while cursor < end:
chunk_end = min(cursor + pd.Timedelta(days=chunk_days), end)
try:
results += kite.historical_data(token, cursor, chunk_end, interval)
except NetworkException:
time.sleep(2)
continue # retry same chunk, don't advance cursor
cursor = chunk_end
time.sleep(sleep_s)
return results
Persist progress so a crash doesn't cost you the whole backfill
For a multi-instrument, multi-year backfill, write each instrument's fetched range to a small progress table/file as you go, so a restart can skip already-completed work instead of re-downloading everything:
def already_fetched(instrument_token: int, progress_file="data/backfill_progress.json") -> bool:
import json, pathlib
p = pathlib.Path(progress_file)
if not p.exists():
return False
return str(instrument_token) in json.loads(p.read_text())