From Backtest to Production: a Bot Architecture Path
A four-stage path from historical tick files to a live trading bot: research on day bundles, API cross-checks, SDK backtesting, and production deployment.
A strategy that looks great in a notebook and loses money live usually did not get worse — it got measured differently. The backtest ran on cleaned candles while production sees raw ticks; the backtest filled instantly at mid while production queues at the touch; the backtest's clock was the candle close while production's is a feed with jitter. This guide describes the architecture path we recommend to close those gaps one stage at a time — using self-serve pieces where possible.
Everything below runs on the CryptoStruct archive: every message Kalshi, Polymarket and 35+ crypto venues publish — every Level-2 snapshot and update at full depth, every trade, every contract, every day since we added the venue — captured co-located with nanosecond venue and receive timestamps and a gap-audited event-id chain, in one normalized schema. It is the same capture our own trading engine and enterprise feeds run on, sold as €1 day bundles you buy as a guest in the Data Shop with instant download — no subscription, no minimum, no sales call. Free full-day samples let you run every command in this guide before paying.
The four stages
Stage 1 — research on the archive, at tick resolution
Start where iteration is cheapest: historical files. Archive days cost €1 per instrument-day (or per series-day bundle on prediction venues) and contain the full L2 order-book path plus every trade at nanosecond timestamps — enough to build queue-aware fill models and charge real spread-crossing costs instead of assuming mid-price fills. (L2 data cannot tell you your exact place in the queue; it lets you estimate it with an explicit queue model rather than ignore it.) Parse with the free Python reader or the one-click CSV/Parquet exports, and validate walk-forward (the training-data guide covers splits and label hygiene).
If your simulated fill logic never touches the order book, your backtest measures a different market than the one you will trade. A deliberately conservative baseline: fill limit orders only when the touch trades through your price, and charge takers the real spread from the file. Real fills additionally depend on traded volume ahead of you, your estimated queue position and the venue's matching rules — a queue-aware model refines this heuristic; it does not make it exact.
Stage 2 — cross-check against an independent implementation
Before trusting your parser and aggregation code, reconcile them against numbers you did not compute. The free per-minute statistics API serves OHLC, VWAP, trade counts, turnover, spreads and book depth per minute for every instrument — computed by an independent implementation from the same recordings. The comparison validates your parser and aggregation path, not the original exchange capture, which both sides share. Aggregate your tick-level results to minutes and diff them: a persistent mismatch usually indicates a parsing, normalization or aggregation difference that should be investigated (aggressor-side mix-ups and quote-currency confusion are the classics). The live analytics give the same numbers visually for spot checks.
Stage 3 — port into an event model that also runs live
The subtle production killer is the programming model: research code is usually vectorized over a whole day, while live code reacts to one event at a time. Port your strategy into an event-driven harness before going live — ours is the Strategy SDK (Java): you implement callbacks for book changes, trades, order updates and timers, and the same class runs in the backtest runner and in production. The tutorial series walks from first strategy to local backtesting. Whatever SDK you use, the invariant to enforce is: one strategy artifact, two runners — never two implementations.
Stage 4 — shadow, then go live on the feed the archive came from
In production, data quality becomes latency-shaped: you care about jitter, gaps and tails. The archive you researched on was recorded from our co-located realtime feeds — going live on those feeds keeps capture and normalization identical between research and production, which removes one major source of train/live distribution shift (normalized schema, exchange timestamps, feed arbitrage against jitter; the realtime tiers are on the pricing page). Deployment placement — regions, links, how to measure your own feed delay — has its own guide: Where to run your prediction-market bot.
Ramp up, don't switch on
- Shadow mode: run the strategy on the live feed with orders logged, not sent — and diff its decisions and hypothetical fills against the backtest runner on the same days.
- Canary size: send real orders at minimum size; measure fill ratio, slippage, reject rate and feed delay against your backtest assumptions.
- Scale up only when the live metrics sit inside the tolerance you set in research — and keep the shadow comparison running as a regression check.
Production risks the backtest never sees
An event-driven port and a careful ramp-up still leave a class of failure modes no historical file contains — the live order path itself. Budget engineering time for at least these before scaling up:
- Order latency and rejects — the venue can answer slowly, or with an error; every reject path needs a tested handler.
- Rate limits — order, cancel and data requests are throttled per venue; bursts that were free in the backtest can lock you out live.
- Disconnects and reconnects — sessions drop; resyncing the book and your open-order state after a gap must be automatic.
- Stale or degraded feeds — detect when data stops being trustworthy and stop quoting instead of trading on it.
- Order-state reconciliation — periodically confirm positions and open orders against the venue; local state drifts after any hiccup.
- Callback and event ordering — live events can interleave in ways the backtest never produced; strategy logic must not depend on a fortunate ordering.
This is why the loop below starts with recording your live fills, quotes and rejects — those logs are also your incident forensics. The trading API documentation covers the live order path in detail.
Close the loop
- Record your live fills, quotes and rejects with the same timestamp discipline as the market data.
- Rebuild your backtest's fill model from measured live slippage — not the other way around.
- Re-run the research pipeline on fresh archive days every few weeks; models drift with regimes.
- Alert on distribution shifts between backtest assumptions and live measurements (spread, fill ratio, feed delay p99).
- Keep one metric dashboard for both runners — if backtest and live are not comparable at a glance, the loop is open.
Limitations
This is an architecture guide, not a strategy recipe: following it makes results measurable and transferable, not profitable. Pricing facts (€1 archive days, realtime terms) are current as of August 2026 — the pricing page is authoritative. The SDK examples reference our Java stack; the staging logic applies to any event-driven framework. Nothing here is investment advice.
Why buy the data in this guide here
Four things every page on this site is built on — and the reason the numbers above exist at all.
We record everything
The complete public feed of each venue as it was published: every Level-2 snapshot and update at the venue's full book depth, every trade with its aggressor side, every quote, funding, mark-price and liquidation event — for every instrument the venue lists, every UTC day since we added the venue. Nothing sampled, no top-N cut, no on-demand capture.
Institutional grade
Captured co-located at the venue with the exchange timestamp and our receive timestamp in integer nanoseconds, an event-id chain that makes any gap visible, and one normalized schema across 35+ venues — the same capture our own high-frequency trading engine and enterprise feeds run on.
€1 per instrument-day
Any instrument-day is €1, series-day bundles start at €1 — no subscription, no minimum order, no tiers to unlock. Credit packs lower the effective price and never expire, and every venue has free full-day samples to test against first.
Self-service for everyone
Pick the days in the Data Shop, pay by card as a guest and download immediately — no sales call, no enterprise contract, no KYC. Coding agents buy the same files through the MCP server, and the free Agent Skill teaches them the format.
Frequently asked questions
Why do backtested strategies fail in live trading?
Usually because backtest and live differ in data (candles vs. ticks), fill assumptions (mid-price vs. queueing at the touch) or programming model (vectorized vs. event-driven). Closing those three gaps stage by stage is exactly what this path is for.
Can I backtest with the same data the live feed provides?
Yes — that is the design: the archive is recorded from the same co-located, normalized feeds offered in realtime, with exchange timestamps in both. A model validated on archive days sees the same schema and clocks live.
Do I need the Strategy SDK to follow this path?
No. The invariant is framework-agnostic: one event-driven strategy artifact that both a backtest runner and production execute. Our Java SDK implements that pattern with a local backtest runner, but any framework with the same property works.
How do I validate my own tick parser and aggregations?
Reconcile against an independent implementation: aggregate your ticks to minutes and diff them against the free per-minute statistics API, which computes OHLC/VWAP/turnover from the same recordings. That validates your parser and aggregation, not the exchange capture both sides share — persistent mismatches usually point to a parsing, normalization or aggregation difference, not market noise.
Should I go straight from backtest to live trading?
No — ramp up. Run the strategy in shadow mode on the live feed first (orders logged, not sent), then trade minimum size while comparing fill ratio, slippage and rejects against backtest assumptions, and only then scale. Each step is a cheap chance to catch a gap the backtest could not show.