A market data feed for a crypto trading bot: from historical tick files to the live WebSocket
Crypto trading bot data feed: backtest on €1 tick files, then stream the same schema live over WebSocket: event ids, one connection, latency per endpoint, cost.
A crypto trading bot lives on two feeds that are usually two products: the historical data it is researched and backtested on, and the live data it trades on. When the two come from different vendors with different schemas, the bot that passed its backtest is not the bot that goes live — a second parser, a second clock convention and a second idea of what "the book" means sit between them. This guide walks the other way round: start on €1 tick files, go live on the WebSocket feed those files are recordings of, and keep one parser, one schema and one set of assumptions from the first backtest to production. It covers what a bot needs from a feed, how the two sides share a schema, the event-id rules that tell you when you missed data, why one connection per host is the right number, where the latency comes from at each endpoint, and what it costs by the day.
Everything below runs on the CryptoStruct archive: every message Kalshi, Polymarket and 35+ crypto venues publish — every Level-2 snapshot and update at full depth, every trade, every contract, every day since we added the venue — captured co-located with nanosecond venue and receive timestamps and a gap-audited event-id chain, in one normalized schema. It is the same capture our own trading engine and enterprise feeds run on, sold as €1 day bundles you buy as a guest in the Data Shop with instant download — no subscription, no minimum, no sales call. Free full-day samples let you run every command in this guide before paying.
What a trading bot actually needs from a market data feed
Strip the marketing away and a bot needs five things from its feed. The order book at the depth the venue publishes, as a snapshot followed by every change — not a top-of-book sample on a timer, because queue position, imbalance and the shape behind the touch are where the signal lives. Every trade with its aggressor side, size and the venue's trade id, because the tape is what your fills will look like. On derivatives the mark price, the index and the funding rate, because a perpetual's carry is part of the P&L, and liquidations if your model watches forced flow. Two clocks per event — the venue's and the receive time — so you can measure the path instead of guessing it. And a way to know when a message was lost, because a book you cannot trust is worse than no book.
The self-serve feed covers Binance, Coinbase, Kalshi, Polymarket and OKX, and what each market carries follows its capabilities, never the venue name: Binance USDT-M and OKX carry every topic in the table below, Coinbase Spot, Kalshi and Polymarket carry the order book and the trades only. A topic a market does not publish never emits, so a bot subscribes to what exists and the parser handles the union.
| Topic | Messages | What a bot receives |
|---|---|---|
| Order book | Book snapshot (type 0), Book update (type 1) | Full-depth snapshot on subscribe, then changed levels only; every event chained by prevEventId → eventId. |
| Top of book | Top of book (type 6) | Best bid and ask as its own stream; optional coalescing for slow consumers. |
| Trades | Trades (type 2) | Every public trade: aggressor side, price, quantity, trade id, exchange timestamp. |
| Mark price | Mark price (type 7) | The venue's mark price as it is republished. |
| Index price | Index price (type 8) | The underlying index price. |
| Funding rate | Funding rate (type 9) | Current rate, next settlement time and the venue's predicted next rate. |
| Liquidations (opt-in) | Liquidations (type 17) | Liquidation orders as the venue publishes them: side, price, quantity, exchange timestamp — off by default, opt in per subscribe. |
The feed topics and their message types — the same tags the historical day files carry.
One schema: the day files are recordings of the live stream
The historical files in the data shop are not a separate product built from candles or REST pulls — they are the live stream, recorded co-located at the venue and written to disk as it arrived. Same topics, same message types, same field names, same normalized schema across every venue in the historical archive. A book snapshot in a day file is the same snapshot a live subscriber gets on connect; the update that follows it is the same update, with the same event id and the same previous id. The practical consequence is the point of this guide: the parser you write for the backtest is the parser you run in production, and a bug you fix in one place is fixed in both.
Both sides carry two clocks on every event, as integer nanoseconds UTC: the venue's own timestamp as it was published, and the receive timestamp of the co-located recorder. In the live feed the second clock is the time the message reached our capture host next to the exchange; your bot adds its own receive time on top. Keep all three in your event model from the first backtest: the difference between the venue clock and the recorder clock is the venue's publication delay, which you can study offline in any day file, and the difference to your own clock is the line plus your on-ramp, which you can only measure live. A bot that was tuned on venue timestamps alone will be surprised by both.
Backtest first — on the same files, for €1 a day
Start with a free full-day sample of the instrument you trade — it is the exact shop format, one UTC day, every message — and get the parser, the book rebuild and the event-id audit right before paying for anything. Then buy the days you need from the shop: one instrument-day is €1, no subscription, no minimum, downloaded instantly. Every purchased or sample day also exports as flat CSV.gz or Parquet through the ?format= parameter — trades, top-of-book quotes and liquidations with microsecond UTC timestamps — which is the quickest route into pandas, DuckDB or Polars for feature research; the native file keeps the full-depth book for the replay itself.
- Rebuild the book from the snapshot and the updates exactly as the live client will — same code path, same data structure — and assert the chain: every update's previous id must equal the last id you applied. The free reader's audit command does this for a whole file in one go.
- Replay on the venue clock, not the recorder clock, and feed your strategy the events in arrival order; a replay sorted by venue timestamp sees a world no live client ever sees.
- Build features on a time grid from the replayed book — depth imbalance, microprice, spread, queue sizes — the order-book dataset guide walks through the pipeline with leakage-safe labels.
- Model fills against the tape you actually have: the trade's aggressor side tells you which side of the book it consumed, so passive-fill assumptions can be checked instead of assumed.
Go live: endpoint, market, scope, key
The live side is one WebSocket per market. Captured at each venue and delivered at our London, Frankfurt, Ashburn, Ohio or Tokyo endpoints over dedicated lines (Binance and OKX from Tokyo at London, Frankfurt, Ashburn or Ohio; Coinbase from Ashburn at Ohio or Tokyo; Kalshi from Ohio at Tokyo; Polymarket from London at Tokyo). You pick the endpoint nearest to the host your bot runs on, the port selects the market, and one API key per account authenticates every connection. Access is bought by the UTC calendar day — paid days start at the next 00:00 UTC and the rest of the purchase day is included free — so the first evening after checkout is already live. The real-time crypto market data hub lists every venue page with its endpoints and prices; the steps are the same for each.
- Pick the venue page from the hub and the endpoint nearest to your bot's host; the market listens on the same port at every endpoint that serves it (Binance USDT-M (port 14004), Binance Spot (port 14003), Coinbase Spot (port 14006), Kalshi (port 14005), Polymarket (port 14007) and OKX (port 14008)).
- Pick the scope: a single instrument or the whole market (Kalshi and Polymarket are sold as whole markets only). Single OKX instruments: spot and perpetual swaps — dated futures and options come with the whole market. A whole market is the right scope for a bot that rotates instruments.
- Pay by card in the configurator; your API key is created with the first purchase and shown in your account next to the endpoint addresses. Keep it in the
CRYPTOSTRUCT_REALTIME_API_KEYenvironment variable — never in a prompt, a repo or a config file you share. - Connect over a plain WebSocket in JSON or SBE following the connection guide — login, subscribe, keep-alive, reconnect — or hand a coding agent the machine-readable spec and let it write the client.
- Run the strategy in shadow mode first: the live book through the backtest parser, orders logged, nothing sent. The distributions you log in that week are the ones the backtest could not show.
Event ids: how a bot knows it missed something
Every order-book update carries an event id and the id of the update it follows. The rule for a bot is the same on every market and in every day file: if the previous id of an update does not equal the last id you applied, data was missed — unsubscribe, subscribe again and rebuild from the fresh snapshot. What differs per market is what the ids are. On Binance USDT-M, Binance Spot and OKX the book ids follow the venue's own sequence; on Coinbase Spot, Kalshi and Polymarket they are opaque. Either way, compare ids for equality only, in arrival order — never sort or subtract them, and never infer a gap size from the distance between two ids.
| Market | Order-book event ids | Trade event ids |
|---|---|---|
| Binance USDT-M | the venue's own sequence (ORDERED) | the venue's own sequence (ORDERED) |
| Binance Spot | the venue's own sequence (ORDERED) | the venue's own sequence (ORDERED) |
| Coinbase Spot | opaque (UNORDERED) | opaque (UNORDERED) |
| Kalshi | opaque (UNORDERED) | opaque (UNORDERED) |
| Polymarket | opaque (UNORDERED) | opaque (UNORDERED) |
| OKX | the venue's own sequence (ORDERED) | opaque (UNORDERED) |
Event-id behaviour per market, as the capabilities report it. ORDERED = the venue's sequence; UNORDERED = opaque ids.
Drop a random update from a day file during replay and check that your client notices on the very next message and resubscribes. A bot that first detects a gap in production usually detects it in its P&L.
One connection per host
Take all your instruments on one connection. Racing several connections to the same host against each other (feed arbitrage) gains nothing: they are served by the same host from the same feed, so no copy arrives earlier, and the duplicate traffic tends to slow your connection down rather than speed it up. That optimization is already done for you: we run feed arbitrage internally, before the data reaches your endpoint, so every connection delivers the fastest copy we have. Use further connections for independent processes only. A second socket from the same host to the same endpoint receives the same messages a few microseconds apart and competes with the first for the same NIC and the same CPU — it adds jitter, never information. The right architecture for a multi-instrument bot is one connection carrying every subscription, one reader thread, and a lock-free hand-off into the strategy. Independent processes on different hosts are the legitimate reason for further connections, and extra connections are added from your account for exactly that.
Where the latency comes from, per endpoint
The latency a bot sees has three parts: the venue's own publication delay, the dedicated line from the capture host next to the venue to the endpoint, and the hop from the endpoint to your host. The feed quotes the second one — the line — as a one-way figure next to every endpoint, and says so; the other two you measure yourself, offline from the two clocks in any day file and live on your own host. The table lists every line on sale; low-latency Binance market data and the low-latency Binance market data guide go deeper into what the figure covers and which clock to measure against.
| Venue | Endpoint | One-way latency | Shop |
|---|---|---|---|
| Binance | London | ~69.8 ms | Buy a day → |
| OKX | London | ~69.8 ms | Buy a day → |
| Binance | Frankfurt | ~68.2 ms | Buy a day → |
| OKX | Frankfurt | ~68.2 ms | Buy a day → |
| Binance | Ashburn | ~67.7 ms | Buy a day → |
| OKX | Ashburn | ~67.7 ms | Buy a day → |
| Binance | Ohio | ~64.1 ms | Buy a day → |
| Coinbase | Ohio | ~4.6 ms | Buy a day → |
| OKX | Ohio | ~64.1 ms | Buy a day → |
| Coinbase | Tokyo | ~67.7 ms | Buy a day → |
| Kalshi | Tokyo | ~64.1 ms | Buy a day → |
| Polymarket | Tokyo | ~69.8 ms | Buy a day → |
Every figure is one-way line latency from the site where the venue is captured to our endpoint — the venue's own latency and the on-ramp to your client come on top. Source: CryptoStruct line measurements, as quoted in the shop and on the venue pages.
- Binance lines: ~69.8 ms Tokyo → London, ~68.2 ms Tokyo → Frankfurt, ~67.7 ms Tokyo → Ashburn and ~64.1 ms Tokyo → Ohio — one-way line latency from the site where the venue is captured to our endpoint — the venue's own latency and the on-ramp to your client come on top.
- OKX lines: ~69.8 ms Tokyo → London, ~68.2 ms Tokyo → Frankfurt, ~67.7 ms Tokyo → Ashburn and ~64.1 ms Tokyo → Ohio — one-way line latency from the site where the venue is captured to our endpoint — the venue's own latency and the on-ramp to your client come on top.
- Coinbase lines: ~4.6 ms Ashburn → Ohio and ~67.7 ms Ashburn → Tokyo — one-way line latency from the site where the venue is captured to our endpoint — the venue's own latency and the on-ramp to your client come on top.
- Kalshi lines: ~64.1 ms Ohio → Tokyo — one-way line latency from the site where the venue is captured to our endpoint — the venue's own latency and the on-ramp to your client come on top.
- Polymarket lines: ~69.8 ms London → Tokyo — one-way line latency from the site where the venue is captured to our endpoint — the venue's own latency and the on-ramp to your client come on top.
Pick the endpoint by where the bot runs, not by the smallest number: the spread between the lines of a venue is a few milliseconds, while a host on the other side of an ocean from its endpoint adds a hop no line figure accounts for. Then measure on your own host — log the receive time of every trade against the venue timestamp and look at the median and the p99, never a single reading — and only then decide on a longer run.
What it costs, by the day
Research days are €1 per instrument-day in the shop, with free full-day samples to start on. Live access is €49 per instrument-day or €99 per market-day, bought per UTC calendar day and per endpoint; longer runs cost less per day: −10 % from 14 days, −25 % from 30, −37 % from 180, −50 % from 365. Nothing renews — extend from your account when you want more days, stop by doing nothing. Realtime access is paid by card — credits apply to historical data only. Enterprise subscriptions already cover further locations and venues — see the enterprise tiers for order entry and T+1 history.
Limitations
The self-serve feed covers Binance, Coinbase, Kalshi, Polymarket and OKX at the endpoints listed above; enterprise subscriptions already cover further locations and venues. Kalshi and Polymarket are sold as whole markets only. Single OKX instruments: spot and perpetual swaps — dated futures and options come with the whole market. The one-way figures are our line measurements on the stated basis, quoted as approximate values — they are not an SLA, they vary with venue load, and what your bot observes also depends on the path from your host to the endpoint and on your own stack. Market data only: there is no order entry on this product, and nothing here is investment advice.
Why buy the data in this guide here
Four things every page on this site is built on — and the reason the numbers above exist at all.
We record everything
The complete public feed of each venue as it was published: every Level-2 snapshot and update at the venue's full book depth, every trade with its aggressor side, every quote, funding, mark-price and liquidation event — for every instrument the venue lists, every UTC day since we added the venue. Nothing sampled, no top-N cut, no on-demand capture.
Institutional grade
Captured co-located at the venue with the exchange timestamp and our receive timestamp in integer nanoseconds, an event-id chain that makes any gap visible, and one normalized schema across 35+ venues — the same capture our own high-frequency trading engine and enterprise feeds run on.
€1 per instrument-day
Any instrument-day is €1, series-day bundles start at €1 — no subscription, no minimum order, no tiers to unlock. Credit packs lower the effective price and never expire, and every venue has free full-day samples to test against first.
Self-service for everyone
Pick the days in the Data Shop, pay by card as a guest and download immediately — no sales call, no enterprise contract, no KYC. Coding agents buy the same files through the MCP server, and the free Agent Skill teaches them the format.
Agents: /llms.txt · MCP server /mcp · every page as markdown via Accept: text/markdown
Frequently asked questions
What market data does a crypto trading bot need?
The order book at the venue's full depth as a snapshot plus every change, every trade with its aggressor side and trade id, on derivatives the mark price, index and funding rate, two timestamps per event and an event-id chain that reveals a missed message. The self-serve feed carries exactly that per market, and the historical day files carry the same.
Can I backtest on the historical files and go live on the same code?
Yes — the day files are recordings of the live stream: same message types, same field names, same normalized schema across Binance, Coinbase, Kalshi, Polymarket and OKX. The parser and the book rebuild you write for a €1 day file read the live WebSocket unchanged; what changes between the two is only the latency your bot observes.
How does my bot know it missed a message?
Every order-book update names the event id it follows. If that previous id does not equal the last id you applied, data was missed: unsubscribe, subscribe again and rebuild the book from the fresh snapshot. Compare ids for equality only, in arrival order — never sort or subtract them.
Do I need one connection per instrument?
No. Take all your instruments on one connection. Racing several connections to the same host against each other (feed arbitrage) gains nothing: they are served by the same host from the same feed, so no copy arrives earlier, and the duplicate traffic tends to slow your connection down rather than speed it up. That optimization is already done for you: we run feed arbitrage internally, before the data reaches your endpoint, so every connection delivers the fastest copy we have. Use further connections for independent processes only. One connection carries every subscription of a market; a second socket from the same host only duplicates traffic and adds jitter.
What does the live feed cost for a bot?
€49 per instrument-day or €99 per market-day, bought per UTC calendar day — paid days start at the next 00:00 UTC and the rest of the purchase day is included free — and longer runs cost less per day: −10 % from 14 days, −25 % from 30, −37 % from 180, −50 % from 365. Nothing renews. Research days in the shop are €1 per instrument-day. Realtime access is paid by card — credits apply to historical data only.