Training Data for AI & Machine-Learning Trading Models — Crypto Order Books, Trades and Event Contracts
Models learn from what happened between the candles: every order placed, changed and cancelled, every trade with its aggressor side, and — on Kalshi and Polymarket — the settlement that labels each window. The CryptoStruct archive records every message of each venue's public feed, co-located and unsampled, in one schema across 35+ venues, and sells it per instrument-day as the native day file — the venue's full book depth, every update, every trade. That file is the training set; a free Python reader streams it event by event or, if you want a tabular view, derives a fixed grid at any depth from it. Any instrument-day is €1, bought as a guest with instant download — no subscription, no sales call.
What a model sees in tick data that candles hide
The complete venue feed as recorded: every Level-2 snapshot and update at the venue's full book depth, every trade, quote, funding and liquidation event — queue size, imbalance, spread and cancellations are features in the data, not lost in aggregation.
Every Kalshi and Polymarket up/down window settles: a 5-minute series yields 288 labeled windows per day, and the last print of each contract implies the outcome the model must learn to anticipate.
A parser written for a Binance perpetual reads Deribit options and Kalshi contracts unchanged — one feature pipeline for spot, derivatives, options and event contracts.
Venue timestamp and our receive time at the venue on every event, so latency-aware features and leak-free alignment of book and tape are possible.
How big a training set is — rows, windows and bytes
Across the whole archive our capture records about 646.4M trade prints per UTC day (19.4B in the 30 days to 2026-10-08, $5.6T of USD turnover across 33 venues) — every one of them a labelled event with its side, size, price and the order-book state around it. The rows below are what a single instrument contributes.
The native day file is the dataset: the venue's full book depth and every update, exactly as published. The grids are optional views the free reader derives from it at whatever depth and sampling you choose.
| Dataset | Per day | Per year | Notes |
|---|---|---|---|
| Native day file, Binance BTCUSDT perpetual — full book depth, every update | tens of millions of events | billions | 1–2 GB compressed per day; this file IS the dataset — stream it with the reader, never load it whole |
| Native day file, BitMEX XRPUSDT perpetual — full book depth, every update | 1.87 M events | ~680 M | 59 MB compressed per day, 1.83 M of them book updates |
| Derived L2 grid (optional), 1-second grid — depth is a parameter (20 shown) | 86,400 rows | 31.5 M rows | a tabular view the free reader builds from the native file; the grid drops the updates between grid points — the native file keeps them |
| Derived L2 grid (optional), 100-millisecond grid — depth is a parameter | 864,000 rows | 315 M rows | per instrument; choose depth and grid from the horizon you predict, or skip the grid and read events |
| Labeled windows, 5-minute up/down series | 288 | 105,120 | e.g. Polymarket BTC 5m: ~577 contract files and ~554k trades per day |
| Labeled windows, 15-minute / hourly series | 96 / 24 | 35,040 / 8,760 | the label budget is fixed by the market clock, not by your features |
Where the labels come from — and how not to leak them
For crypto instruments the label is yours to define — a forward mid-price move over a horizon with a dead-band, the sign of the next N-second return, whether a level gets consumed — and the L2 grid gives you the inputs at any sampling rate. Because every event carries both the venue timestamp and the co-located receive timestamp, features can be built strictly from information available before the label's start, which is where most leakage in published crypto models comes from.
For event contracts the label comes with the data: every Kalshi and Polymarket up/down window settles, and the last print of a contract implies the outcome (in our statistics a close at or above 0.97 reads as Yes/Up, at or below 0.03 as No/Down — always described as "implied", the venue's official resolution is a separate record you should join). A 5-minute series produces 288 labeled windows per UTC day, 105,120 per year; the training-data guide works through that budget, walk-forward splits and the purge/embargo needed around each window.
Split by time, never at random: adjacent windows share regime and order flow, and a random split leaks that structure across train and test. Trim the ~8-minute overlap at the start of every day file before concatenating days, audit each file (stats FILE --deep: parse errors, chain gaps, crossed states, coverage close to 86,400 s) and drop grid rows flagged suspect before drawing conclusions from microstructure features.
Read the native file — or derive a feature grid from it in two commands
Sequence and event models read the native day file directly — iter_events in the free reader from the AI toolkitstreams a multi-gigabyte day at the venue's full depth, every update included. If you want a tabular view instead, the same reader writes a fixed-grid order-book table (depth and sampling are parameters) or the trade tape as Parquet; polars or pandas takes it from there. Days are independent, so the conversion parallelizes trivially.
# optional: a fixed-grid view of the book (depth and sampling are parameters —
# the native day file keeps every level and every update)
for f in data/*.txt.zst; do
out="parquet/book_$(basename "$f" .txt.zst).parquet"
[ -f "$out" ] || python3 cryptostruct_reader.py book "$f" --every 1s --depth 20 --out "$out"
done
# trades of the same days, for flow features and labels
for f in data/*.txt.zst; do
out="parquet/trades_$(basename "$f" .txt.zst).parquet"
[ -f "$out" ] || python3 cryptostruct_reader.py trades "$f" --out "$out"
done
import polars as pl
book = pl.scan_parquet("parquet/book_*.parquet")
# columns: ts, mid, spread, spread_bps, bid_px_1..20, bid_qty_1..20,
# ask_px_1..20, ask_qty_1..20, n_events, suspect
feat = (
book.filter(pl.col("suspect") == 0)
.with_columns(
imb1=(pl.col("bid_qty_1") - pl.col("ask_qty_1"))
/ (pl.col("bid_qty_1") + pl.col("ask_qty_1")),
micro=(pl.col("bid_px_1") * pl.col("ask_qty_1") + pl.col("ask_px_1") * pl.col("bid_qty_1"))
/ (pl.col("bid_qty_1") + pl.col("ask_qty_1")),
)
.with_columns(y=(pl.col("mid").shift(-60) / pl.col("mid") - 1)) # 60 s forward return
)
df = feat.collect()
The full walkthrough — grid choice, features, labels, splits and audit — is the order-book dataset guide.
How to assemble and buy a training set
- Start with the free samples: a pinned high-volatility Bitcoin day (Binance BTCUSDT perpetual, 2026-07-06, ~1.5 GB), rolling full days on major venues and the busiest Kalshi and Polymarket contracts of the day are on the downloads page — build the pipeline against them before buying anything.
- Size the set from the label budget, not from gigabytes: pick the instruments and the number of days that give you enough independent labeled windows, then buy exactly those instrument-days in the Data Shop at €1 each, or whole series-days of an event-contract family on the bundles page.
- Buy in bulk when the set is large: "Select to buy" in the instrument scanner adds a shared date window across many instruments, credit packs bring the effective per-day price down (€100 → 140 credits · €250 → 375 credits · €1,000 → 1,750 credits) and never expire, and Premium adds 50 credits a month for €20.
- Download everything you own as a batch ZIP from your account, bundle days as one archive each, or date ranges as one resumable tar; convert with the free reader (grids, trades, funding, greeks to Parquet) or use
?format=trades.parquetdirectly on the download link. - Let an agent do it: the Agent Skill teaches Claude Code or Cursor the format and recipes, and the MCP server searches instruments and bundles, previews and — with your consent — buys and downloads on your behalf.
What you may do with the data
Purchased data is for your own use — backtesting, research, analytics and training your own models, in-house or in your own products; the trained models and derived data you build from it are yours. The raw files and tick streams themselves may not be redistributed, resold or published. The full terms are on the data license page.
The archive itself — venues, message types, prices and bulk delivery — is described on the historical data overview.
Why buy training data here
Four things every page on this site is built on — and the reason the numbers above exist at all.
We record everything
The complete public feed of each venue as it was published: every Level-2 snapshot and update at the venue's full book depth, every trade with its aggressor side, every quote, funding, mark-price and liquidation event — for every instrument the venue lists, every UTC day since we added the venue. Nothing sampled, no top-N cut, no on-demand capture.
Institutional grade
Captured co-located at the venue with the exchange timestamp and our receive timestamp in integer nanoseconds, an event-id chain that makes any gap visible, and one normalized schema across 35+ venues — the same capture our own high-frequency trading engine and enterprise feeds run on.
€1 per instrument-day
Any instrument-day is €1, series-day bundles start at €1 — no subscription, no minimum order, no tiers to unlock. Credit packs lower the effective price and never expire, and every venue has free full-day samples to test against first.
Self-service for everyone
Pick the days in the Data Shop, pay by card as a guest and download immediately — no sales call, no enterprise contract, no KYC. Coding agents buy the same files through the MCP server, and the free Agent Skill teaches them the format.
Agents: /llms.txt · MCP server /mcp · every page as markdown via Accept: text/markdown
Training data — FAQ
Can I train AI or machine-learning models on this data?
Yes — purchased data may be used in-house for research, backtesting and training your own models, and the trained models and derived analytics are yours. The raw files and tick streams themselves may not be redistributed, resold or published.
Why is tick data better training data than OHLCV candles?
Candles summarize a bar and discard the order book, the sequence of trades and every cancellation. Tick files keep the full Level-2 evolution and every print with its aggressor side and nanosecond timestamps, so imbalance, microprice, depth, spread dynamics and trade flow are features in the data, and labels can be aligned to the exact moment information became available.
Are the datasets labeled?
Event-contract data is: every Kalshi and Polymarket up/down window settles, so each window is a labeled example — a 5-minute series yields 288 per day — with the outcome implied by the last print in our statistics and the official resolution as the authoritative label. Crypto order-book and trade data is unlabeled; forward returns or level-consumption labels are built from the same file.
How much training data does a trading model need?
Count independent labeled windows, not gigabytes: a 5-minute up/down series gives 105,120 windows per year, a 15-minute series 35,040, an hourly one 8,760, and for crypto instruments an L2 grid at one second yields 86,400 rows per instrument-day. The up/down training-data guide works through the arithmetic and the walk-forward splits.
Is the data available as Parquet or CSV?
Yes — every free sample and every purchased tick day exports as gzipped CSV or Parquet at no extra cost: trades and liquidations (derivatives, 2026 onward) in both formats, top-of-book quotes as CSV where the venue publishes a BBO stream, with integer microsecond UTC timestamps. Order-book grids at any sampling rate come out of the free reader as Parquet. The native day file stays the complete record — full book depth, every update, nanosecond timestamps — and is what a sequence model should read; the exports and grids are convenience views of it.
How much does a training set cost?
€1 per instrument-day and series-day bundles from €1 per day, no minimum; credit packs (€100 → 140 credits · €250 → 375 credits · €1,000 → 1,750 credits) never expire, and Premium adds 50 credits a month for €20. A year of one liquid perpetual is 365 instrument-days; a year of a 5-minute event-contract series is 365 series-days holding every contract of every window.