Polymarket Data (SII dataset)

The "107 GB" open Parquet dump of Polymarket CLOB fills, November 2022 to March 2026.

by SII-WANGZJ

Last updated

US persons
Yes
Taker fee
None
Settlement
Crypto
Liquidity
CLOB

What it is

The dataset people mean when they say "the open 107 GB Polymarket dump". A team at the Shanghai Innovation Institute, with co-authors from Westlake, Shanghai Jiao Tong, Harbin Institute of Technology and Fudan, decoded every OrderFilled event from Polymarket's two original CLOB exchange contracts on Polygon, linked each fill to its market through the Gamma API, and published the result as Parquet on Hugging Face together with the MIT-licensed Python pipeline that produced it.

It is published from SII-WANGZJ, the personal account of its first author, Zhengjie Wang, on both Hugging Face and GitHub, not from an account of the institute, and the licence names "Polymarket Data Contributors" as the copyright holder. The institute is an affiliation, not the publisher.

Five files, each a different cut of the same fills:

FileWhat it is
orderfilled.parquetThe raw decoded events: block, transaction, contract, maker and taker, asset IDs, filled amounts, fees, order hash.
trades.parquetThe same fills with market linkage, outcome name, price and USD amount — the card recommends starting here.
quant.parquetFills normalised to the YES token, contract-side rows filtered out, for price and time-series work.
users.parquetEach fill split into a maker row and a taker row, all expressed as buys with signed amounts, sorted by wallet.
markets.parquetMarket metadata from Gamma.

It is a Polymarket settlement record, and nothing else: no Kalshi, no order books, no quotes. The dataset card's own coverage statement is the CLOB era only — first record 21 November 2022, earlier AMM-era trading from 2020 to 2022 not included — through 4 March 2026.

Availability

A public, ungated Hugging Face dataset: hf download SII-WANGZJ/Polymarket_data --repo-type dataset, no account required, no jurisdiction attached. The practical gate is disk. The Hugging Face file listing on 30 September 2026 showed orderfilled.parquet at 110.3 GB, users.parquet at 47.7 GB, trades.parquet at 37.5 GB, quant.parquet at 36.7 GB and markets.parquet at 0.3 GB — about 232 GB in all, against the 107 GB the GitHub README still advertises and the 163 GB on the card. The files can be downloaded individually, and trades.parquet alone answers most questions.

Pricing

Free, with nothing for sale. The dataset card says "MIT License - Free for commercial and research use" and links the GitHub repository's LICENSE, which reads "MIT License, Copyright (c) 2026 Polymarket Data Contributors" — checked on 30 September 2026. Hugging Face's own licence metadata field is empty, so the grant for the data rests on that sentence and that link. The README adds that users are responsible for complying with Polymarket's terms.

Markets & resolution

Nothing here resolves anything, and there is no resolution column in the documented schemas: a fill is a fill. Which way a market settled has to come from Polymarket — the Gamma API for the market record or the Data API for resolution state — joined on the market ID.

Integrations

Parquet, read with anything that reads Parquet; the README's examples are pandas. The GitHub repository is the collection pipeline: a Polygon log fetcher, an event decoder, the Gamma market fetcher and the cleaning steps that produce the derived files, with a "continuous" mode that polls new blocks every two seconds. An Alchemy key is optional for a faster RPC.

Limitations

The pipeline cannot see anything after 28 April 2026. It decodes the two exchange contracts listed in the README (0x4bFb…982E and 0xC5d5…f80a). Polymarket moved to new exchange contracts and a new collateral token that day. Issue #4, "Migrate to CTF Exchange V2", has been open since 30 April 2026 with no reply. A second request of the same kind (#8), opened on 19 May 2026 and reporting that the pipeline now extracts zero trades, was closed by its own reporter the next day with no comment from the maintainer and no code change. The "continuous mode" in the README is therefore a way to record nothing.

The code has not changed since the day it was published. Every commit on the default branch is dated 1 January 2026, and there are no tags or releases. A pull request fixing silently skipped block ranges (#7) was closed by its own author a minute after opening it, on 17 May 2026, and never reviewed. The dataset on Hugging Face has moved on — files were replaced on 20 and 21 July 2026 — but the repository that is supposed to reproduce it has not.

Completeness has been disputed in public, and the fix is a comment. Issue #1 on GitHub, opened 12 March 2026, compared the raw files with Goldsky's subgraph and found whole runs of 100 blocks missing, with up to 26.8% of a day's fills absent. A Hugging Face discussion opened on 12 April 2026 reported about 21% of 2024's blocks missing from the raw files. On 18 June 2026 the maintainer replied there "I have solved this problem", with no changelog, no version number and no note on the card. The card still says "no missing blocks or gaps". If completeness matters to your result, count fills against another source for the window you use.

Timestamps are reported as wrong. Issue #3, open since 9 April 2026, found the timestamp in orderfilled.parquet disagreeing with the Polygon block timestamp in 100 of 100 sampled fills. Use block_number and look the time up yourself.

The size, record counts and coverage dates disagree across its own pages. README, card and file listing give three different totals; the card's "Last Updated" of 5 March 2026 predates the July file replacement, and what the replaced files cover is not stated anywhere.

No order books. A Hugging Face discussion asking for them, opened 21 April 2026, has no reply. Depth is not recoverable from fills.

Alternatives

Prediction Market Analysis is the other free archive, smaller, covering Kalshi as well as Polymarket, with analysis scripts attached. For Polymarket fills that continue past April 2026, the Goldsky Polymarket datasets are the maintained version of the same decoding, paid by the row, and Polymarket dashboards on Dune put it behind SQL. For order-book history, which no fill dataset contains, see Marketlens.

Specs

Interfaces
Python
Export
Parquet
Available in
Global
KYC required
No
Market subjects
Politics, Sports, Crypto, Macro, Culture, Business
Resolved by
Resolves nothing
Maker fee
None
Platforms
Library, CLI
AI features
None
Pricing verified
Availability verified
Markets verified

Background

How this part of the sector works, rather than which product to pick.

Also worth comparing

  • Goldsky Polymarket Datasets — Polymarket's on-chain fills and positions, backfilled and streamed into your own database.
  • Marketlens — Recorded Polymarket order books across every market type, replayable to the tick.
  • Polymarket Data API — Polymarket's keyless read API for trades, positions, holders and price history.
  • DepthFeed — The recorded bid-ask ladder for crypto up-or-down markets, which no venue keeps itself.
  • PolyOrderbooks — Recorded L2 books for Polymarket crypto markets, priced by how far back you may look.
  • Predexon — Tick-level book history as Parquet, billed by the gigabyte, plus a mempool-aware feed.

FAQ

Is this the same dataset as Prediction Market Analysis?

No. Prediction Market Analysis is Jonathan Becker's framework with a roughly 36 GiB archive of Kalshi and Polymarket trades. This one is a separate Polymarket-only dataset from a team at the Shanghai Innovation Institute, decoded from Polygon OrderFilled events, and it has no Kalshi data at all.

How big is it really?

Bigger than the headline. The GitHub README says 107 GB and 1.1 billion records; the Hugging Face card says 163 GB and 1.9 billion; the five files listed on Hugging Face on 30 September 2026 add up to about 232 GB, most of it the raw OrderFilled file at 110 GB.

Does it include trades after Polymarket's April 2026 contract upgrade?

Not from the published pipeline. The code decodes the two original exchange contracts, and an issue asking for CTF Exchange V2 support has been open since 30 April 2026 without a reply. The dataset card's stated coverage ends on 4 March 2026.

Can I use it commercially?

The dataset card says MIT, free for commercial and research use, and links the repository's MIT licence. The Hugging Face metadata carries no licence field, so the grant rests on that prose and that file. The underlying rows are public blockchain events.