Goldsky Polymarket Datasets
Polymarket's on-chain fills and positions, backfilled and streamed into your own database.
by Goldsky
Last updated
What it is
Polymarket's settlement layer, decoded and delivered to a database you run. Goldsky reads Polygon, decodes the events Polymarket's exchange and token contracts emit, and publishes them as named datasets that a Goldsky Turbo pipeline can stream — history first, then live — into PostgreSQL, ClickHouse, MySQL, Kafka, S3 as Parquet, SQS, Pub/Sub or a webhook. Polymarket's own data-resources page lists Goldsky first among the providers it points builders at for trades, balances, positions and redemptions.
Four Polymarket datasets are published:
| Dataset | What one row is |
|---|---|
polymarket.order_filled | One side of a fill — so a single trade between two orders emits two rows. Carries price, USDC and share amounts, side, maker or taker, counterparty, fee and builder address. |
polymarket.orders_matched | One taker order matched against however many makers — the high-level view of the same trade. |
polymarket.user_balances | The current outcome-token balance per holder and token. |
polymarket.user_positions | Holdings with average entry price, realised PnL and total bought. |
This card is about those datasets, not about Goldsky the company, which also sells RPC, generic subgraph hosting and a compute product. A reader of this site hires Goldsky for one reason: to own a complete, current copy of Polymarket's on-chain trading record without writing and operating the indexer that produces it.
The subgraphs are not the product any more. Goldsky's Polymarket page opens with a notice: on 28 April 2026 Polymarket migrated to v2 contracts "and is no longer using subgraphs going forward", and existing public subgraph endpoints "will return incomplete or incorrect data". Anything still querying a Polymarket subgraph on Goldsky is reading the pre-cutover world — the failure described in where a dashboard gets its data, where the chart goes flat and nothing errors.
Availability
No geographic restriction appears in Goldsky's customer agreement (last updated 16 April 2026, Endless Sky, Inc. doing business as Goldsky, governed by California law), and there is no identity step: an account, a login through the CLI, and a card when you move past the Starter credit. The data describes Polymarket's international platform, whose trading is closed to US persons; reading its public chain record is a different act, and nothing on Goldsky's side distinguishes the two.
Pricing
Usage-metered, per hour, on two meters that matter for these datasets: pipeline workers and rows written to your sink.
- Starter gives every new team a one-time 100 USD credit, priced from the first unit with no free allowances. When the credit hits zero the pipelines pause; you are never billed on Starter. The marketing table at goldsky.com/pricing, read on 1 October 2026, still shows Starter with 750 pipeline worker-hours and 1 million events written free, which contradicts the billing page in the documentation; this card follows the documentation.
- Scale starts when you add a card. Each month includes one free pipeline worker and one million free writes. After that a worker is 0.10 USD an hour — about 73 USD for an always-on month — and a medium pipeline is four workers, a large one ten. Writes cost 1 USD per 100,000 up to 100 million a month and 0.10 USD per 100,000 beyond that.
- Enterprise is a contract.
The write meter is the one to model, and Goldsky gives you the number. Its Polymarket page warns that user positions "may be up to 1.2B entities to backfill, and up to 150M entities monthly to maintain". Worked through the published write rates, and assuming the backfill lands in one billing month: the first million free, 99 million at 1 USD per 100,000 (990 USD), and 1.1 billion at 0.10 USD per 100,000 (1,100 USD) — about 2,090 USD in writes for the positions backfill alone. Maintaining 150 million a month is about 1,040 USD. Worker hours and your own database are on top. Filtering in the pipeline's SQL transform before the sink is how that bill comes down.
Integrations
Pipelines are YAML files deployed with the Goldsky CLI (goldsky turbo apply), with a sql
transform stage between source and sink — which is the one place you write code. The v2 Polymarket
datasets support a fast scan over a block_number range, so a specific period can be backfilled
without replaying from genesis; filtering by timestamp is not supported for that scan.
There is no query API to call. The product is the pipeline; the interface you read from is whatever database you pointed it at. Goldsky does publish an MCP server, but it serves Goldsky's documentation to an assistant, not these datasets.
Delivery is at-least-once. Goldsky's delivery-guarantees page says so and tells you to make the
sink idempotent — a primary key on Postgres, ReplacingMergeTree on ClickHouse — because a crash
between a sink write and a source commit replays the in-flight batch. Webhook, S3, SQS and Pub/Sub
sinks have no deduplication of their own. Pipelines are also reorg-aware: a reorganised block is
walked back with deletes and updates, so a naive append-only consumer will see rows disappear.
Limitations
- On-chain only. No order book, no quotes, no cancelled orders, no market titles or rules —
Polymarket matches off chain and settles on chain, and these datasets are the settlement half.
Market metadata comes from the Gamma API, keyed by the token IDs in
the
assetcolumn. - Duplicates are your problem. At-least-once delivery into a sink with no key produces double counts in exactly the volume figures people build on this.
- Two rows per fill.
order_filledreports each fill from both sides; summing it without choosing a side doubles volume. - The PnL is Goldsky's arithmetic.
user_positionscarries average price and realised PnL computed by the dataset; the page gives the columns, not the cost-basis method. - Cost scales with the whole platform. There is no per-market product: the datasets are Polymarket-wide, and the meter runs on rows written.
- The subgraph era is over. Any tutorial that points you at a Polymarket subgraph URL predates 28 April 2026.
Alternatives
For the same on-chain record queried in SQL without running anything, Polymarket dashboards on Dune is the batch form. For a free, already-decoded answer to wallet and position questions, Polymarket's own Data API serves positions, trades and PnL per wallet, rate-limited but keyless. For a one-off historical copy rather than a live pipeline, the SII Polymarket dataset is a free Parquet download — of the pre-cutover contracts only.
Specs
- Interfaces
- SQL
- Export
- Parquet, JSON
- Available in
- Global
- KYC required
- No
- Market subjects
- Politics, Sports, Crypto, Macro, Culture, Business
- Resolved by
- Resolves nothing
- Maker fee
- None
- Platforms
- CLI, Web
- AI features
- None
- Pricing verified
- Availability verified
- Capabilities verified
Background
How this part of the sector works, rather than which product to pick.
- Where a dashboard gets its data, and what it loses when that moves — A venue API, an indexer or the chain itself. Which one a dashboard reads decides how far back its history goes and what vanishes when that source changes.
- What a trade actually costs on a prediction market — Venues publish trading fees in four incompatible units, so 3.00% on one can be cheaper than 1.75% on another. How to convert them, and what else takes a cut.
- Where weather data comes from and what you may do with it — Observations, model output and archives are three products with three licences. The forecast is the cheap half; a clean observation history is not.
Also worth comparing
- Marketlens — Recorded Polymarket order books across every market type, replayable to the tick.
- Polymarket Data (SII dataset) — The "107 GB" open Parquet dump of Polymarket CLOB fills, November 2022 to March 2026.
- Polymarket Data API — Polymarket's keyless read API for trades, positions, holders and price history.
- DepthFeed — The recorded bid-ask ladder for crypto up-or-down markets, which no venue keeps itself.
- PolyOrderbooks — Recorded L2 books for Polymarket crypto markets, priced by how far back you may look.
- Predexon — Tick-level book history as Parquet, billed by the gigabyte, plus a mempool-aware feed.
FAQ
Do the Polymarket subgraphs on Goldsky still work?
Not reliably. Goldsky's own Polymarket page says that on 28 April 2026 Polymarket moved to v2 contracts and stopped using subgraphs, and that existing public subgraph endpoints return incomplete or incorrect data. The recommended route is a Turbo pipeline reading the v2 datasets.
What Polymarket data does Goldsky provide?
Four decoded datasets — Order Filled, one row per side of every fill; Orders Matched, one row per taker match; User Balances, every outcome-token holding; and User Positions, holdings with average price and realised PnL. All are on-chain events, so there is no order book and no market metadata.
How much does it cost to backfill Polymarket positions?
Goldsky warns that user positions can be up to 1.2 billion entities to backfill and up to 150 million a month to maintain. At the published write rates that is roughly 2,090 USD of writes for the backfill and about 1,040 USD a month to maintain, before worker hours.
Is there a free way to try it?
Every new team gets a one-time 100 USD Starter credit. It is drawn from the first unit of usage and does not reset; when it runs out the pipelines pause and nothing is charged. A full Polymarket backfill will not fit in it.