# Goldsky Polymarket Datasets

Polymarket's on-chain fills and positions, backfilled and streamed into your own database.

*https://predictionmarkets.tools/tools/goldsky-polymarket-datasets · Prediction Market Data APIs*

## Facts

### At a glance

| Field | Value |
| --- | --- |
| Vendor | Goldsky |
| Category | Prediction Market Data APIs |
| Job | historical |
| Website | https://docs.goldsky.com/chains/polymarket |
| Pricing model | usage |
| Free tier | true |
| Open source | false |
| Licence | none |
| Self-hosted | false |
| Tested hands-on | false |
| Last updated | 2026-10-01 |

### Pricing

| Tier | USD | Period |
| --- | --- | --- |
| Starter | 0 USD | — |
| Scale, per extra pipeline worker | 73 USD | usage |
| Scale, per 100,000 rows written | 1 USD | usage |
| Enterprise | on request | — |

### Availability

| Field | Value |
| --- | --- |
| Jurisdictions | global |
| Open to US persons | true |
| KYC required | false |

### Markets

| Field | Value |
| --- | --- |
| Subjects | politics, sports, crypto, macro, culture, business |
| Settlement | crypto |
| Resolved by | none |

### Economics

| Field | Value |
| --- | --- |
| Taker fee | none |
| Maker fee | none |
| Liquidity model | clob |
| Platforms | cli, web |
| AI features | none |

### Interfaces

| Field | Value |
| --- | --- |
| API | false |
| WebSocket | false |
| Scripting | SQL |
| Python | false |
| MCP server | false |
| Export | parquet, json |

### Capabilities

Yes: none

No: charting, screening, order_book, backtesting, automation, live_trading, paper_trading, portfolio_tracking, calibration_scoring, cross_venue, alerts, news, tax_reporting

*Verified: pricing 2026-09-30; availability 2026-09-30; capabilities 2026-09-30.*

## What it is

Polymarket's settlement layer, decoded and delivered to a database you run. Goldsky reads Polygon,
decodes the events Polymarket's exchange and token contracts emit, and publishes them as named
datasets that a Goldsky **Turbo pipeline** can stream — history first, then live — into PostgreSQL,
ClickHouse, MySQL, Kafka, S3 as Parquet, SQS, Pub/Sub or a webhook. Polymarket's own data-resources
page lists Goldsky first among the providers it points builders at for trades, balances, positions
and redemptions.

Four Polymarket datasets are published:

| Dataset | What one row is |
| --- | --- |
| `polymarket.order_filled` | One side of a fill — so a single trade between two orders emits two rows. Carries price, USDC and share amounts, side, maker or taker, counterparty, fee and builder address. |
| `polymarket.orders_matched` | One taker order matched against however many makers — the high-level view of the same trade. |
| `polymarket.user_balances` | The current outcome-token balance per holder and token. |
| `polymarket.user_positions` | Holdings with average entry price, realised PnL and total bought. |

This card is about those datasets, not about Goldsky the company, which also sells RPC, generic
subgraph hosting and a compute product. A reader of this site hires Goldsky for one reason: to own
a complete, current copy of Polymarket's on-chain trading record without writing and operating the
indexer that produces it.

**The subgraphs are not the product any more.** Goldsky's Polymarket page opens with a notice: on
28 April 2026 Polymarket migrated to v2 contracts "and is no longer using subgraphs going forward",
and existing public subgraph endpoints "will return incomplete or incorrect data". Anything still
querying a Polymarket subgraph on Goldsky is reading the pre-cutover world — the failure described in
[where a dashboard gets its data](https://predictionmarkets.tools/guides/where-a-dashboard-gets-its-data), where the chart goes flat
and nothing errors.

## Availability

No geographic restriction appears in Goldsky's customer agreement (last updated 16 April 2026,
Endless Sky, Inc. doing business as Goldsky, governed by California law), and there is no identity
step: an account, a login through the CLI, and a card when you move past the Starter credit. The
data describes Polymarket's international platform, whose trading is closed to US persons; reading
its public chain record is a different act, and nothing on Goldsky's side distinguishes the two.

## Pricing

Usage-metered, per hour, on two meters that matter for these datasets: **pipeline workers** and
**rows written to your sink**.

- **Starter** gives every new team a one-time 100 USD credit, priced from the first unit with no
  free allowances. When the credit hits zero the pipelines pause; you are never billed on Starter.
  The marketing table at goldsky.com/pricing, read on 1 October 2026, still shows Starter with 750
  pipeline worker-hours and 1 million events written free, which contradicts the billing page in
  the documentation; this card follows the documentation.
- **Scale** starts when you add a card. Each month includes one free pipeline worker and one million
  free writes. After that a worker is 0.10 USD an hour — about 73 USD for an always-on month — and a
  medium pipeline is four workers, a large one ten. Writes cost 1 USD per 100,000 up to 100 million
  a month and 0.10 USD per 100,000 beyond that.
- **Enterprise** is a contract.

**The write meter is the one to model, and Goldsky gives you the number.** Its Polymarket page warns
that user positions "may be up to 1.2B entities to backfill, and up to 150M entities monthly to
maintain". Worked through the published write rates, and assuming the backfill lands in one billing
month: the first million free, 99 million at 1 USD per 100,000 (990 USD), and 1.1 billion at 0.10
USD per 100,000 (1,100 USD) — about 2,090 USD in writes for the positions backfill alone. Maintaining
150 million a month is about 1,040 USD. Worker hours and your own database are on top. Filtering in
the pipeline's SQL transform before the sink is how that bill comes down.

## Integrations

Pipelines are YAML files deployed with the Goldsky CLI (`goldsky turbo apply`), with a `sql`
transform stage between source and sink — which is the one place you write code. The v2 Polymarket
datasets support a fast scan over a `block_number` range, so a specific period can be backfilled
without replaying from genesis; filtering by timestamp is not supported for that scan.

There is no query API to call. The product is the pipeline; the interface you read from is whatever
database you pointed it at. Goldsky does publish an MCP server, but it serves Goldsky's documentation
to an assistant, not these datasets.

Delivery is **at-least-once**. Goldsky's delivery-guarantees page says so and tells you to make the
sink idempotent — a primary key on Postgres, `ReplacingMergeTree` on ClickHouse — because a crash
between a sink write and a source commit replays the in-flight batch. Webhook, S3, SQS and Pub/Sub
sinks have no deduplication of their own. Pipelines are also reorg-aware: a reorganised block is
walked back with deletes and updates, so a naive append-only consumer will see rows disappear.

## Limitations

- **On-chain only.** No order book, no quotes, no cancelled orders, no market titles or rules —
  Polymarket matches off chain and settles on chain, and these datasets are the settlement half.
  Market metadata comes from the [Gamma API](https://predictionmarkets.tools/tools/polymarket-gamma-api), keyed by the token IDs in
  the `asset` column.
- **Duplicates are your problem.** At-least-once delivery into a sink with no key produces double
  counts in exactly the volume figures people build on this.
- **Two rows per fill.** `order_filled` reports each fill from both sides; summing it without
  choosing a side doubles volume.
- **The PnL is Goldsky's arithmetic.** `user_positions` carries average price and realised PnL
  computed by the dataset; the page gives the columns, not the cost-basis method.
- **Cost scales with the whole platform.** There is no per-market product: the datasets are
  Polymarket-wide, and the meter runs on rows written.
- **The subgraph era is over.** Any tutorial that points you at a Polymarket subgraph URL predates
  28 April 2026.

## Alternatives

For the same on-chain record queried in SQL without running anything,
[Polymarket dashboards on Dune](https://predictionmarkets.tools/tools/dune-polymarket-dashboards) is the batch form. For a free,
already-decoded answer to wallet and position questions, Polymarket's own
[Data API](https://predictionmarkets.tools/tools/polymarket-data-api) serves positions, trades and PnL per wallet, rate-limited but
keyless. For a one-off historical copy rather than a live pipeline, the
[SII Polymarket dataset](https://predictionmarkets.tools/tools/sii-polymarket-data) is a free Parquet download — of the pre-cutover
contracts only.

## FAQ

### Do the Polymarket subgraphs on Goldsky still work?

Not reliably. Goldsky's own Polymarket page says that on 28 April 2026 Polymarket moved to v2 contracts and stopped using subgraphs, and that existing public subgraph endpoints return incomplete or incorrect data. The recommended route is a Turbo pipeline reading the v2 datasets.

### What Polymarket data does Goldsky provide?

Four decoded datasets — Order Filled, one row per side of every fill; Orders Matched, one row per taker match; User Balances, every outcome-token holding; and User Positions, holdings with average price and realised PnL. All are on-chain events, so there is no order book and no market metadata.

### How much does it cost to backfill Polymarket positions?

Goldsky warns that user positions can be up to 1.2 billion entities to backfill and up to 150 million a month to maintain. At the published write rates that is roughly 2,090 USD of writes for the backfill and about 1,040 USD a month to maintain, before worker hours.

### Is there a free way to try it?

Every new team gets a one-time 100 USD Starter credit. It is drawn from the first unit of usage and does not reset; when it runs out the pipelines pause and nothing is charged. A full Polymarket backfill will not fit in it.

## Background

- [Where a dashboard gets its data, and what it loses when that moves](https://predictionmarkets.tools/guides/where-a-dashboard-gets-its-data.md) — A venue API, an indexer or the chain itself. Which one a dashboard reads decides how far back its history goes and what vanishes when that source changes.

- [What a trade actually costs on a prediction market](https://predictionmarkets.tools/guides/what-a-trade-actually-costs.md) — Venues publish trading fees in four incompatible units, so 3.00% on one can be cheaper than 1.75% on another. How to convert them, and what else takes a cut.
- [Where weather data comes from and what you may do with it](https://predictionmarkets.tools/guides/where-weather-data-comes-from.md) — Observations, model output and archives are three products with three licences. The forecast is the cheap half; a clean observation history is not.

## Also worth comparing

- [Marketlens](https://predictionmarkets.tools/tools/marketlens.md) — Recorded Polymarket order books across every market type, replayable to the tick.
- [Polymarket Data (SII dataset)](https://predictionmarkets.tools/tools/sii-polymarket-data.md) — The "107 GB" open Parquet dump of Polymarket CLOB fills, November 2022 to March 2026.
- [Polymarket Data API](https://predictionmarkets.tools/tools/polymarket-data-api.md) — Polymarket's keyless read API for trades, positions, holders and price history.
- [DepthFeed](https://predictionmarkets.tools/tools/depthfeed.md) — The recorded bid-ask ladder for crypto up-or-down markets, which no venue keeps itself.
- [PolyOrderbooks](https://predictionmarkets.tools/tools/polyorderbooks.md) — Recorded L2 books for Polymarket crypto markets, priced by how far back you may look.
- [Predexon](https://predictionmarkets.tools/tools/predexon.md) — Tick-level book history as Parquet, billed by the gigabyte, plus a mempool-aware feed.
