# Why your bot gets throttled, and what the limit is counted against

Kalshi meters tokens per account tier; Polymarket's order book meters requests per IP and queues the excess. What each counts, and what the libraries assume.

*https://predictionmarkets.tools/guides/why-your-bot-gets-throttled · background to Trading Clients, SDKs & Bots*

**Answer:** Because each venue counts something different, and your client library usually counts something else again. Kalshi gives every account two token buckets, read and write, sized by a tier earned from trading volume; an order costs 10 tokens and a cancel 2. Polymarket's international order book limits requests per IP address per endpoint and delays the excess rather than refusing it. The Python and TypeScript clients here throttle by a single request rate, or not at all.

A process that ran cleanly for a week starts getting `429` back, or starts taking two seconds to
place an order that used to take forty milliseconds. Nothing in the code changed. What changed is
that the process crossed a line it did not know it was near, and it did not know because the line
is not drawn in requests per second. It is drawn in tokens, or per IP address, or per endpoint, or
per tier that moved — and the library underneath the process may be counting in a unit the venue
does not use at all.

This page covers the two venues most bots in this catalogue are written against, Kalshi and
Polymarket's international order book, and the client libraries that sit between your code and
them. Polymarket US, Adjacent, Manifold and Limitless publish limits of their own, counted four
more ways; those are read in [what running a bot does not solve](https://predictionmarkets.tools/guides/what-a-bot-cannot-fix),
together with the retry loop that backs off when it should not. Everything below was read from the
venues' documentation and the libraries' source on 27 September 2026. None of it has been run
against a funded account by this site.

## How it works

### Kalshi: two buckets per account, and a tier you earn

The [Kalshi API](https://predictionmarkets.tools/tools/kalshi-api)'s rate-limit page does not count requests. It counts tokens: "Every
authenticated request costs **tokens**. Your tier sets your **budget**: the rate, in tokens per
second, at which your balance refills." Most requests cost 10. A cancel costs 2 — the page's own
example prices a 25-order batch cancel at 50 tokens against 250 for a 25-order batch create. The
authoritative list of anything that is not 10 is an endpoint, `GET /account/endpoint_costs`, not
a table on the page, so it can change without the page changing.

There are **two buckets, and the split is by what the request does**, not by how it arrives. Read
covers `GET` endpoints and anything not routed elsewhere. Write covers order placement, amends,
cancels, order groups, the request-for-quote flow and block-trade accepts. "REST and FIX requests
drain the same buckets", so moving the order path to FIX buys no headroom by itself. Perpetual
futures run in separate buckets that event-contract traffic never touches.

**The buckets are small at the bottom and large at the top.** Per-second budgets, read and write:

| Tier | Read | Write | How you get it |
| --- | --- | --- | --- |
| Basic | 200 | 100 | Complete account signup |
| Advanced | 300 | 300 | One call to an upgrade endpoint |
| Expert | 600 | 600 | 0.075% volume share to earn, 0.05% to keep |
| Premier | 1,200 | 1,200 | 0.125% to earn, 0.10% to keep |
| Paragon | 2,400 | 2,400 | 0.25% to earn, 0.20% to keep |
| Prime | 4,800 | 4,800 | 0.50% to earn, 0.40% to keep |
| Prestige | 12,000 | 9,600 | 1.00% to earn, 0.80% to keep |

Divide by the cost. At Basic, 100 write tokens a second is **10 orders a second** or 50 cancels;
200 read tokens is 20 reads. The step to Advanced is almost free and most people never take it:
the upgrade endpoint grants a permanent Advanced tier if "at least 1 of the user's last 100
Predictions orders was created via API", which triples the write budget to 30 orders a second.

Everything above Advanced is earned from volume, and the arithmetic is worth reading slowly. Your
share is your trailing 30-day volume, "counting both sides of every trade you are part of, as
maker and as taker", divided by twice the previous calendar month's exchange volume. Kalshi reviews
it once a day and grants the tier for 30 days; falling below the keep threshold does not drop you
at once, it lets the current grant run out. Your present tier and every active grant are readable
from `GET /account/limits`. **Read it at start-up** rather than hard-coding a number, because the
number your process was tuned against last month can lapse by next month.

**Bursting is real, but not everywhere.** Each budget is a token bucket that refills continuously —
"There are no fixed windows and no per-second resets." Basic and Advanced read buckets, and write
buckets above Basic, hold three seconds of budget, so a client that sat idle can spend three times
its per-second rate at once. The Basic write bucket holds one second. So the one tier where a
beginner's bot lives is the one tier where a quiet minute banks no more than one second of orders
for the moment the market moves.

**Batching saves round trips, not tokens.** "A batch request costs the same as making each call
individually", and "the whole batch must fit in the bucket at once": a 25-order create needs 250
tokens on arrival or the entire batch is rejected. At Basic, with a 100-token bucket, a batch of
more than ten orders can never succeed, however long you wait. The batch-create reference says the
maximum batch size "scales with your tier's write budget", which is the same fact from the other
side.

**Sharding changes which bucket a write draws from.** A single order sent with an explicit
`exchange_index` of 1 or more bills that shard's own write bucket, and each shard's bucket carries
the full tier budget. Shard 0 bills the unscoped bucket. An auto-routed order — `exchange_index`
of -1, or omitted when a `market_ticker` is given — "is billed to every shard's Write bucket". Batch
creates and cancels always bill the unscoped bucket, whatever index they carry. A bot that leaves
routing to the exchange is therefore spending from every bucket at once.

**When you run out, the response tells you almost nothing.** "429 responses do not currently
include `Retry-After` or `X-RateLimit-*` headers. There is no penalty or cooldown." The bucket keeps
refilling, and at 1,200 tokens a second an empty bucket has the 10 tokens for one order back in
about 8.3 milliseconds. The venue's instruction is exponential backoff, and the absence of headers
means the only accurate model of your balance is the one you keep yourself.

### Polymarket: requests per IP address, delayed rather than refused

Polymarket's [international order book](https://predictionmarkets.tools/tools/polymarket-clob-api) is built on the opposite
assumptions, and its rate-limit page says so in its first sentence: "The limits on this page are
IP-based and enforced using Cloudflare's throttling system. When you exceed the limit for any
endpoint, requests are throttled (delayed/queued) rather than immediately rejected. Limits reset on
sliding time windows."

Three words in that sentence decide how a bot behaves.

**IP-based.** The unit is the address your traffic leaves from, not your account and not your key.
Two processes on one machine share one allowance, and a second key buys nothing. A hosted runner or a shared office connection shares its allowance with whoever else is behind it.

**Per endpoint.** There is no single number. The CLOB API has a general ceiling of 9,000 requests per
10 seconds, and under it a separate allowance for nearly every route: `/book`, `/price` and
`/midpoint` at 1,500 per 10 seconds, their batch forms `/books`, `/prices` and `/midpoints` at 500,
`/prices-history` at 1,000, the tick-size lookup at 200. The trading endpoints carry two limits
each, a burst and a sustained:

| Endpoint | Burst | Sustained |
| --- | --- | --- |
| `POST /order` | 5,000 per 10 s | 120,000 per 10 min |
| `DELETE /order` | 5,000 per 10 s | 120,000 per 10 min |
| `POST /orders` | 2,000 per 10 s | 21,000 per 10 min |
| `DELETE /orders` | 2,000 per 10 s | 15,000 per 10 min |
| `DELETE /cancel-all` | 250 per 10 s | 6,000 per 10 min |
| `DELETE /cancel-market-orders` | 1,500 per 10 s | 21,000 per 10 min |

The sustained column is the one that binds. 120,000 per 10 minutes is 200 single orders a second,
held; the burst allows 500 a second for a few seconds. The batch endpoint takes up to 15 signed
orders per call, and `DELETE /orders` up to 1,000 order ids, a cap set on 15 June 2026. The market
metadata host and the wallet-data host are metered separately again — the Gamma API's `/markets`
at 300 per 10 seconds, the Data API's `/v2/trades` at 300.

**Delayed.** This is the part that surprises people who have only met Kalshi's model. Over the
limit, the request is not answered with an error; it waits. A bot that measures only success or
failure sees every order succeed and never learns it was over the line — what it sees instead is
latency climbing, on the path where latency is the thing it was built to avoid. The error-code
reference does list a `429 Too Many Requests` with the instruction to back off exponentially, so
both behaviours need handling. But the first symptom is usually time, not a status code.

**The published numbers move, and upward.** The changelog records a rate-limit increase on 8 April
2026 and another on 1 June 2026, when the sustained order limits went to 200 a second; a year
earlier, in May 2025, `POST /order` was published at 500 per 10 seconds burst and 3,000 per
10 minutes sustained. A limit hard-coded from a tutorial is more likely to be too cautious than too
bold — which costs speed rather than orders, and is the cheaper mistake.

**A second limiter is visible in the official client before it is on the page.** Polymarket's
[unified Python client](https://predictionmarkets.tools/tools/polymarket-client), at 0.11.0, parses a set of `Poly-RateLimit-*`
headers "sent with order and cancellation responses": the token balance left "in the applicable
rate-limit bucket", when the current wait ends, which tier was applied, and a `warning` flag that
is true "when the limiter runs in warning mode and the request would have been rejected under live
enforcement; monitor it to adjust request patterns before enforcement begins". The class is
documented as "per-signer". The public rate-limit page, as read on 27 September 2026, describes
only the per-IP limits. Treat that as what it is: the vendor's own code describing a per-signer
token budget on the order path that is not yet in the vendor's own table — and a header worth
logging now, so that you know where you stand if it is switched on.

**Above all of this sit the builder tiers.** An application routing orders under its own builder
code sits in one of three tiers — Unverified, Verified or Partner — and the tier table lists API
rate limits as "Standard" for the first two and "Highest" for Partner, with no figure attached. The
tier that has a number is the relayer, which submits the gasless wallet transactions: 100 a day
unverified, 10,000 verified, unlimited for partners, and the relayer's `/submit` endpoint is
separately limited to 25 requests a minute.

### A stream is metered differently from a request

The obvious way out of a read budget is to stop asking and start listening, and on both venues
that is the right move — but the stream has limits of its own, and they are not a request count.

**On Polymarket the old subscription cap is gone.** The changelog for 28 May 2025 records that "the
100 token subscription limit has been removed for the Markets channel. You can now subscribe to as
many token IDs as needed for your use case." Tokens can be added to and removed from an open
market-channel connection without reconnecting. What the connection does require is a heartbeat:
send the text frame `PING` every 10 seconds and the server replies `PONG`. A process that stops
sending it — because its event loop is blocked processing the burst it subscribed to — is a process
that loses the connection at the worst moment.

**On Kalshi the limit is how fast you read.** Every WebSocket connection is authenticated at the
handshake, including the one that carries only public channels. The error table has no
subscription-count error at all. It has code 25, "Subscription buffer overflow": "The subscription's
event buffer overflowed during a message burst. Subscribe to a smaller subset of data, or ensure
that your connection read throughput is optimized." So on Kalshi the stream throttles you by
falling off when your consumer cannot keep up — which happens in exactly the minutes when every
book on the exchange moves at once.

**Neither venue says, in the pages read for this one, whether stream traffic draws on the REST
budget.** Kalshi's token page covers requests; its WebSocket pages do not mention tokens.
Polymarket's limits are per HTTP endpoint and say nothing about sockets. Assume nothing either way,
and measure: count what you send down the socket, and watch whether your REST latency moves when a
subscription grows.

**A hosted layer adds a meter of its own.** A unified service in front of several venues bills its
own budget on top of theirs. [PMXT](https://predictionmarkets.tools/tools/pmxt)'s hosted tiers, as recorded on its card, run
from 60 requests a minute and 5 WebSocket streams free to 1,000 a minute and 100 streams on Pro,
with a WebSocket message costing 0.1 of a credit — so a book that ticks ten times a second spends
the free month's 25,000 credits in about seven hours. The venue's limit is then rarely the one you
hit first.

### What the client libraries assume

Most bots never call a venue directly. They call a library, and the library has already decided
what a limit is. Read from each project's source on 27 September 2026:

**[pykalshi](https://predictionmarkets.tools/tools/pykalshi) 2.0.0** retries by default: `max_retries=3`, on `429`, `500`, `502`,
`503` and `504`, and on timeouts and connection errors, waiting half a second doubling to a cap of
30 seconds, or whatever a `Retry-After` header says — which Kalshi's page says its 429 responses do not
carry. Its optional
`RateLimiter` is off unless you pass one, and when on it counts **requests per second**, default 10,
not tokens: an order and a cancel are the same unit to it, though Kalshi charges five times as much
for the first. It also reads `X-RateLimit-Remaining` and `X-RateLimit-Reset` to correct itself,
headers Kalshi's page says 429 responses do not currently carry. And the retry loop runs for
`POST` as well as `GET`. A create-order request that timed out is sent again, and the
`client_order_id` that would let the exchange recognise the second copy is an optional argument
the library does not fill in for you.

**[polymarket-client](https://predictionmarkets.tools/tools/polymarket-client) 0.11.0**, Polymarket's own unified SDK, goes the
other way. Its error class states that "The SDK does not retry automatically; callers decide how to
react." The one exception is Data API reads, retried on `429` up to twice, only when the server's
suggested wait is five seconds or less. Everything on the order path raises `RateLimitError`
carrying the parsed `Retry-After` and the `Poly-RateLimit-*` state, and hands you a listener,
`on_rate_limit_update`, to watch the budget before it runs out. Nothing is hidden, and nothing is
done for you.

**[PMXT](https://predictionmarkets.tools/tools/pmxt) and [CCXT](https://predictionmarkets.tools/tools/ccxt)** throttle the way CCXT has always throttled crypto
exchanges: one number per venue, the minimum gap between requests, with rate limiting on by
default. PMXT sets 100 milliseconds for Kalshi and 200 for Polymarket, with a bucket capacity of
one, so there is no burst at all. CCXT's prediction classes set the reverse — 200 milliseconds for
Kalshi, 100 for Polymarket — and give every endpoint a cost of 1. Neither knows about Kalshi's
tiers, its read and write split, the cancel discount, or Polymarket's per-endpoint table. The two
libraries disagree by a factor of two on the same venue because neither number is taken from the
venue: CCXT's 5 requests a second to Kalshi is half of a Basic account's 10 orders, and a small
fraction of what a Premier account may send; PMXT's 5 a second to Polymarket is a fortieth of the
published sustained order limit.

**Kalshi's own generated clients** carry no retry policy of their own, as the
[kalshi-python-sync](https://predictionmarkets.tools/tools/kalshi-python-sync) card records. Whatever you want to happen on a `429`
is yours to write.

Put together: the library either leaves limits to you — Polymarket's own client at least reports
the venue's bucket state as it goes — or it enforces a single request rate it chose itself, and
none of the five keeps a budget in the venue's actual unit. That is not a defect in any one
of them; a generic throttle cannot know your tier. It is a reason to know which kind you are
running before you tune anything.

## What it costs

No venue here sells headroom for money. The currencies are volume, paperwork and time.

- **Kalshi headroom is priced in trading volume.** Expert needs 0.075% of the exchange's volume on
  the formula above, Prestige 1.00%, both measured on both sides of your own trades against twice
  last month's exchange total. What that is in dollars moves with the exchange every month, and so
  your tier can move without you trading any differently.
- **Kalshi's first step up is priced in one API order.** Advanced needs one of your last 100 orders
  to have come through the API, and a call to the upgrade endpoint. It is permanent.
- **Polymarket's order-path limits are the same for everyone at the published level.** The builder
  tiers lift the relayer's daily transaction count by application and review, and describe the
  Partner tier's API limits as the "highest", without a number.
- **A throttled Polymarket request costs time, not a rejection.** The price is paid in queue delay
  on the order you most wanted placed, and nothing in the response says so.
- **A duplicated order costs a position.** A library that retries a timed-out `POST` without an
  idempotency key can place the same order twice, and the second fill is charged the same fee as
  the first — [what a trade actually costs](https://predictionmarkets.tools/guides/what-a-trade-actually-costs) is that
  arithmetic.
- **A hosted layer costs its own meter.** PMXT's free tier is 25,000 credits a month; a streaming
  bot spends it in hours rather than weeks.

## What you can do about it

Each of these is an hour or less, and each is cheaper before a funded run than during one.

**Write down the unit before the number.** For every venue the bot touches: is the limit counted in
tokens or requests; per account, per signer or per IP; per endpoint or overall; rejected or
delayed when exceeded. Kalshi is tokens, per account, read and write, rejected. Polymarket's order
book is requests, per IP, per endpoint, delayed. A library setting that does not match that
sentence is a guess.

**On Kalshi, read your tier at start-up and budget in tokens.** Call `GET /account/limits` for the
tier and the grants, `GET /account/endpoint_costs` for anything that does not cost 10, and keep a
local bucket that refills at your budget and holds your tier's capacity — one second at Basic write,
three seconds above it. Price a batch before you send it: at Basic, a batch that needs more than 100
tokens is rejected whole.

**Take the free upgrade.** If you trade on Kalshi through the API at all, the Advanced tier costs one
API-placed order and one call, and triples the write budget.

**Set a client order id on every order, whatever the library does.** It is the only thing that lets
a retried create be recognised as a retry. If your library retries `POST` on a timeout — pykalshi
does — this is not optional.

**On Polymarket, measure latency as a rate-limit signal.** Log the round-trip time of every order
call alongside its status. A delay that climbs while every call still succeeds is the Cloudflare
throttle doing exactly what its page says. Log the `Poly-RateLimit-*` headers too, especially
`warning`, which reports requests that "would have been rejected under live enforcement".

**Budget cancels as their own thing.** On Kalshi a cancel costs a fifth of an order. On Polymarket
`DELETE /order` has its own allowance, `DELETE /orders` takes up to 1,000 ids, and
`DELETE /cancel-all` has the smallest burst on the table at 250 per 10 seconds — a panic button to
press once, not in a loop.

**Move reads to the stream, and keep the stream fed.** Subscribe instead of polling books. Then keep
the heartbeat on its own timer, so that a burst of messages cannot starve it, and treat Kalshi's
buffer-overflow code 25 as a sign to narrow the subscription or read faster, as its own error text
says, not to resubscribe to the same set.

**Set the library's throttle yourself, from the venue's page.** PMXT and CCXT both expose
`rateLimit` and `enableRateLimit`; pykalshi takes a `RateLimiter` and a `max_retries`. The defaults
were chosen for no tier in particular. Replace them with numbers you can point at on the venue's own
page, and write the date you read it next to them — the Polymarket table has moved twice in 2026
already.

Rate limits are one of the things a bot inherits from the venue rather than fixes. The others —
rehearsal, silent API changes, partial fills, a market paused under a resting order — are in
[what running a bot does not solve](https://predictionmarkets.tools/guides/what-a-bot-cannot-fix), and the libraries themselves
are compared in the [trading clients and bots](https://predictionmarkets.tools/categories/trading-clients-and-bots) listing.

## Tools this bears on

- [Kalshi API](https://predictionmarkets.tools/tools/kalshi-api.md) — REST, WebSocket and FIX access to a CFTC-regulated event exchange.
- [Polymarket CLOB API](https://predictionmarkets.tools/tools/polymarket-clob-api.md) — Order books, prices and order placement on Polymarket's matching engine.
- [pykalshi](https://predictionmarkets.tools/tools/pykalshi.md) — Unofficial Kalshi client with what the generated SDK leaves out - streams and retries.
- [polymarket-client](https://predictionmarkets.tools/tools/polymarket-client.md) — Polymarket's own unified Python SDK - sync and async, data through order signing.
- [PMXT](https://predictionmarkets.tools/tools/pmxt.md) — CCXT-shaped client for prediction markets, with a hosted API and a self-hosted mode.

## FAQ

### How many orders a second can I send to Kalshi?

Divide your tier's write budget by the cost of an order, which is 10 tokens. A Basic account refills 100 write tokens a second, so 10 orders; Advanced, granted permanently once one of your last 100 orders came through the API, is 30; Premier is 120. Your current tier is returned by the account limits endpoint, and tiers above Advanced are earned from volume.

### Why do my Polymarket orders get slow instead of failing?

Because Polymarket's published limits are enforced by Cloudflare per IP address, and over the limit a request is delayed or queued rather than refused. A bot that only counts errors will see every order succeed and never learn it was over the line. Log round-trip time next to status, and back off on 429 as the error reference instructs.

### Does a batch order save rate limit on Kalshi?

No. A batch costs the same tokens as the orders sent one by one, and the whole batch has to fit in the bucket when it arrives or it is rejected entirely. At the Basic tier the write bucket holds 100 tokens, so a create batch of more than ten orders cannot succeed however long you wait.

### Does my client library handle rate limits for me?

Partly, and usually in a unit the venue does not use. pykalshi retries and can throttle by requests per second; PMXT and CCXT enforce one fixed gap per venue; Polymarket's own Python client does not retry order calls at all and leaves the decision to you. None of them knows your Kalshi tier or Polymarket's per-endpoint table.

### Will more API keys give me more throughput?

Not on either venue as documented. Kalshi's budgets belong to the account and its tier, and Polymarket's published limits are counted per IP address, so a second key on the same account or the same machine draws on the same allowance.

## Sources

1. [Rate Limits and Tiers](https://docs.kalshi.com/getting_started/rate_limits) — Kalshi, read 2026-09-27
2. [Upgrade Account API Usage Level](https://docs.kalshi.com/api-reference/account/upgrade-account-api-usage-level) — Kalshi, read 2026-09-27
3. [List Non-Default Endpoint Costs, live response](https://external-api.kalshi.com/trade-api/v2/account/endpoint_costs) — Kalshi, read 2026-09-27
4. [Quick Start: WebSockets](https://docs.kalshi.com/getting_started/quick_start_websockets) — Kalshi, read 2026-09-27
5. [Rate Limits](https://docs.polymarket.com/api-reference/rate-limits) — Polymarket, read 2026-09-27
6. [Real-Time Data, market channel](https://docs.polymarket.com/market-data/realtime-data) — Polymarket, read 2026-09-27
7. [Predictions Changelog](https://docs.polymarket.com/changelog/predictions) — Polymarket, read 2026-09-27
8. [Builder Program Tiers](https://docs.polymarket.com/programs/builders/tiers) — Polymarket, read 2026-09-27
9. [Error Codes](https://docs.polymarket.com/resources/error-codes) — Polymarket, read 2026-09-27
10. [polymarket-client 0.11.0, rate_limit.py and _internal/retry.py](https://github.com/Polymarket/py-sdk/blob/main/src/polymarket/rate_limit.py) — Polymarket, read 2026-09-27
11. [pykalshi 2.0.0, _sync/client.py and rate_limiter.py](https://github.com/arshka/pykalshi/blob/main/pykalshi/_sync/client.py) — arshka (pykalshi), read 2026-09-27
12. [PMXT core 2.17.1, BaseExchange.ts](https://github.com/pmxt-dev/pmxt/blob/main/core/src/BaseExchange.ts) — PMXT, read 2026-09-27
13. [CCXT, prediction/kalshi.ts and prediction/polymarket.ts](https://github.com/ccxt/ccxt/blob/master/ts/src/prediction/kalshi.ts) — CCXT, read 2026-09-27

*Last updated 2026-09-27. A reference page, corrected in place — not a dated post.*
