Why your bot gets throttled, and what the limit is counted against
Kalshi meters tokens per account tier; Polymarket's order book meters requests per IP and queues the excess. What each counts, and what the libraries assume.
Because each venue counts something different, and your client library usually counts something else again. Kalshi gives every account two token buckets, read and write, sized by a tier earned from trading volume; an order costs 10 tokens and a cancel 2. Polymarket's international order book limits requests per IP address per endpoint and delays the excess rather than refusing it. The Python and TypeScript clients here throttle by a single request rate, or not at all.
A process that ran cleanly for a week starts getting 429 back, or starts taking two seconds to
place an order that used to take forty milliseconds. Nothing in the code changed. What changed is
that the process crossed a line it did not know it was near, and it did not know because the line
is not drawn in requests per second. It is drawn in tokens, or per IP address, or per endpoint, or
per tier that moved — and the library underneath the process may be counting in a unit the venue
does not use at all.
This page covers the two venues most bots in this catalogue are written against, Kalshi and Polymarket's international order book, and the client libraries that sit between your code and them. Polymarket US, Adjacent, Manifold and Limitless publish limits of their own, counted four more ways; those are read in what running a bot does not solve, together with the retry loop that backs off when it should not. Everything below was read from the venues' documentation and the libraries' source on 27 September 2026. None of it has been run against a funded account by this site.
How it works
Kalshi: two buckets per account, and a tier you earn
The Kalshi API's rate-limit page does not count requests. It counts tokens: "Every
authenticated request costs tokens. Your tier sets your budget: the rate, in tokens per
second, at which your balance refills." Most requests cost 10. A cancel costs 2 — the page's own
example prices a 25-order batch cancel at 50 tokens against 250 for a 25-order batch create. The
authoritative list of anything that is not 10 is an endpoint, GET /account/endpoint_costs, not
a table on the page, so it can change without the page changing.
There are two buckets, and the split is by what the request does, not by how it arrives. Read
covers GET endpoints and anything not routed elsewhere. Write covers order placement, amends,
cancels, order groups, the request-for-quote flow and block-trade accepts. "REST and FIX requests
drain the same buckets", so moving the order path to FIX buys no headroom by itself. Perpetual
futures run in separate buckets that event-contract traffic never touches.
The buckets are small at the bottom and large at the top. Per-second budgets, read and write:
| Tier | Read | Write | How you get it |
|---|---|---|---|
| Basic | 200 | 100 | Complete account signup |
| Advanced | 300 | 300 | One call to an upgrade endpoint |
| Expert | 600 | 600 | 0.075% volume share to earn, 0.05% to keep |
| Premier | 1,200 | 1,200 | 0.125% to earn, 0.10% to keep |
| Paragon | 2,400 | 2,400 | 0.25% to earn, 0.20% to keep |
| Prime | 4,800 | 4,800 | 0.50% to earn, 0.40% to keep |
| Prestige | 12,000 | 9,600 | 1.00% to earn, 0.80% to keep |
Divide by the cost. At Basic, 100 write tokens a second is 10 orders a second or 50 cancels; 200 read tokens is 20 reads. The step to Advanced is almost free and most people never take it: the upgrade endpoint grants a permanent Advanced tier if "at least 1 of the user's last 100 Predictions orders was created via API", which triples the write budget to 30 orders a second.
Everything above Advanced is earned from volume, and the arithmetic is worth reading slowly. Your
share is your trailing 30-day volume, "counting both sides of every trade you are part of, as
maker and as taker", divided by twice the previous calendar month's exchange volume. Kalshi reviews
it once a day and grants the tier for 30 days; falling below the keep threshold does not drop you
at once, it lets the current grant run out. Your present tier and every active grant are readable
from GET /account/limits. Read it at start-up rather than hard-coding a number, because the
number your process was tuned against last month can lapse by next month.
Bursting is real, but not everywhere. Each budget is a token bucket that refills continuously — "There are no fixed windows and no per-second resets." Basic and Advanced read buckets, and write buckets above Basic, hold three seconds of budget, so a client that sat idle can spend three times its per-second rate at once. The Basic write bucket holds one second. So the one tier where a beginner's bot lives is the one tier where a quiet minute banks no more than one second of orders for the moment the market moves.
Batching saves round trips, not tokens. "A batch request costs the same as making each call individually", and "the whole batch must fit in the bucket at once": a 25-order create needs 250 tokens on arrival or the entire batch is rejected. At Basic, with a 100-token bucket, a batch of more than ten orders can never succeed, however long you wait. The batch-create reference says the maximum batch size "scales with your tier's write budget", which is the same fact from the other side.
Sharding changes which bucket a write draws from. A single order sent with an explicit
exchange_index of 1 or more bills that shard's own write bucket, and each shard's bucket carries
the full tier budget. Shard 0 bills the unscoped bucket. An auto-routed order — exchange_index
of -1, or omitted when a market_ticker is given — "is billed to every shard's Write bucket". Batch
creates and cancels always bill the unscoped bucket, whatever index they carry. A bot that leaves
routing to the exchange is therefore spending from every bucket at once.
When you run out, the response tells you almost nothing. "429 responses do not currently
include Retry-After or X-RateLimit-* headers. There is no penalty or cooldown." The bucket keeps
refilling, and at 1,200 tokens a second an empty bucket has the 10 tokens for one order back in
about 8.3 milliseconds. The venue's instruction is exponential backoff, and the absence of headers
means the only accurate model of your balance is the one you keep yourself.
Polymarket: requests per IP address, delayed rather than refused
Polymarket's international order book is built on the opposite assumptions, and its rate-limit page says so in its first sentence: "The limits on this page are IP-based and enforced using Cloudflare's throttling system. When you exceed the limit for any endpoint, requests are throttled (delayed/queued) rather than immediately rejected. Limits reset on sliding time windows."
Three words in that sentence decide how a bot behaves.
IP-based. The unit is the address your traffic leaves from, not your account and not your key. Two processes on one machine share one allowance, and a second key buys nothing. A hosted runner or a shared office connection shares its allowance with whoever else is behind it.
Per endpoint. There is no single number. The CLOB API has a general ceiling of 9,000 requests per
10 seconds, and under it a separate allowance for nearly every route: /book, /price and
/midpoint at 1,500 per 10 seconds, their batch forms /books, /prices and /midpoints at 500,
/prices-history at 1,000, the tick-size lookup at 200. The trading endpoints carry two limits
each, a burst and a sustained:
| Endpoint | Burst | Sustained |
|---|---|---|
POST /order | 5,000 per 10 s | 120,000 per 10 min |
DELETE /order | 5,000 per 10 s | 120,000 per 10 min |
POST /orders | 2,000 per 10 s | 21,000 per 10 min |
DELETE /orders | 2,000 per 10 s | 15,000 per 10 min |
DELETE /cancel-all | 250 per 10 s | 6,000 per 10 min |
DELETE /cancel-market-orders | 1,500 per 10 s | 21,000 per 10 min |
The sustained column is the one that binds. 120,000 per 10 minutes is 200 single orders a second,
held; the burst allows 500 a second for a few seconds. The batch endpoint takes up to 15 signed
orders per call, and DELETE /orders up to 1,000 order ids, a cap set on 15 June 2026. The market
metadata host and the wallet-data host are metered separately again — the Gamma API's /markets
at 300 per 10 seconds, the Data API's /v2/trades at 300.
Delayed. This is the part that surprises people who have only met Kalshi's model. Over the
limit, the request is not answered with an error; it waits. A bot that measures only success or
failure sees every order succeed and never learns it was over the line — what it sees instead is
latency climbing, on the path where latency is the thing it was built to avoid. The error-code
reference does list a 429 Too Many Requests with the instruction to back off exponentially, so
both behaviours need handling. But the first symptom is usually time, not a status code.
The published numbers move, and upward. The changelog records a rate-limit increase on 8 April
2026 and another on 1 June 2026, when the sustained order limits went to 200 a second; a year
earlier, in May 2025, POST /order was published at 500 per 10 seconds burst and 3,000 per
10 minutes sustained. A limit hard-coded from a tutorial is more likely to be too cautious than too
bold — which costs speed rather than orders, and is the cheaper mistake.
A second limiter is visible in the official client before it is on the page. Polymarket's
unified Python client, at 0.11.0, parses a set of Poly-RateLimit-*
headers "sent with order and cancellation responses": the token balance left "in the applicable
rate-limit bucket", when the current wait ends, which tier was applied, and a warning flag that
is true "when the limiter runs in warning mode and the request would have been rejected under live
enforcement; monitor it to adjust request patterns before enforcement begins". The class is
documented as "per-signer". The public rate-limit page, as read on 27 September 2026, describes
only the per-IP limits. Treat that as what it is: the vendor's own code describing a per-signer
token budget on the order path that is not yet in the vendor's own table — and a header worth
logging now, so that you know where you stand if it is switched on.
Above all of this sit the builder tiers. An application routing orders under its own builder
code sits in one of three tiers — Unverified, Verified or Partner — and the tier table lists API
rate limits as "Standard" for the first two and "Highest" for Partner, with no figure attached. The
tier that has a number is the relayer, which submits the gasless wallet transactions: 100 a day
unverified, 10,000 verified, unlimited for partners, and the relayer's /submit endpoint is
separately limited to 25 requests a minute.
A stream is metered differently from a request
The obvious way out of a read budget is to stop asking and start listening, and on both venues that is the right move — but the stream has limits of its own, and they are not a request count.
On Polymarket the old subscription cap is gone. The changelog for 28 May 2025 records that "the
100 token subscription limit has been removed for the Markets channel. You can now subscribe to as
many token IDs as needed for your use case." Tokens can be added to and removed from an open
market-channel connection without reconnecting. What the connection does require is a heartbeat:
send the text frame PING every 10 seconds and the server replies PONG. A process that stops
sending it — because its event loop is blocked processing the burst it subscribed to — is a process
that loses the connection at the worst moment.
On Kalshi the limit is how fast you read. Every WebSocket connection is authenticated at the handshake, including the one that carries only public channels. The error table has no subscription-count error at all. It has code 25, "Subscription buffer overflow": "The subscription's event buffer overflowed during a message burst. Subscribe to a smaller subset of data, or ensure that your connection read throughput is optimized." So on Kalshi the stream throttles you by falling off when your consumer cannot keep up — which happens in exactly the minutes when every book on the exchange moves at once.
Neither venue says, in the pages read for this one, whether stream traffic draws on the REST budget. Kalshi's token page covers requests; its WebSocket pages do not mention tokens. Polymarket's limits are per HTTP endpoint and say nothing about sockets. Assume nothing either way, and measure: count what you send down the socket, and watch whether your REST latency moves when a subscription grows.
A hosted layer adds a meter of its own. A unified service in front of several venues bills its own budget on top of theirs. PMXT's hosted tiers, as recorded on its card, run from 60 requests a minute and 5 WebSocket streams free to 1,000 a minute and 100 streams on Pro, with a WebSocket message costing 0.1 of a credit — so a book that ticks ten times a second spends the free month's 25,000 credits in about seven hours. The venue's limit is then rarely the one you hit first.
What the client libraries assume
Most bots never call a venue directly. They call a library, and the library has already decided what a limit is. Read from each project's source on 27 September 2026:
pykalshi 2.0.0 retries by default: max_retries=3, on 429, 500, 502,
503 and 504, and on timeouts and connection errors, waiting half a second doubling to a cap of
30 seconds, or whatever a Retry-After header says — which Kalshi's page says its 429 responses do not
carry. Its optional
RateLimiter is off unless you pass one, and when on it counts requests per second, default 10,
not tokens: an order and a cancel are the same unit to it, though Kalshi charges five times as much
for the first. It also reads X-RateLimit-Remaining and X-RateLimit-Reset to correct itself,
headers Kalshi's page says 429 responses do not currently carry. And the retry loop runs for
POST as well as GET. A create-order request that timed out is sent again, and the
client_order_id that would let the exchange recognise the second copy is an optional argument
the library does not fill in for you.
polymarket-client 0.11.0, Polymarket's own unified SDK, goes the
other way. Its error class states that "The SDK does not retry automatically; callers decide how to
react." The one exception is Data API reads, retried on 429 up to twice, only when the server's
suggested wait is five seconds or less. Everything on the order path raises RateLimitError
carrying the parsed Retry-After and the Poly-RateLimit-* state, and hands you a listener,
on_rate_limit_update, to watch the budget before it runs out. Nothing is hidden, and nothing is
done for you.
PMXT and CCXT throttle the way CCXT has always throttled crypto exchanges: one number per venue, the minimum gap between requests, with rate limiting on by default. PMXT sets 100 milliseconds for Kalshi and 200 for Polymarket, with a bucket capacity of one, so there is no burst at all. CCXT's prediction classes set the reverse — 200 milliseconds for Kalshi, 100 for Polymarket — and give every endpoint a cost of 1. Neither knows about Kalshi's tiers, its read and write split, the cancel discount, or Polymarket's per-endpoint table. The two libraries disagree by a factor of two on the same venue because neither number is taken from the venue: CCXT's 5 requests a second to Kalshi is half of a Basic account's 10 orders, and a small fraction of what a Premier account may send; PMXT's 5 a second to Polymarket is a fortieth of the published sustained order limit.
Kalshi's own generated clients carry no retry policy of their own, as the
kalshi-python-sync card records. Whatever you want to happen on a 429
is yours to write.
Put together: the library either leaves limits to you — Polymarket's own client at least reports the venue's bucket state as it goes — or it enforces a single request rate it chose itself, and none of the five keeps a budget in the venue's actual unit. That is not a defect in any one of them; a generic throttle cannot know your tier. It is a reason to know which kind you are running before you tune anything.
What it costs
No venue here sells headroom for money. The currencies are volume, paperwork and time.
- Kalshi headroom is priced in trading volume. Expert needs 0.075% of the exchange's volume on the formula above, Prestige 1.00%, both measured on both sides of your own trades against twice last month's exchange total. What that is in dollars moves with the exchange every month, and so your tier can move without you trading any differently.
- Kalshi's first step up is priced in one API order. Advanced needs one of your last 100 orders to have come through the API, and a call to the upgrade endpoint. It is permanent.
- Polymarket's order-path limits are the same for everyone at the published level. The builder tiers lift the relayer's daily transaction count by application and review, and describe the Partner tier's API limits as the "highest", without a number.
- A throttled Polymarket request costs time, not a rejection. The price is paid in queue delay on the order you most wanted placed, and nothing in the response says so.
- A duplicated order costs a position. A library that retries a timed-out
POSTwithout an idempotency key can place the same order twice, and the second fill is charged the same fee as the first — what a trade actually costs is that arithmetic. - A hosted layer costs its own meter. PMXT's free tier is 25,000 credits a month; a streaming bot spends it in hours rather than weeks.
What you can do about it
Each of these is an hour or less, and each is cheaper before a funded run than during one.
Write down the unit before the number. For every venue the bot touches: is the limit counted in tokens or requests; per account, per signer or per IP; per endpoint or overall; rejected or delayed when exceeded. Kalshi is tokens, per account, read and write, rejected. Polymarket's order book is requests, per IP, per endpoint, delayed. A library setting that does not match that sentence is a guess.
On Kalshi, read your tier at start-up and budget in tokens. Call GET /account/limits for the
tier and the grants, GET /account/endpoint_costs for anything that does not cost 10, and keep a
local bucket that refills at your budget and holds your tier's capacity — one second at Basic write,
three seconds above it. Price a batch before you send it: at Basic, a batch that needs more than 100
tokens is rejected whole.
Take the free upgrade. If you trade on Kalshi through the API at all, the Advanced tier costs one API-placed order and one call, and triples the write budget.
Set a client order id on every order, whatever the library does. It is the only thing that lets
a retried create be recognised as a retry. If your library retries POST on a timeout — pykalshi
does — this is not optional.
On Polymarket, measure latency as a rate-limit signal. Log the round-trip time of every order
call alongside its status. A delay that climbs while every call still succeeds is the Cloudflare
throttle doing exactly what its page says. Log the Poly-RateLimit-* headers too, especially
warning, which reports requests that "would have been rejected under live enforcement".
Budget cancels as their own thing. On Kalshi a cancel costs a fifth of an order. On Polymarket
DELETE /order has its own allowance, DELETE /orders takes up to 1,000 ids, and
DELETE /cancel-all has the smallest burst on the table at 250 per 10 seconds — a panic button to
press once, not in a loop.
Move reads to the stream, and keep the stream fed. Subscribe instead of polling books. Then keep the heartbeat on its own timer, so that a burst of messages cannot starve it, and treat Kalshi's buffer-overflow code 25 as a sign to narrow the subscription or read faster, as its own error text says, not to resubscribe to the same set.
Set the library's throttle yourself, from the venue's page. PMXT and CCXT both expose
rateLimit and enableRateLimit; pykalshi takes a RateLimiter and a max_retries. The defaults
were chosen for no tier in particular. Replace them with numbers you can point at on the venue's own
page, and write the date you read it next to them — the Polymarket table has moved twice in 2026
already.
Rate limits are one of the things a bot inherits from the venue rather than fixes. The others — rehearsal, silent API changes, partial fills, a market paused under a resting order — are in what running a bot does not solve, and the libraries themselves are compared in the trading clients and bots listing.
Tools this bears on
Cards in the catalogue where what is above changes the decision.
Kalshi API
REST, WebSocket and FIX access to a CFTC-regulated event exchange.
Free tier onlyFree tier
Polymarket CLOB API
Order books, prices and order placement on Polymarket's matching engine.
Free tier onlyFree tier
pykalshi
Unofficial Kalshi client with what the generated SDK leaves out - streams and retries.
FreeFree tierOpen source
polymarket-client
Polymarket's own unified Python SDK - sync and async, data through order signing.
FreeFree tierOpen source
PMXT
CCXT-shaped client for prediction markets, with a hosted API and a self-hosted mode.
$29.99/moFree tierOpen source
FAQ
How many orders a second can I send to Kalshi?
Divide your tier's write budget by the cost of an order, which is 10 tokens. A Basic account refills 100 write tokens a second, so 10 orders; Advanced, granted permanently once one of your last 100 orders came through the API, is 30; Premier is 120. Your current tier is returned by the account limits endpoint, and tiers above Advanced are earned from volume.
Why do my Polymarket orders get slow instead of failing?
Because Polymarket's published limits are enforced by Cloudflare per IP address, and over the limit a request is delayed or queued rather than refused. A bot that only counts errors will see every order succeed and never learn it was over the line. Log round-trip time next to status, and back off on 429 as the error reference instructs.
Does a batch order save rate limit on Kalshi?
No. A batch costs the same tokens as the orders sent one by one, and the whole batch has to fit in the bucket when it arrives or it is rejected entirely. At the Basic tier the write bucket holds 100 tokens, so a create batch of more than ten orders cannot succeed however long you wait.
Does my client library handle rate limits for me?
Partly, and usually in a unit the venue does not use. pykalshi retries and can throttle by requests per second; PMXT and CCXT enforce one fixed gap per venue; Polymarket's own Python client does not retry order calls at all and leaves the decision to you. None of them knows your Kalshi tier or Polymarket's per-endpoint table.
Will more API keys give me more throughput?
Not on either venue as documented. Kalshi's budgets belong to the account and its tier, and Polymarket's published limits are counted per IP address, so a second key on the same account or the same machine draws on the same allowance.
Sources
- Rate Limits and Tiers — Kalshi, read
- Upgrade Account API Usage Level — Kalshi, read
- List Non-Default Endpoint Costs, live response — Kalshi, read
- Quick Start: WebSockets — Kalshi, read
- Rate Limits — Polymarket, read
- Real-Time Data, market channel — Polymarket, read
- Predictions Changelog — Polymarket, read
- Builder Program Tiers — Polymarket, read
- Error Codes — Polymarket, read
- polymarket-client 0.11.0, rate_limit.py and _internal/retry.py — Polymarket, read
- pykalshi 2.0.0, _sync/client.py and rate_limiter.py — arshka (pykalshi), read
- PMXT core 2.17.1, BaseExchange.ts — PMXT, read
- CCXT, prediction/kalshi.ts and prediction/polymarket.ts — CCXT, read
The catalogue next door
This page is background, not a listing. The products it bears on are in Trading Clients, SDKs & Bots, each filled in against the same schema, with the fields to narrow it yourself.
Last updated . Corrected in place: this is a reference page, not a dated post.