Why your bot gets throttled, and what the limit is counted against

Kalshi meters tokens per account tier; Polymarket's order book meters requests per IP and queues the excess. What each counts, and what the libraries assume.

Because each venue counts something different, and your client library usually counts something else again. Kalshi gives every account two token buckets, read and write, sized by a tier earned from trading volume; an order costs 10 tokens and a cancel 2. Polymarket's international order book limits requests per IP address per endpoint and delays the excess rather than refusing it. The Python and TypeScript clients here throttle by a single request rate, or not at all.

A process that ran cleanly for a week starts getting 429 back, or starts taking two seconds to place an order that used to take forty milliseconds. Nothing in the code changed. What changed is that the process crossed a line it did not know it was near, and it did not know because the line is not drawn in requests per second. It is drawn in tokens, or per IP address, or per endpoint, or per tier that moved — and the library underneath the process may be counting in a unit the venue does not use at all.

This page covers the two venues most bots in this catalogue are written against, Kalshi and Polymarket's international order book, and the client libraries that sit between your code and them. Polymarket US, Adjacent, Manifold and Limitless publish limits of their own, counted four more ways; those are read in what running a bot does not solve, together with the retry loop that backs off when it should not. Everything below was read from the venues' documentation and the libraries' source on 27 September 2026. None of it has been run against a funded account by this site.

How it works

Kalshi: two buckets per account, and a tier you earn

The Kalshi API's rate-limit page does not count requests. It counts tokens: "Every authenticated request costs tokens. Your tier sets your budget: the rate, in tokens per second, at which your balance refills." Most requests cost 10. A cancel costs 2 — the page's own example prices a 25-order batch cancel at 50 tokens against 250 for a 25-order batch create. The authoritative list of anything that is not 10 is an endpoint, GET /account/endpoint_costs, not a table on the page, so it can change without the page changing.

There are two buckets, and the split is by what the request does, not by how it arrives. Read covers GET endpoints and anything not routed elsewhere. Write covers order placement, amends, cancels, order groups, the request-for-quote flow and block-trade accepts. "REST and FIX requests drain the same buckets", so moving the order path to FIX buys no headroom by itself. Perpetual futures run in separate buckets that event-contract traffic never touches.

The buckets are small at the bottom and large at the top. Per-second budgets, read and write:

TierReadWriteHow you get it
Basic200100Complete account signup
Advanced300300One call to an upgrade endpoint
Expert6006000.075% volume share to earn, 0.05% to keep
Premier1,2001,2000.125% to earn, 0.10% to keep
Paragon2,4002,4000.25% to earn, 0.20% to keep
Prime4,8004,8000.50% to earn, 0.40% to keep
Prestige12,0009,6001.00% to earn, 0.80% to keep

Divide by the cost. At Basic, 100 write tokens a second is 10 orders a second or 50 cancels; 200 read tokens is 20 reads. The step to Advanced is almost free and most people never take it: the upgrade endpoint grants a permanent Advanced tier if "at least 1 of the user's last 100 Predictions orders was created via API", which triples the write budget to 30 orders a second.

Everything above Advanced is earned from volume, and the arithmetic is worth reading slowly. Your share is your trailing 30-day volume, "counting both sides of every trade you are part of, as maker and as taker", divided by twice the previous calendar month's exchange volume. Kalshi reviews it once a day and grants the tier for 30 days; falling below the keep threshold does not drop you at once, it lets the current grant run out. Your present tier and every active grant are readable from GET /account/limits. Read it at start-up rather than hard-coding a number, because the number your process was tuned against last month can lapse by next month.

Bursting is real, but not everywhere. Each budget is a token bucket that refills continuously — "There are no fixed windows and no per-second resets." Basic and Advanced read buckets, and write buckets above Basic, hold three seconds of budget, so a client that sat idle can spend three times its per-second rate at once. The Basic write bucket holds one second. So the one tier where a beginner's bot lives is the one tier where a quiet minute banks no more than one second of orders for the moment the market moves.

Batching saves round trips, not tokens. "A batch request costs the same as making each call individually", and "the whole batch must fit in the bucket at once": a 25-order create needs 250 tokens on arrival or the entire batch is rejected. At Basic, with a 100-token bucket, a batch of more than ten orders can never succeed, however long you wait. The batch-create reference says the maximum batch size "scales with your tier's write budget", which is the same fact from the other side.

Sharding changes which bucket a write draws from. A single order sent with an explicit exchange_index of 1 or more bills that shard's own write bucket, and each shard's bucket carries the full tier budget. Shard 0 bills the unscoped bucket. An auto-routed order — exchange_index of -1, or omitted when a market_ticker is given — "is billed to every shard's Write bucket". Batch creates and cancels always bill the unscoped bucket, whatever index they carry. A bot that leaves routing to the exchange is therefore spending from every bucket at once.

When you run out, the response tells you almost nothing. "429 responses do not currently include Retry-After or X-RateLimit-* headers. There is no penalty or cooldown." The bucket keeps refilling, and at 1,200 tokens a second an empty bucket has the 10 tokens for one order back in about 8.3 milliseconds. The venue's instruction is exponential backoff, and the absence of headers means the only accurate model of your balance is the one you keep yourself.

Polymarket: requests per IP address, delayed rather than refused

Polymarket's international order book is built on the opposite assumptions, and its rate-limit page says so in its first sentence: "The limits on this page are IP-based and enforced using Cloudflare's throttling system. When you exceed the limit for any endpoint, requests are throttled (delayed/queued) rather than immediately rejected. Limits reset on sliding time windows."

Three words in that sentence decide how a bot behaves.

IP-based. The unit is the address your traffic leaves from, not your account and not your key. Two processes on one machine share one allowance, and a second key buys nothing. A hosted runner or a shared office connection shares its allowance with whoever else is behind it.

Per endpoint. There is no single number. The CLOB API has a general ceiling of 9,000 requests per 10 seconds, and under it a separate allowance for nearly every route: /book, /price and /midpoint at 1,500 per 10 seconds, their batch forms /books, /prices and /midpoints at 500, /prices-history at 1,000, the tick-size lookup at 200. The trading endpoints carry two limits each, a burst and a sustained:

EndpointBurstSustained
POST /order5,000 per 10 s120,000 per 10 min
DELETE /order5,000 per 10 s120,000 per 10 min
POST /orders2,000 per 10 s21,000 per 10 min
DELETE /orders2,000 per 10 s15,000 per 10 min
DELETE /cancel-all250 per 10 s6,000 per 10 min
DELETE /cancel-market-orders1,500 per 10 s21,000 per 10 min

The sustained column is the one that binds. 120,000 per 10 minutes is 200 single orders a second, held; the burst allows 500 a second for a few seconds. The batch endpoint takes up to 15 signed orders per call, and DELETE /orders up to 1,000 order ids, a cap set on 15 June 2026. The market metadata host and the wallet-data host are metered separately again — the Gamma API's /markets at 300 per 10 seconds, the Data API's /v2/trades at 300.

Delayed. This is the part that surprises people who have only met Kalshi's model. Over the limit, the request is not answered with an error; it waits. A bot that measures only success or failure sees every order succeed and never learns it was over the line — what it sees instead is latency climbing, on the path where latency is the thing it was built to avoid. The error-code reference does list a 429 Too Many Requests with the instruction to back off exponentially, so both behaviours need handling. But the first symptom is usually time, not a status code.

The published numbers move, and upward. The changelog records a rate-limit increase on 8 April 2026 and another on 1 June 2026, when the sustained order limits went to 200 a second; a year earlier, in May 2025, POST /order was published at 500 per 10 seconds burst and 3,000 per 10 minutes sustained. A limit hard-coded from a tutorial is more likely to be too cautious than too bold — which costs speed rather than orders, and is the cheaper mistake.

A second limiter is visible in the official client before it is on the page. Polymarket's unified Python client, at 0.11.0, parses a set of Poly-RateLimit-* headers "sent with order and cancellation responses": the token balance left "in the applicable rate-limit bucket", when the current wait ends, which tier was applied, and a warning flag that is true "when the limiter runs in warning mode and the request would have been rejected under live enforcement; monitor it to adjust request patterns before enforcement begins". The class is documented as "per-signer". The public rate-limit page, as read on 27 September 2026, describes only the per-IP limits. Treat that as what it is: the vendor's own code describing a per-signer token budget on the order path that is not yet in the vendor's own table — and a header worth logging now, so that you know where you stand if it is switched on.

Above all of this sit the builder tiers. An application routing orders under its own builder code sits in one of three tiers — Unverified, Verified or Partner — and the tier table lists API rate limits as "Standard" for the first two and "Highest" for Partner, with no figure attached. The tier that has a number is the relayer, which submits the gasless wallet transactions: 100 a day unverified, 10,000 verified, unlimited for partners, and the relayer's /submit endpoint is separately limited to 25 requests a minute.

A stream is metered differently from a request

The obvious way out of a read budget is to stop asking and start listening, and on both venues that is the right move — but the stream has limits of its own, and they are not a request count.

On Polymarket the old subscription cap is gone. The changelog for 28 May 2025 records that "the 100 token subscription limit has been removed for the Markets channel. You can now subscribe to as many token IDs as needed for your use case." Tokens can be added to and removed from an open market-channel connection without reconnecting. What the connection does require is a heartbeat: send the text frame PING every 10 seconds and the server replies PONG. A process that stops sending it — because its event loop is blocked processing the burst it subscribed to — is a process that loses the connection at the worst moment.

On Kalshi the limit is how fast you read. Every WebSocket connection is authenticated at the handshake, including the one that carries only public channels. The error table has no subscription-count error at all. It has code 25, "Subscription buffer overflow": "The subscription's event buffer overflowed during a message burst. Subscribe to a smaller subset of data, or ensure that your connection read throughput is optimized." So on Kalshi the stream throttles you by falling off when your consumer cannot keep up — which happens in exactly the minutes when every book on the exchange moves at once.

Neither venue says, in the pages read for this one, whether stream traffic draws on the REST budget. Kalshi's token page covers requests; its WebSocket pages do not mention tokens. Polymarket's limits are per HTTP endpoint and say nothing about sockets. Assume nothing either way, and measure: count what you send down the socket, and watch whether your REST latency moves when a subscription grows.

A hosted layer adds a meter of its own. A unified service in front of several venues bills its own budget on top of theirs. PMXT's hosted tiers, as recorded on its card, run from 60 requests a minute and 5 WebSocket streams free to 1,000 a minute and 100 streams on Pro, with a WebSocket message costing 0.1 of a credit — so a book that ticks ten times a second spends the free month's 25,000 credits in about seven hours. The venue's limit is then rarely the one you hit first.

What the client libraries assume

Most bots never call a venue directly. They call a library, and the library has already decided what a limit is. Read from each project's source on 27 September 2026:

pykalshi 2.0.0 retries by default: max_retries=3, on 429, 500, 502, 503 and 504, and on timeouts and connection errors, waiting half a second doubling to a cap of 30 seconds, or whatever a Retry-After header says — which Kalshi's page says its 429 responses do not carry. Its optional RateLimiter is off unless you pass one, and when on it counts requests per second, default 10, not tokens: an order and a cancel are the same unit to it, though Kalshi charges five times as much for the first. It also reads X-RateLimit-Remaining and X-RateLimit-Reset to correct itself, headers Kalshi's page says 429 responses do not currently carry. And the retry loop runs for POST as well as GET. A create-order request that timed out is sent again, and the client_order_id that would let the exchange recognise the second copy is an optional argument the library does not fill in for you.

polymarket-client 0.11.0, Polymarket's own unified SDK, goes the other way. Its error class states that "The SDK does not retry automatically; callers decide how to react." The one exception is Data API reads, retried on 429 up to twice, only when the server's suggested wait is five seconds or less. Everything on the order path raises RateLimitError carrying the parsed Retry-After and the Poly-RateLimit-* state, and hands you a listener, on_rate_limit_update, to watch the budget before it runs out. Nothing is hidden, and nothing is done for you.

PMXT and CCXT throttle the way CCXT has always throttled crypto exchanges: one number per venue, the minimum gap between requests, with rate limiting on by default. PMXT sets 100 milliseconds for Kalshi and 200 for Polymarket, with a bucket capacity of one, so there is no burst at all. CCXT's prediction classes set the reverse — 200 milliseconds for Kalshi, 100 for Polymarket — and give every endpoint a cost of 1. Neither knows about Kalshi's tiers, its read and write split, the cancel discount, or Polymarket's per-endpoint table. The two libraries disagree by a factor of two on the same venue because neither number is taken from the venue: CCXT's 5 requests a second to Kalshi is half of a Basic account's 10 orders, and a small fraction of what a Premier account may send; PMXT's 5 a second to Polymarket is a fortieth of the published sustained order limit.

Kalshi's own generated clients carry no retry policy of their own, as the kalshi-python-sync card records. Whatever you want to happen on a 429 is yours to write.

Put together: the library either leaves limits to you — Polymarket's own client at least reports the venue's bucket state as it goes — or it enforces a single request rate it chose itself, and none of the five keeps a budget in the venue's actual unit. That is not a defect in any one of them; a generic throttle cannot know your tier. It is a reason to know which kind you are running before you tune anything.

What it costs

No venue here sells headroom for money. The currencies are volume, paperwork and time.

  • Kalshi headroom is priced in trading volume. Expert needs 0.075% of the exchange's volume on the formula above, Prestige 1.00%, both measured on both sides of your own trades against twice last month's exchange total. What that is in dollars moves with the exchange every month, and so your tier can move without you trading any differently.
  • Kalshi's first step up is priced in one API order. Advanced needs one of your last 100 orders to have come through the API, and a call to the upgrade endpoint. It is permanent.
  • Polymarket's order-path limits are the same for everyone at the published level. The builder tiers lift the relayer's daily transaction count by application and review, and describe the Partner tier's API limits as the "highest", without a number.
  • A throttled Polymarket request costs time, not a rejection. The price is paid in queue delay on the order you most wanted placed, and nothing in the response says so.
  • A duplicated order costs a position. A library that retries a timed-out POST without an idempotency key can place the same order twice, and the second fill is charged the same fee as the first — what a trade actually costs is that arithmetic.
  • A hosted layer costs its own meter. PMXT's free tier is 25,000 credits a month; a streaming bot spends it in hours rather than weeks.

What you can do about it

Each of these is an hour or less, and each is cheaper before a funded run than during one.

Write down the unit before the number. For every venue the bot touches: is the limit counted in tokens or requests; per account, per signer or per IP; per endpoint or overall; rejected or delayed when exceeded. Kalshi is tokens, per account, read and write, rejected. Polymarket's order book is requests, per IP, per endpoint, delayed. A library setting that does not match that sentence is a guess.

On Kalshi, read your tier at start-up and budget in tokens. Call GET /account/limits for the tier and the grants, GET /account/endpoint_costs for anything that does not cost 10, and keep a local bucket that refills at your budget and holds your tier's capacity — one second at Basic write, three seconds above it. Price a batch before you send it: at Basic, a batch that needs more than 100 tokens is rejected whole.

Take the free upgrade. If you trade on Kalshi through the API at all, the Advanced tier costs one API-placed order and one call, and triples the write budget.

Set a client order id on every order, whatever the library does. It is the only thing that lets a retried create be recognised as a retry. If your library retries POST on a timeout — pykalshi does — this is not optional.

On Polymarket, measure latency as a rate-limit signal. Log the round-trip time of every order call alongside its status. A delay that climbs while every call still succeeds is the Cloudflare throttle doing exactly what its page says. Log the Poly-RateLimit-* headers too, especially warning, which reports requests that "would have been rejected under live enforcement".

Budget cancels as their own thing. On Kalshi a cancel costs a fifth of an order. On Polymarket DELETE /order has its own allowance, DELETE /orders takes up to 1,000 ids, and DELETE /cancel-all has the smallest burst on the table at 250 per 10 seconds — a panic button to press once, not in a loop.

Move reads to the stream, and keep the stream fed. Subscribe instead of polling books. Then keep the heartbeat on its own timer, so that a burst of messages cannot starve it, and treat Kalshi's buffer-overflow code 25 as a sign to narrow the subscription or read faster, as its own error text says, not to resubscribe to the same set.

Set the library's throttle yourself, from the venue's page. PMXT and CCXT both expose rateLimit and enableRateLimit; pykalshi takes a RateLimiter and a max_retries. The defaults were chosen for no tier in particular. Replace them with numbers you can point at on the venue's own page, and write the date you read it next to them — the Polymarket table has moved twice in 2026 already.

Rate limits are one of the things a bot inherits from the venue rather than fixes. The others — rehearsal, silent API changes, partial fills, a market paused under a resting order — are in what running a bot does not solve, and the libraries themselves are compared in the trading clients and bots listing.

Tools this bears on

Cards in the catalogue where what is above changes the decision.

  • Kalshi API

    REST, WebSocket and FIX access to a CFTC-regulated event exchange.

    Free tier onlyFree tier

  • Polymarket CLOB API

    Order books, prices and order placement on Polymarket's matching engine.

    Free tier onlyFree tier

  • pykalshi

    Unofficial Kalshi client with what the generated SDK leaves out - streams and retries.

    FreeFree tierOpen source

  • polymarket-client

    Polymarket's own unified Python SDK - sync and async, data through order signing.

    FreeFree tierOpen source

  • PMXT

    CCXT-shaped client for prediction markets, with a hosted API and a self-hosted mode.

    $29.99/moFree tierOpen source

FAQ

How many orders a second can I send to Kalshi?

Divide your tier's write budget by the cost of an order, which is 10 tokens. A Basic account refills 100 write tokens a second, so 10 orders; Advanced, granted permanently once one of your last 100 orders came through the API, is 30; Premier is 120. Your current tier is returned by the account limits endpoint, and tiers above Advanced are earned from volume.

Why do my Polymarket orders get slow instead of failing?

Because Polymarket's published limits are enforced by Cloudflare per IP address, and over the limit a request is delayed or queued rather than refused. A bot that only counts errors will see every order succeed and never learn it was over the line. Log round-trip time next to status, and back off on 429 as the error reference instructs.

Does a batch order save rate limit on Kalshi?

No. A batch costs the same tokens as the orders sent one by one, and the whole batch has to fit in the bucket when it arrives or it is rejected entirely. At the Basic tier the write bucket holds 100 tokens, so a create batch of more than ten orders cannot succeed however long you wait.

Does my client library handle rate limits for me?

Partly, and usually in a unit the venue does not use. pykalshi retries and can throttle by requests per second; PMXT and CCXT enforce one fixed gap per venue; Polymarket's own Python client does not retry order calls at all and leaves the decision to you. None of them knows your Kalshi tier or Polymarket's per-endpoint table.

Will more API keys give me more throughput?

Not on either venue as documented. Kalshi's budgets belong to the account and its tier, and Polymarket's published limits are counted per IP address, so a second key on the same account or the same machine draws on the same allowance.

Sources

  1. Rate Limits and Tiers — Kalshi, read
  2. Upgrade Account API Usage Level — Kalshi, read
  3. List Non-Default Endpoint Costs, live response — Kalshi, read
  4. Quick Start: WebSockets — Kalshi, read
  5. Rate Limits — Polymarket, read
  6. Real-Time Data, market channel — Polymarket, read
  7. Predictions Changelog — Polymarket, read
  8. Builder Program Tiers — Polymarket, read
  9. Error Codes — Polymarket, read
  10. polymarket-client 0.11.0, rate_limit.py and _internal/retry.py — Polymarket, read
  11. pykalshi 2.0.0, _sync/client.py and rate_limiter.py — arshka (pykalshi), read
  12. PMXT core 2.17.1, BaseExchange.ts — PMXT, read
  13. CCXT, prediction/kalshi.ts and prediction/polymarket.ts — CCXT, read

The catalogue next door

This page is background, not a listing. The products it bears on are in Trading Clients, SDKs & Bots, each filled in against the same schema, with the fields to narrow it yourself.

Last updated . Corrected in place: this is a reference page, not a dated post.