Skip to main content

Rate Limits

Every API-key-authenticated request returns rate-limit headers so your integration can throttle proactively.

Headers

Every API-key-authenticated response on /api carries these:
X-RateLimit-Policy is set by the success-path header middleware, which never runs on a rejection: the limiter’s own handler ends the response. Do not parse the policy string out of a 429 — it is not there. The other four are set on the 429, alongside Retry-After.
Read X-RateLimit-Remaining on each response and slow down as it approaches 0 to avoid 429s entirely. Session/JWT-authenticated requests use separate role-based limits and do not receive these headers.

Defaults

The range is the same for every key type — there is no lower ceiling on PATs. A key created without rateLimitPerMin gets 60 req/min. Configure it when creating or updating keys via the API Keys dashboard or POST /api/api-keys (rateLimitPerMin, an integer from 1 to 6000).

Handling 429 Responses

When rate-limited, the API returns 429 with a Retry-After header (seconds) plus X-RateLimit-Limit, -Window, -Remaining, and -Reset (but not -Policy, see above), and this body:
Prefer honoring Retry-After / X-RateLimit-Reset when present; otherwise fall back to exponential backoff with jitter:
Everything above describes the per-key limiter on /api. The MCP server has a separate limiter with its own budget, bucketing, headers, and error shape — read the next section before wiring backoff into an MCP client.

MCP rate limits

The MCP server is governed by its own limiter, applied to both POST /mcp and the stream-resume endpoint GET /mcp/stream/{streamId}. It is mounted after authentication, so the bucket always keys off the resolved identity rather than the connection.

Budget

The 120/min default is deliberately higher than the 60/min /api default — agents fan out tool calls, and one user turn can be a dozen tools/call requests. Operators can retune it fleet-wide with the MCP_RATE_LIMIT_PER_MIN environment variable; when the caller presents an API key, that key’s own rateLimitPerMin takes precedence over the fleet default.

Bucketing

One 60-second window per identity, in this order:
  1. Per API key — mcp:apikey:<id>
  2. Per user — mcp:user:<id> (OAuth connector tokens)
  3. Per IP — mcp:ip:<ip> (unauthenticated probes, which only ever see the 401 challenge anyway)
Two MCP clients sharing one API key share one budget. Give each agent its own key if you want them isolated.

Headers

The MCP limiter emits Retry-After and the IETF standard RateLimit-* headers (RateLimit-Limit, RateLimit-Remaining, RateLimit-Reset).
It does not emit the legacy X-RateLimit-* trio documented above — those are REST-only. An MCP client that reads X-RateLimit-Remaining will find nothing. Read RateLimit-Remaining (no X- prefix), or honor Retry-After.

The 429 body

HTTP 429, with a JSON-RPC error envelope rather than the REST error shape:
Note id is always null — the limiter rejects before dispatch and does not echo your JSON-RPC request id, so a client that matches responses strictly by id will not find this one. The message restates the limit that was applied, which is the quickest way to tell whether you are on the fleet default or your key’s own budget.