Rate limits
Guide

Rate limits

Requests to https://mcp.jethost.bg/mcp are metered per connection and per account. Normal day-to-day use with an AI assistant stays well within the limits; they exist to keep the service fast and fair when a client loops or runs away.

HTTP 429Retry-AfterRateLimit-* headers

Limits #

BucketAllowance
Per connection100 requests per 60 s — one access token. Starts afresh when the client refreshes its token.
Per account & application200 requests per 60 s — every connection one application holds to your account, combined. Survives token refresh.
  • Every POST to the MCP endpoint counts — tool calls and protocol messages alike (initialize, tools/list, ping, notifications).
  • Windows are fixed: a window opens with the first request and closes 60 s later, when the allowance is restored in full.
  • Signed-in customers see the same numbers in the client area under Usage limits.

Response headers #

Every authenticated response from the MCP endpoint reports the bucket closest to running out, so a client can slow down before it is refused:

RateLimit-LimitAllowance of that bucket per window.
RateLimit-RemainingRequests left in the current window.
RateLimit-ResetSeconds until the window resets.

When you hit a limit #

  • The server answers 429 Too Many Requests with a Retry-After header (seconds), the RateLimit-* headers of the bucket that refused, and a JSON body — see the example.
  • Wait at least Retry-After seconds, then continue. Retrying sooner is refused again and does not shorten the wait.
  • A rate limit is always a 429. A timeout, a dropped connection or a 5xx is not a rate limit — treat it as a transient failure and retry with exponential backoff.