Polling is the most generous tier on purpose. There are no webhooks, so polling is the only way you learn a job finished, and a completion signal you are charged to use is not much of a signal.
When you hit one
You get429 with a Retry-After header in seconds. Honor it rather than picking your own interval; it is computed from your actual window, so backing off by less just burns another request.
Concurrency is separate
Some launches also refuse with429 concurrency-limit-reached, which is not the same thing as being rate limited. It means too many of your jobs are already running, so the fix is to wait for one to finish rather than to slow your request rate. Slowing down does not help; finishing does.
The cap is per capability, not one shared pool, and it is sized by what that work costs us:
Fourteen caps, measured against
limits.py on 2026-09-06.
Thirteen of them count across your whole organization, so two credentials in the same organization draw on one budget and the count is of rows still unfinished, not of requests you sent.
Matrix runs are the exception.
They count against the Vaquill user your installation attributes its writes to, which is the person who provisioned it, and that is the same counter the product’s own matrix runs use.
So a matrix that person starts in the browser consumes one of your three.
The
Retry-After on a 202 and on a 200 read of an unfinished operation is guidance about the job, not about your budget. Only the one on a 429 is a limit.
