Thrive AI Engineering
← All questions

models · Editorial starter

How should an agent retry rate limits without exceeding its total deadline?

By Thrive AI Editorial · Created 2026-10-07 · Updated 2026-10-07

Thrive AI Editorial · AI-assisted worked example, not a member conversation.

A model request receives a rate-limit response with a retry delay. The application has only a few seconds left, while both the SDK and application may retry. How should retry hints, total deadlines and nested retry loops interact?

This editorial example computes retry decisions without sleeping or calling a provider. It does not assume a particular model's quota or pricing.

0 score

Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.

1 published answer

A selected solution is chosen by the question author or moderator, not an independent certification. Check assumptions before using any code in production.

Thrive AI Editorial

Thrive AI Editorial · AI-assisted worked example, not a member conversation.

2026-10-07

Budget the whole operation, not just each attempt

A ten-second timeout per attempt does not imply a ten-second operation. Several attempts plus backoff can consume much longer. Nested retry loops can multiply requests: three application attempts with three internal attempts each can produce up to nine requests.

Choose one retry owner or account for all layers. Retry only eligible transient failures. A billing or exhausted-quota error is not repaired by repeatedly sending the same request.

When a valid server retry hint exceeds your remaining budget, do not clamp it downward and retry early. Defer or fail with an explicit retry-later outcome. A request should not claim it completed when it was merely queued.

Code / output
retryable? -> attempts left? -> delay + attempt budget fits deadline?
                                   yes -> wait then retry
                                   no  -> defer / report deadline

Test the scheduling decision

Python 3, standard library only. Durations are illustrative seconds, not provider recommendations. Deterministic jitter makes the example reproducible.

Code / output
def next_delay(attempt, remaining, server_delay=None, jitter=0.25):
    if attempt >= 3:
        return None
    fallback = min(8.0, 2.0 ** attempt)
    minimum = fallback if server_delay is None else server_delay
    if minimum < 0 or not 0 <= jitter <= 1:
        raise ValueError("invalid delay")
    delay = minimum + jitter
    attempt_budget = 2.0
    return delay if delay + attempt_budget <= remaining else None

assert next_delay(0, 10) == 1.25
assert next_delay(1, 10, server_delay=4) == 4.25
assert next_delay(1, 5, server_delay=4) is None
assert next_delay(3, 100) is None
print("fallback delay:", next_delay(0, 10))
print("server hint:", next_delay(1, 10, server_delay=4))
print("insufficient deadline:", next_delay(1, 5, server_delay=4))

Expected output:

Code / output
fallback delay: 1.25
server hint: 4.25
insufficient deadline: None

What must be added in production

Use a monotonic clock for elapsed budgets. Parse the provider's supported Retry-After format, including an HTTP date if supported, validate values, and handle invalid hints explicitly. This example accepts a numeric delay already parsed by a trusted caller.

Randomize jitter independently across clients; deterministic jitter is used here only for testing. Apply bounded concurrency and shared admission control when multiple workers consume one project-level quota. Backoff in one worker does not coordinate the whole fleet.

Inspect the retry defaults of the exact SDK version you deploy. Disable internal retries if the application owns the complete policy, or measure and include their attempts and delays. Recheck cancellation after sleeping and before starting another request.

Acceptance criteria

Assert no retry begins before a valid server minimum, no attempt begins without sufficient budget, and terminal errors do not trigger retry loops. Include missing, malformed and excessively long hints. Record attempt counts, terminal category and total elapsed time without logging sensitive prompts.

Our scheduler does not estimate real request duration, enforce distributed quotas or call a model. Treat its two-second attempt reserve as a test assumption, not a safe production constant.

Source checked 2026-10-08: OpenAI rate-limit guidance. Server hints, jitter, bounded attempts/deadlines and SDK retry interactions are documented there. This executable scheduling policy is our own AI-assisted example.

0 score

Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.

Verify your sign-in to contribute an answer

Reading is free. Posting requires a verified account and moderation; no course purchase is required.