Handling 429s: Backoff and Queue Design for Sends
A 429 from your email API isn't an error to swallow — it's a signal to slow down and retry safely. How to build backoff, jitter, idempotency and a queue that survives rate limits.
The first time a signup spike hits, a lot of teams discover their email code has no answer for 429 Too Many Requests. The send throws, the job dies, and the welcome email — or worse, the password reset — never goes out. The fix isn't a bigger rate limit. It's treating 429 as what it is: a normal, expected part of talking to any rate-limited API, and building the small amount of machinery that turns it into a brief delay instead of a lost message.
A 429 is not an error to log and forget. It's the API asking you to slow down — and a request that's still worth completing.
Know which errors to retry — and which to fix#
Every response comes back in the same envelope — { success, code, message, data, errors } — and the code tells you what to do. The single most important distinction in retry logic is retryable vs terminal.
| Code | Meaning | Retry? |
|---|---|---|
| 202 | Accepted | Done — you're finished |
| 400 | Malformed request | No — fix the payload |
| 401 | Bad/missing API key | No — fix auth |
| 403 | Key lacks the scope (e.g. not send) | No — fix the key's scope |
| 409 | Idempotency conflict — a request with this key is in flight/complete | No — treat as already-sent |
| 422 | Valid JSON, invalid values (bad address, missing field) | No — fix the data |
| 429 | Rate limited | Yes — back off and retry |
| 5xx | Transient server-side issue | Yes — back off and retry |
Retrying a 422 in a loop just burns quota and delays the send that would actually work. Retrying a 429 is exactly right. Baking this table into a single is_retryable(code) function is the cleanest way to keep the logic honest.
Back off exponentially — with jitter#
Naive retries make things worse. If a thousand jobs all hit a 429 and all retry after exactly one second, they collide again at second one — a "thundering herd." The fix is exponential backoff with jitter: increase the delay each attempt, and randomise it so retries spread out instead of synchronising.
- Honour
Retry-Afterfirst. If the response includes it, wait at least that long. The server is telling you when it'll be ready. - Otherwise back off exponentially: roughly base × 2^attempt, capped at a ceiling.
- Add jitter: randomise within the window so clients desynchronise.
- Cap attempts. After a handful, stop and dead-letter the job for inspection rather than retrying forever.
Retry safely: the idempotency key is non-negotiable#
Here's the trap. You send a request, the network times out, and you never got the response — but the send may have succeeded on the server. Retry it naively and you risk two password-reset codes, two receipts, two of everything. The Idempotency-Key header closes this: reuse the same key on the retry, and the server recognises the duplicate and returns the original result instead of sending again. (A 409 means a request with that key is already in flight or complete — treat it as "already sent," not as a failure.) We go deeper on this in idempotent sends and safe retries.
The rule: one logical email = one idempotency key, reused across every retry of that email. Generate it from something stable — the event that triggered the send — not from a fresh random value per attempt.
Put a queue in front of sends#
Backoff handles the individual request. A queue handles the shape of your traffic. Instead of firing sends inline in the web request, enqueue a job and let workers drain it at a controlled rate. This gives you three things you can't get from retry logic alone:
- Backpressure. A signup spike fills the queue; workers drain it steadily. The API sees a smooth rate, not a wall.
- Durability. If a worker crashes mid-retry, the job is still on the queue. Nothing is lost to a process restart.
- A place for dead letters. Requests that exhaust retries land in a dead-letter queue you can inspect and replay, rather than vanishing into a log line.
Hands-on: a worker that does it right#
The pattern below retries 429/5xx with jittered exponential backoff, honours Retry-After, reuses one idempotency key across attempts, and refuses to retry terminal errors. Sandbox key (nv_test_…) shown.
// Node — a single send with correct retry behaviour
const RETRYABLE = new Set([429, 500, 502, 503, 504]);
async function sendWithRetry(payload, idempotencyKey, { maxAttempts = 6 } = {}) {
let attempt = 0;
while (true) {
const res = await fetch("https://api.notifiva.com/v1/send", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.NOTIFIVA_KEY}`, // nv_test_… while building
"Idempotency-Key": idempotencyKey, // SAME key every attempt
"Content-Type": "application/json",
},
body: JSON.stringify(payload),
});
if (res.status === 202) return (await res.json()).data; // messageId is here
if (res.status === 409) return { deduped: true }; // already sent — not a failure
const body = await res.json().catch(() => ({}));
if (!RETRYABLE.has(res.status) || attempt >= maxAttempts - 1) {
throw new Error(`Terminal ${res.status}: ${body.message || "send failed"}`);
}
// Honour Retry-After, else exponential backoff with full jitter
const retryAfter = Number(res.headers.get("Retry-After"));
const backoff = Math.min(1000 * 2 ** attempt, 30000);
const waitMs = retryAfter ? retryAfter * 1000 : Math.random() * backoff;
await new Promise(r => setTimeout(r, waitMs));
attempt++;
}
}Note what this doesn't do: it never retries a 400, 401, 403, or 422, and it never mints a new idempotency key on retry. Those two disciplines are what separate a safe retry loop from a duplicate-email generator.
Example scenario: the launch-day duplicate storm#
(Illustrative scenario, not a customer case.) A product launch drives a signup surge. The send code has a retry loop — but it generates a new idempotency key each attempt and retries on any non-2xx. Under load, the API returns 429; timeouts also fire. The loop retries with fresh keys, so every "maybe it sent" becomes a definite second send. Users get three welcome emails and two of them mark it as spam, which dents sender reputation right when volume is highest. The bug wasn't the rate limit; it was retrying unsafely. One stable idempotency key per email, plus a retryable-code check, would have made the same surge a non-event.
Editorial disclosure: Drafted with AI assistance; reviewed, fact-checked and edited by Alkım Kaplaner. Next scheduled review: .
Frequently asked questions
What does a 429 from an email API mean?
It means you've exceeded the allowed request rate temporarily. It's not a permanent failure — the correct response is to wait (honouring any Retry-After header) and retry the same request. Persistent 429s at low volume suggest a too-tight client loop or a burst you should queue.
How should I back off after a 429?
Honour Retry-After if present; otherwise use exponential backoff (roughly doubling the delay per attempt) with jitter to avoid a thundering herd, capped at a ceiling and a maximum number of attempts. After the cap, dead-letter the job rather than retrying forever.
Why do I need an idempotency key when retrying?
Because a network timeout can hide a send that actually succeeded. Reusing the same idempotency key across retries lets the server recognise the duplicate and return the original result instead of sending a second email. Without it, retries create duplicate resets, receipts and codes.
Which HTTP errors should I never retry?
400, 401, 403, and 422 — these are client-side problems (bad payload, bad auth, wrong scope, invalid data) that won't fix themselves. Retrying them wastes quota and delays real sends. Retry 429 and 5xx; treat 409 as "already sent."
Do I really need a queue for transactional email?
For anything with bursts — signups, launches, batch notifications — yes. A queue provides backpressure so the API sees a smooth rate, durability so crashes don't lose jobs, and a dead-letter path for requests that exhaust retries. Low, steady volume can start without one and add it as you grow.
Send transactional email you can rely on
Notifiva gives you a REST API and an SMTP relay, SPF/DKIM signing, delivery webhooks and per-message logs — so the mail your users are waiting on actually arrives.