Retries & Exponential Backoff
TL;DR
In distributed systems, network requests fail due to temporary hiccups. Retries ensure your application tries the request again automatically. Exponential Backoff with Jitter ensures that when you retry, you don’t accidentally overload the recovering system by sending thousands of retries at the exact same millisecond.
Mental Model
How It Works
If a downstream API returns a 5xx (Server Error) or a network timeout occurs, it might succeed a second later.
Instead of retrying instantly, you wait a base time (e.g., 1s), then double it for the next try (2s, 4s, 8s). This is Exponential Backoff.
To prevent all your server’s instances from retrying at the exact same synchronized moment (a “Thundering Herd” attack), you add random Jitter (e.g., wait 2.3s instead of exactly 2.0s).
Example
const axios = require('axios');
const axiosRetry = require('axios-retry').default;
const client = axios.create();
// Configure axios to automatically retry on network errors or 5xx status codes
axiosRetry(client, {
retries: 3, // Number of retries
retryDelay: axiosRetry.exponentialDelay, // Uses exponential backoff (100ms, 200ms, 400ms)
retryCondition: (error) => {
// Only retry if it's a network error or a 5xx response
return axiosRetry.isNetworkOrIdempotentRequestError(error) || error.response?.status >= 500;
}
});
async function fetchExternalData() {
// If the server returns a 502, it will automatically wait and retry 3 times!
const response = await client.get('https://api.example.com/data');
return response.data;
}
Common Interview Questions
Should you retry HTTP POST requests?
Only if they are idempotent. A GET or PUT request is idempotent (doing it twice doesn’t change the outcome). A standard POST request (like charging a credit card) is usually NOT idempotent. If the connection drops after the server processed it but before you received the response, an automatic retry will charge the user twice.
What is the “Thundering Herd” problem?
If a popular downstream service goes offline for 5 seconds, thousands of requests will queue up. If they all use a static 1-second retry delay, all thousands of requests will hit the recovering server at the exact same millisecond, instantly crashing it again. Jitter solves this by randomizing the retry times slightly.