Thirty Seconds Is Not a Reasonable Timeout
The most common timeout value I see in production codebases is 30 seconds. It's almost never the right choice. Here's why most timeout configs silently make outages worse, and what to set instead.
I have a consulting habit that annoys people. When I'm onboarding onto a new codebase, I grep for timeout. Not the test suite timeouts — the HTTP client timeouts, the database query timeouts, the connection pool timeouts. The numbers I find there tell me more about the system's resilience than any architecture diagram.
The most common value? 30 seconds. Sometimes 30000 (milliseconds, same thing). Occasionally 60. Often, no timeout at all — which means the library default, which is frequently infinite.
Here's the thing about 30 seconds: it's not a decision. It's the absence of one. Someone needed a number, typed 30, and moved on. That was three years ago and nobody has thought about it since.
Why 30 seconds destroys your system
Picture a typical web request. A user hits your API. Your API calls a downstream service to fetch some data. That downstream service is having a bad day — not down, just slow. Instead of responding in 150ms, it's taking 28 seconds.
Your timeout is 30 seconds, so the request doesn't fail. It succeeds. Slowly.
Meanwhile, every one of those requests is holding a thread (or a connection, or a goroutine, depending on your stack). Your thread pool has 200 slots. At normal traffic of 50 requests per second, those threads cycle in and out quickly — 150ms each, plenty of headroom. But now each thread is occupied for 28 seconds. Within four seconds, your entire thread pool is saturated. Request 201 starts queuing. Your API's response time goes from 150ms to "however long it takes for a slot to free up," which is also 28 seconds, plus queue time.
Your service isn't down. It's worse than down — it's slow. Your load balancer's health check still passes because your health endpoint doesn't call the downstream service. Your monitoring doesn't page because there are no 5xx errors. Users just see a spinner. For half a minute. Per click.
This is what I mean when I say 30 seconds is not a reasonable timeout. The timeout was supposed to be your safety net. Instead, it gave the failure exactly enough room to cascade.
The math you should actually do
A good timeout isn't a round number you feel comfortable with. It's derived from two things: the normal response time of the dependency, and how long your callers can wait.
Start from the caller side. If your API has an SLO of 500ms at p99, and you call three downstream services sequentially, none of them can take 30 seconds. Even if the calls are parallel, 30 seconds is absurd when your users expect sub-second responses.
Here's my rule of thumb: set the timeout to roughly 3-5x the p99 latency of the dependency under normal conditions. If your database queries normally complete in 50ms at p99, a 250ms timeout is reasonable. If a downstream API responds in 200ms at p99, try 800ms to 1 second.
# Don't do this
http_client:
timeout: 30s
# Do this — and document why
http_client:
connect_timeout: 500ms # TCP handshake; if it takes longer, the host is unreachable
read_timeout: 1200ms # p99 is ~300ms; 4x gives room for GC pauses and slow queries
write_timeout: 500ms # request bodies are small; slow writes indicate network troubleSplitting connect, read, and write timeouts matters. A connect timeout of 5 seconds means you'll wait 5 seconds to discover that a host is unreachable — information you could have in 500ms. A read timeout of 30 seconds means you'll hold a connection open for half a minute waiting for bytes that aren't coming.
Tip
"But what about slow legitimate requests?"
This is the objection I always get. "Some of our queries genuinely take 10 seconds." OK. Why?
Nine times out of ten, the slow request is a report generation, a bulk export, or a search query with no index. These should not share a timeout config with your regular traffic. The fix isn't to raise the timeout for everything — it's to route slow work differently.
Move the heavy operation to a background job. Give it its own connection pool with its own timeout. Put it behind a queue. Return a 202 and let the client poll for results. Whatever you do, don't make every API call wait 30 seconds because one endpoint needs it.
# Regular API calls — tight timeout
api_client = HttpClient(timeout=1.2)
# Report generation — separate client, separate pool, separate timeout
report_client = HttpClient(
timeout=45.0,
max_connections=5, # bounded so it can't starve the main pool
)The five connections matter as much as the 45 seconds. If your report path shares the main pool and each request takes 45 seconds, a handful of report requests will eat all your connections.
The timeout you forgot about
I see teams carefully tune their HTTP timeouts and completely forget about database statement timeouts. Your PostgreSQL statement_timeout defaults to 0, which means "wait forever." A single unoptimized query — maybe a new feature that accidentally scans a full table — can hold a connection indefinitely while the pool drains.
-- Set at the connection level for your application's pool
SET statement_timeout = '5s';Same principle: know your normal query times, set a limit that gives headroom but doesn't let runaway queries silently consume resources. Five seconds is generous for most OLTP workloads. If you need longer for batch operations, use a separate connection with a separate timeout.
Retry and timeout: the dangerous combination
There's a special circle of timeout hell reserved for configs that combine a 30-second timeout with 3 retries. That's a single user request potentially tying up resources for two full minutes before it fails. If the downstream service is slow, you've just tripled the pressure on it while also keeping your own threads busy for 90 seconds.
If your timeout fires, think carefully about whether a retry makes sense. Timeouts usually mean the dependency is overwhelmed. Retrying immediately adds load to an already struggling service. If you do retry, use a much shorter timeout on the retry — and set a budget. "I'll spend at most 2 seconds total on this call, retries included" is more useful than "I'll try 3 times with a 30-second timeout each."
How I audit timeout configs
When I start with a new client, I build a rough map of every network boundary — service-to-service, service-to-database, service-to-cache, service-to-third-party — and collect the timeout value at each one. Then I ask three questions:
- Is there a timeout at all? No timeout means infinite wait, which means a slow dependency can freeze your service permanently.
- Is the timeout derived from something real? If nobody can explain why it's 30 seconds, it's wrong.
- Is the total timeout budget sensible? If Service A calls B calls C, and each has a 30-second timeout, a failure in C can cascade 30 seconds of latency all the way to the user — compounded if retries are involved.
The answers are rarely encouraging. Most systems I look at have at least one path where a slow dependency can silently saturate a connection pool, and the team doesn't know it until it happens in production at the worst possible time.
What's the most creative timeout-related failure you've run into? I keep a collection at this point, and the patterns are always the same — a number someone picked years ago, never measured against reality, quietly waiting to make a bad day worse.