Our Monitoring Cost More Than What It Monitored

A client's Datadog bill quietly grew to $38K/month — nearly double their actual compute spend. The culprit wasn't one big mistake but a dozen small defaults nobody revisited.


I was six weeks into an infrastructure cost review for a Series B startup when I found the number that made everyone in the room go quiet. Their Datadog bill had hit $38,400 the previous month. Their entire compute spend on AWS — EC2, ECS, RDS, the lot — was $21,000.

They were paying nearly twice as much to watch their services as to run them.

Nobody had noticed because the observability bill was buried across three cost centers and two billing accounts. Finance saw "software subscriptions." Engineering saw "monitoring." Neither team had the full picture. It took a spreadsheet and a Wednesday afternoon to put the pieces together.

How costs compound when nobody owns them

The company ran about 40 services across ECS, processing somewhere around 2,000 requests per second at peak. Reasonable workload, reasonable team size — sixteen engineers. They'd adopted Datadog two years earlier when they had eight services and five engineers. At the time, the bill was around $4,000 a month and nobody thought twice about it.

The problem with observability pricing is that costs scale with volume, not value. Every new service added hosts, logs, traces, and custom metrics. Nobody was subtracting anything. The bill grew 12-15% per quarter, but so did revenue, so it never triggered a review.

Here's what I found when I dug into the line items.

Debug logging in production, forever

The single biggest cost driver was log ingestion: $14,200 per month. The platform was ingesting about 1.8 TB of logs daily. For 40 services doing 2,000 req/s, that felt high. I pulled a sample and started reading.

Roughly 40% of the log volume was DEBUG-level output. One service alone — a payment reconciliation worker — was producing 380 GB of logs per month. It had been deployed with LOG_LEVEL=debug eighteen months earlier when someone was chasing an intermittent bug. The bug was fixed within a week. The log level was never changed back.

This was the pattern everywhere. Seven of the forty services were running at debug level in production. Not because anyone decided to — because nobody decided not to.

# What I found in most service configs
LOG_LEVEL: debug
 
# What three of them didn't even have
# (defaulting to the library's own default, which was also debug)

The fix was straightforward. I set every service to INFO in production and added a shared config default that teams could override with a documented reason. Log volume dropped 62% the first week.

Tracing at 100% sample rate

The second surprise was APM costs: $11,800 per month. They were tracing every single request. Not 10%, not 1% — every one. The DD_TRACE_SAMPLE_RATE had been set to 1.0 in their base Docker image and never touched.

For debugging production issues, you rarely need more than 5-10% of traces. The math is simple: at 2,000 req/s, a 10% sample rate still gives you 200 traced requests per second. That's plenty to spot latency patterns, find slow endpoints, and trace errors. You're not doing statistical analysis on individual requests — you're looking for patterns.

I set the sample rate to 0.1 and added head-based sampling rules to keep 100% of error traces and anything over 2 seconds. The APM bill dropped to around $1,900 the next month.

Tip

Always keep error traces at 100% sample rate. Drop the baseline to 5-10% and use rules to capture slow or erroring requests in full. You won't miss anything that matters.

Custom metrics nobody queried

The third line item was custom metrics: $7,600 per month. They had 4,200 custom metrics. I asked the team which dashboards used them. Awkward silence.

I exported the metric names and cross-referenced them with every dashboard, monitor, and alert in the account. 2,800 of the 4,200 metrics were not referenced anywhere. Not in a single dashboard. Not in a single alert. They were being emitted by application code, ingested by Datadog, and stored — for no reason.

Some were leftovers from A/B tests that ended a year ago. Some were per-customer counters that made sense when they had 30 customers and became a cardinality explosion at 1,200. One team had instrumented a feature with 47 distinct metrics during a performance investigation six months prior, found the bottleneck, fixed it, and never removed the instrumentation.

I wrote a script that tagged every unreferenced metric and gave teams two weeks to claim anything they still needed. Nobody claimed the vast majority. We removed 2,400 metrics and the custom metrics bill dropped to $2,100.

The retention policy that didn't exist

Log retention was set to 30 days, which is Datadog's default. But nobody had made that decision — it was just what came out of the box. When I asked the team what they actually needed, the answer was interesting. Most debugging happens within the first 48 hours of an incident. Compliance required 90 days of certain audit logs. Everything else was a middle ground nobody had thought about.

We moved to a tiered approach: hot storage for 7 days on the expensive indexes, 90-day retention on rehydratable archives for audit logs, and 15 days for everything else. The savings here were modest — about $2,800 per month — but the principle mattered. Every default you don't question is a cost you've silently approved.

The uncomfortable math

After all the changes, the monthly Datadog bill settled at around $11,200. Still not cheap, but a 71% reduction from $38,400. The compute bill was still $21,000. Monitoring was no longer the bigger number.

Total effort: about three days of investigation, one day of implementation, and two weeks of verifying nothing broke. No service went unmonitored. No alerts stopped firing. The team didn't lose any debugging capability they were actually using.

The irony is that all of these issues were individually small decisions — or non-decisions. Nobody set out to spend $38K on monitoring. It happened because observability tooling scales with entropy, and entropy always increases unless someone is actively managing it.

What I'd tell any team

Your observability costs deserve the same scrutiny as your infrastructure costs. If your monitoring bill is a line item nobody owns, it's growing faster than you think. Audit it like you'd audit your AWS bill: look at what you're ingesting, what you're actually querying, and what you're storing but never touching.

The most expensive telemetry is the data you collect, pay for, and never look at. Which, if my consulting experience is any guide, is probably about half of it.