Project Links
Tech Stack
DatadogAWSAPMLog ManagementCost AnalysisSavings Plans
Cost Optimisation
Cloud Cost Optimisation — Datadog and AWS Spend Reduction
Systematic cost reduction across Datadog and AWS — cutting observability spend by 15% ($800+/month) through log retention tuning, APM sampling, orphaned resource cleanup, and Savings Plan commitment. All changes made without reducing production visibility.
The Problem
Our Datadog bill had crept up over 18 months without anyone reviewing it. We were paying for things we no longer used, storing logs we never searched, and sampling 100% of traces for services that handled thousands of requests per minute.
What We Did
Log Management
- Identified the top 3 services by log volume — they accounted for 60% of total logs
- Configured log pipelines to drop DEBUG level logs entirely in production
- Moved access logs and debug logs to S3 (rehydrate on demand when needed)
- Changed retention from 15 days hot to 7 days hot + 90 days archive
APM Sampling
- Reduced default trace sampling from 100% to 10% for successful requests
- Kept 100% sampling for errors and requests slower than 500ms
- Result: same visibility into problems, 40% lower APM cost
Infrastructure Cleanup
- Identified 12 decommissioned hosts still reporting to Datadog
- Removed agents from terminated EC2 instances
- Cleaned up unused dashboards and their associated custom metrics
Results
- Datadog monthly spend: reduced from ~$5,400 to ~$4,600
- AWS compute: additional $200/month saved through Reserved Instance purchases for stable workloads
- No alerts removed, no dashboards degraded, no visibility gaps introduced
The full breakdown of each change and the reasoning behind it is in How I Cut Our Cloud Bill by $800 a Month.
Project Links
Tech Stack
DatadogAWSAPMLog ManagementCost AnalysisSavings Plans