Cloud Cost Optimisation — Datadog and AWS Spend Reduction

Project Links

Tech Stack

DatadogAWSAPMLog ManagementCost AnalysisSavings Plans
Cost Optimisation

Cloud Cost Optimisation — Datadog and AWS Spend Reduction

Systematic cost reduction across Datadog and AWS — cutting observability spend by 15% ($800+/month) through log retention tuning, APM sampling, orphaned resource cleanup, and Savings Plan commitment. All changes made without reducing production visibility.

The Problem

Our Datadog bill had crept up over 18 months without anyone reviewing it. We were paying for things we no longer used, storing logs we never searched, and sampling 100% of traces for services that handled thousands of requests per minute.

Financial cost analysis and savings chart Datadog usage chart showing log volume reduction

What We Did

Log Management

  • Identified the top 3 services by log volume — they accounted for 60% of total logs
  • Configured log pipelines to drop DEBUG level logs entirely in production
  • Moved access logs and debug logs to S3 (rehydrate on demand when needed)
  • Changed retention from 15 days hot to 7 days hot + 90 days archive

APM Sampling

  • Reduced default trace sampling from 100% to 10% for successful requests
  • Kept 100% sampling for errors and requests slower than 500ms
  • Result: same visibility into problems, 40% lower APM cost

Infrastructure Cleanup

  • Identified 12 decommissioned hosts still reporting to Datadog
  • Removed agents from terminated EC2 instances
  • Cleaned up unused dashboards and their associated custom metrics
Analytics dashboard showing metric and cost trends

Results

  • Datadog monthly spend: reduced from ~$5,400 to ~$4,600
  • AWS compute: additional $200/month saved through Reserved Instance purchases for stable workloads
  • No alerts removed, no dashboards degraded, no visibility gaps introduced

The full breakdown of each change and the reasoning behind it is in How I Cut Our Cloud Bill by $800 a Month.