A Cloud Cost Audit That Found More Than Anyone Expected

CLIENT

A large US data center and colocation infrastructure company managing facilities across multiple regions. The company provided colocation, interconnection, and cloud on-ramp services to enterprise customers. Internally, it ran a significant cloud footprint across AWS and Azure supporting operational platforms, customer-facing tools, monitoring systems, and internal business applications. The cloud environment had grown considerably over a three-year period of geographic expansion. 

CHALLENGE

The CFO’s office had raised a concern that was becoming harder to ignore: cloud costs had grown by over 90% in three years while the number of cloud-dependent applications had grown by roughly 60%. The engineering team could offer only a general explanation – more services, more data, more regions – rather than a specific one. 

Some local optimization work had happened, but no one had taken an organization-wide view of where the spend was going or whether the growth was proportionate to actual usage. The infrastructure team was already stretched running a multi-region colocation operation. A cloud cost program wasn’t something they could take on alongside existing commitments. 

SOLUTION

The assessment started with read access to both cloud environments and conversations with the engineering team leads responsible for the main operational platforms. We were looking for structural inefficiencies: architectural decisions that made sense during rapid expansion but were generating recurring cost without proportionate value. 

The assessment took four weeks. 

  • In AWS, the largest finding was compute. EC2 instances supporting operational monitoring and the customer portal were over-provisioned, sized for anticipated peak load that had never materialized. Right-sizing those instances and applying Reserved Instance coverage to stable baseline workloads would reduce compute spend materially. 
  • The second AWS finding was data transfer. Several services were moving data across regions unnecessarily, a result of how the architecture was configured during a fast expansion period. Some cross-region traffic was intentional for redundancy. A larger share was not. 
  • In Azure, the primary finding was storage. A large portion of operational logs and monitoring data sat in hot storage tiers despite not being accessed in months. Lifecycle policies had been discussed by the team but never implemented. 
  • Azure compute had a separate issue: VM scale sets configured with minimum instance counts above the actual workload floor, a conservative decision at initial setup that was never reviewed. 

Across both environments, we found seven cloud services provisioned for projects that had since concluded and never decommissioned. 

We delivered findings in three tiers: immediate actions with no operational risk, changes requiring testing, and architectural changes requiring more planning. The immediate tier represented roughly $390K in annual savings. The full program totaled approximately $640K.

RESULT

The team implemented the first tier in three weeks. Monthly cloud spend dropped by around 24% in the first billing cycle. 

The second tier, including right-sizing and reserved instance commitments, completed over the following five weeks. By end of month two, monthly spend was down approximately 41% from the pre-assessment baseline across both environments. 

The seven decommissioned-but-still-running services had dependencies the internal team wasn’t aware of. One was still receiving periodic writes from a monitoring system that assumed it was active. Removing it cleanly required tracing and updating that dependency, which added a week but prevented a silent data loss issue. 

The tagging work done during the assessment exposed a gap in cloud governance: a significant share of resources had no owner tag and no team attribution. The client established a resource tagging policy as a direct result, making cost attribution by team and project possible for the first time. 

The CFO used the savings projection in the board’s annual infrastructure review. The board approved a follow-on cloud governance program with the tagging initiative as its foundation.

TECHNOLOGY STACK 

  • AWS Cost Explorer + Cost and Usage Reports – baseline spend analysis across all AWS accounts and regions
  • AWS Compute Optimizer – right-sizing recommendations for EC2, RDS, and Lambda workloads
  • AWS Reserved Instances + Savings Plans – commitment-based pricing applied to stable baseline workloads
  • Azure Cost Management + Advisor – spend analysis and optimization recommendations across the Azure environment
  • Azure Blob Storage lifecycle policies – automated tiering for operational logs and archival data 
  • Terraform – all infrastructure changes version-controlled and peer-reviewed before deployment
  • CloudWatch + Azure Monitor – post-optimization monitoring to validate performance under right-sized configurations

 

If your cloud costs have grown faster than your application footprint over the last two years, the gap is usually structural. A targeted assessment across both environments gives you a clear picture before committing to a remediation plan.

• Cloud Computing, IT services

Share:

Similar Case study

Scroll to Top

Upload your CV, and we will get in touch with you soon.