When Your Serverless Bill Becomes Very Serverfull
Three months ago, a junior developer on our team deployed what they thought was a simple Lambda function to resize user profile images. Standard stuff. Except they configured it to run with 3008MB of memory and forgot to set a timeout. The function would occasionally hang on corrupted images, burning through our AWS budget like a cryptocurrency mining operation disguised as image processing.
Our monthly Lambda bill went from $312 to $47,000 before anyone noticed. The function was technically working—users got their resized images eventually—but each hung execution was costing us roughly $23 per image. I’ve seen plenty of expensive bugs in my career, but this one had the distinction of being both catastrophically expensive and hilariously preventable.
That incident became our team’s introduction to the harsh reality that cloud infrastructure cost optimization isn’t just about finding cheaper instances. It’s about understanding the dozen different ways your perfectly functional code can bankrupt your startup.
The Unsexy Truth About Reserved Instances and Savings Plans
Everyone talks about reserved instances like they’re some mystical art form, but the math is actually straightforward once you strip away the AWS marketing speak. Reserved instances work when you can predict your baseline compute needs for the next one to three years. The catch is that most teams grossly overestimate their predictability.
Here’s what actually works: Start with a conservative 30% reservation coverage on your stable workloads. For us, that meant our primary application servers running our API and our always-on database instances. We avoided reserving anything related to batch processing, machine learning training, or development environments. These workloads are inherently spiky and unpredictable.
Savings Plans are AWS’s newer, more flexible alternative, and they’re genuinely better for most use cases. Unlike reserved instances that lock you into specific instance types and availability zones, Savings Plans give you the same discount rates while letting you shift between different services and regions. The 1-year Compute Savings Plan is our go-to approach for anything we’re confident will stick around.
Spot Instances Are Not Just for Batch Jobs Anymore
The conventional wisdom says spot instances are only suitable for fault-tolerant batch processing. That’s increasingly wrong. With the right architecture, you can run production web applications on spot instances and achieve 60-80% cost savings without sacrificing reliability.
The key is mixing spot instances with on-demand instances in your auto-scaling groups. We run 70% of our web tier on spot instances across multiple instance types and availability zones. When AWS reclaims spot capacity, our load balancer shifts traffic to the remaining instances while new spot instances spin up. The occasional brief capacity reduction is invisible to users but saves us about $8,000 monthly.
Auto Scaling Groups with mixed instance policies make this simple to implement. You define multiple instance types that can handle your workload—say, m5.large, m5a.large, and m4.large—and AWS automatically spreads your spot requests across them. Diversification reduces interruption rates dramatically. Our current spot interruption rate sits around 2%, and when interruptions happen, they’re gracefully handled by our existing high availability setup.
Storage Costs Hide in Plain Sight
Nobody gets excited about storage optimization, which is exactly why it’s such a goldmine. EBS snapshots are the worst offender. Teams religiously create daily snapshots for backup purposes, then forget they exist. We discovered 847 snapshots dating back three years, costing us $1,200 monthly to store complete disk images of long-deleted development instances.
The solution isn’t just deletion, it’s lifecycle automation. AWS Data Lifecycle Manager can automatically create snapshots on a schedule and delete old ones. We set up a policy that keeps daily snapshots for seven days, weekly snapshots for four weeks, and monthly snapshots for one year. Our snapshot storage costs dropped 78%.
S3 Intelligent Tiering deserves more attention than it gets. Enable it on any bucket where you’re unsure about access patterns. It automatically moves objects between storage classes based on actual usage patterns, and the monitoring costs are typically offset by savings within the first month. We enabled it across our user-generated content buckets and saved 45% on storage costs without changing a single line of code.
The Hidden Expense of Data Transfer
Data transfer charges are the cloud equivalent of death by a thousand cuts. They’re small enough to ignore individually but add up to substantial monthly expenses. Cross-region data transfer costs $0.02 per GB, which sounds insignificant until you realize your application is transferring 2TB daily between regions for no good reason.
We discovered our biggest data transfer culprit was our monitoring setup. Our metrics collection was configured to send data from our us-west-2 application instances to our monitoring infrastructure in us-east-1. Moving our monitoring to the same region as our applications eliminated $890 in monthly transfer charges.
CloudFront isn’t just for static assets anymore. We put our API behind CloudFront and saw a 23% reduction in data transfer costs. The caching benefits were minimal for our mostly dynamic API responses, but the geographic distribution meant user requests got served from edge locations instead of always hitting our origin servers across continents.
The Tools That Actually Matter
AWS Cost Explorer is serviceable but limited. For serious cost optimization, you need tools that can break down costs by feature, team, or customer. We use Kubecost for our Kubernetes workloads and CloudHealth for everything else. Both provide the granular visibility needed to identify optimization opportunities that AWS’s native tools miss.
The underrated hero is AWS Compute Optimizer. It analyzes your actual resource utilization and recommends right-sizing opportunities. Unlike generic monitoring that shows you CPU and memory graphs, Compute Optimizer specifically suggests moving from m5.xlarge to m5.large based on your actual usage patterns. Following its recommendations reduced our EC2 costs by 31% over six months.
Cost optimization isn’t a one-time project, it’s an ongoing engineering discipline. The most expensive infrastructure is often the most invisible: the perfectly functional systems quietly burning money in the background while everyone focuses on the next feature sprint. Sometimes the most elegant solution is simply remembering to turn things off.