ULA · Sequoia-backed B2B marketplace
100% Spot in production e-commerce
Ran all production compute on EC2 Spot Instances with graceful interruption handling, at a ~$33K/month AWS baseline. Most teams abandon Spot after their first interruption incident.
- Stateless services with graceful drain on interruption notice; zero stateful dependencies on Spot nodes.
- Queue-backed async workloads with idempotent consumers; mixed-instance ASGs for diversification.
- Held the infrastructure line during a 2-month, ~10x traffic burst (AWS spend peaked at ~$66K/month).
100%
Of production compute on EC2 Spot
~$33K/mo
AWS baseline for a regional B2B marketplace
2 months
Traffic burst held (Ramadan campaign, ~$66K/mo peak)
How 100% Spot works in production
Stateless services with graceful drain on interruption notice. Queue-backed async workloads with idempotent consumers. Autoscaling groups with mixed instance policies for diversification. And the rule that makes it all safe: zero stateful dependencies on Spot nodes.
The traffic burst — and scope discipline
During a Ramadan sale campaign, traffic spiked roughly 10x and AWS spend peaked at ~$66K/month for about two months. The root cause of the bottleneck was a Lambda authorizer — an application-layer problem the engineering team solved with bucket-based rate limiting.
My job was different, and I kept it that way: observability during the incident, on-call, and coordinating with AWS Enterprise Support to raise service limits and enable max scale-out — giving the team room to breathe while they built the fix. Knowing what’s yours and what isn’t is what keeps incident response fast.