cloud
Kubernetes Autoscaling Without the Surprise Bill
Autoscaling is usually sold as a cost-saving feature, but misconfigured it does the opposite: pods scale up under load and never scale back down because resource requests are set too conservatively, or a single noisy deployment keeps utilization metrics artificially high across the whole node pool.
Start with accurate resource requests and limits before touching the Horizontal Pod Autoscaler. If requests are set well above actual usage "just to be safe," the cluster autoscaler will provision nodes to match those inflated numbers, and no amount of HPA tuning downstream will fix a bin-packing problem that starts at the pod spec.
Cluster autoscaler and Karpenter (or equivalent node-provisioning tools) behave differently under scale-down: Karpenter tends to consolidate more aggressively by default, which is usually what you want for cost, but can cause more pod churn than teams running stateful or latency-sensitive workloads are comfortable with. Test scale-down behavior in staging under realistic load before trusting it in production.
Set a monthly review of actual versus requested resource usage per namespace — this single habit catches more waste than any autoscaler configuration change, because it surfaces the services nobody remembered were still running at 3x their real footprint.
Apeniq builds and tunes Kubernetes platforms for teams that outgrew their initial cluster setup — reach out if your cloud bill keeps climbing despite adding autoscaling.