Cloud cost optimization is a perennial challenge for infrastructure teams, especially when it comes to sizing virtual machines (VMs). Inefficiently sized VMs lead to wasted spend, degraded performance, or both. Azure provides Azure Advisor to guide VM resize recommendations, while AWS offers AWS Compute Optimizer for similar purposes.
In this blog post, I’ll walk through how to effectively leverage Azure Advisor for resizing decisions. I’ll also highlight critical considerations that are often overlooked during these evaluations: the difference in shared CPU definitions across cloud providers, the importance Check over here of looking beyond average CPU utilization by analyzing P95/P99 percentiles and spike duration, and how always-on small services can accumulate cloud waste. Finally, I’ll touch on outbound networking's impact on cost and performance considerations.
Introduction: The Myth of Average CPU and VM Resizing
One of the most common pitfalls when resizing VMs based on cloud recommendations is over-reliance on average CPU utilization metrics. The temptation is to look at a 24-hour average CPU and think, “This VM is only using 20% CPU, let’s downsize it.” The reality is far more complex.
- Spike duration and frequency: A VM that averages 20% CPU but regularly spikes to 90% CPU for a few seconds can see significant performance degradation if downsized. P95/P99 CPU utilization metrics: These percentile metrics show what the CPU usage looks like during most of the "worst case" moments, not just the average. Shared CPU cloud instance types: Definition of vCPUs and shared CPUs vary by provider—this impacts perceived performance guarantees. Always-on small services: Small VMs often run 24/7 and may not just waste CPU but also other resources like memory, storage, and networking—which all contribute to cost.
Azure Advisor: What It Is and What It Recommends
Azure Advisor is a native Azure service that provides personalized, actionable best practices recommendations to optimize Azure resources. When it comes to VM resizing, Azure Advisor analyzes historical usage data, including CPU, memory, and disk I/O, to suggest when to resize up or down your VMs.

Azure Advisor’s VM recommendations typically fall into two categories:
Right-size VM Recommendations: Suggest downsizing or upsizing based on observed resource utilization. Stop/Start Recommendations: For VMs that are idle for long periods.What Metrics Does Azure Advisor Use?
Azure Advisor primarily looks at metrics averaged over a period (typically 7 days by default) such as:
- Average CPU utilization Memory utilization (if enabled) Network throughput Disk I/O
However, Azure Advisor currently does not expose detailed percentile or spike-duration analytics directly. This is worth noting before blindly relying on its resize recommendations.
Comparing Azure Advisor with AWS Compute Optimizer
For those familiar with AWS, AWS Compute Optimizer offers VM resizing suggestions with a bit more granularity around performance histograms, including recommendations based on P95/P99 CPU usage distribution. The service considers how long high utilization periods last to avoid downsizing that may cause performance issues during workload spikes.
Azure Advisor’s VM recommendations can be complemented with Azure Monitor metrics that allow you to analyze utilization with similar percentile breakdowns, though this requires manual setup.
Guidelines for Using Azure Advisor Effectively for VM Resize Decisions
1. Look Beyond Averages: Gather P95 and P99 CPU Utilization Metrics
Before resizing, inspect the distribution of CPU utilization. This helps answer:
- Do I see frequent short spikes that would be masked in an average? How long do these spikes last?
How to do this in Azure:

Only after this careful analysis should you act on Azure Advisor’s recommendations.
2. Understand the Shared CPU Definitions
Different cloud providers define a “vCPU” or “shared CPU” differently:
Provider Shared CPU Definition Common Misconception Azure Some VM series (like B-series) provide credit-based bursting on a physical CPU. CPU credits accumulate during idle periods and allow bursts above baseline performance. Assuming all vCPUs behave like reserved physical cores. AWS T-series instances work similarly with CPU credits; C-series have dedicated vCPUs mapped to physical hardware threads. Assuming that a high vCPU count guarantees sustained performance without checking burst budget. Google Cloud vCPUs are more strictly defined as hardware hyperthreads backed by underlying physical cores. Assuming shared CPU equals compromised uptime and performance.When resizing VMs based on Azure Advisor, make sure to consider whether the VM series uses bursting or has dedicated CPU allocation.
3. Factor in Always-on Small Services and Hidden Cost Components
Small, always-on VMs running internal tools, staging environments, or worker queues can quietly accumulate cloud waste:
- Idle CPU or small CPU bursts might look cheap but multiply that cost by thousands of VMs across environments. Storage and outbound networking charges often get overlooked. For Azure, outbound data egress is billed and can be a significant cost factor, especially with VMs communicating intensively with other services.
Always review your costs holistically, including:
- Disk IOPS and capacity utilization. Outbound network egress. Azure charges egress differently depending on target regions and services - evaluate if resizing affects network patterns. Memory Utilization. Downsizing too aggressively on memory-starved VMs causes crashing and restarts, negating any cost savings.
4. Set Rollback Criteria Before Resizing Pilots
Before acting on recommendations, establish clear criteria for rollback to ensure safety:
- What P95 CPU and memory utilization thresholds must hold after resizing? Are application latency and throughput metrics within SLA targets? Is outbound network performance unaffected or improved? Is there a mechanism for automatic rollback if performance deteriorates (e.g., via infrastructure-as-code changes, monitoring alerts)?
Example Workflow: Using Azure Advisor with Custom Metrics for Safe VM Resizing
Identify candidate VMs: Use Azure Advisor to see right-sizing recommendations. Gather metrics: Pull 7 to 14 days of CPU utilization, focusing on P95 and P99 percentiles using Azure Monitor Metrics Explorer. Analyze spike duration: Plot CPU usage over time to understand spike frequency and length. Evaluate memory and network metrics: Include memory pressure and outbound network egress to ensure resizing won’t cause disruptions or additional costs. Select pilot VMs: Choose a small subset for pilot resizing. Deploy resized instances: Downsize or upsize based on combined analysis. Monitor post-deployment: Track SLAs, P95 CPU, memory, latency, and network egress continuously. Go/no-go decision: Use rollback criteria predefined to decide whether to continue or revert.Final Thoughts
Azure Advisor is a powerful tool but should not be viewed as an oracle for VM resizing decisions. It is best used as part of a more comprehensive approach:
- Leverage Azure Advisor recommendations as initial inputs, not final decisions. Incorporate detailed monitoring data, especially P95 and P99 CPU utilization and spike duration analysis. Understand how your VM class semantics (shared CPU, bursting, dedicated) impact performance perception. Consider indirect costs such as outbound networking impacts when resizing. Always design pilot experiments with rollback criteria to minimize risk.
By combining Azure Advisor’s automated insights with rigorous percentile-based observation and provider-specific VM characteristics, you can achieve cost savings without https://smoothdecorator.com/how-do-i-use-p90-p95-and-p99-5-to-classify-cpu-demand/ sacrificing the reliability and performance your services depend on.
Additional Resources
- Azure Advisor Overview Azure Monitor Metrics Guide AWS Compute Optimizer Documentation Azure Outbound Bandwidth Pricing