In the context of computing power shortages, Amazon AWS requires engineers to reduce CPU resource waste
Complete. Here is the key summaryAgainst the backdrop of a shortage of computing power, Amazon AWS management has requested engineers to conserve resources, including general-purpose CPUs and AI chip computing power. To ensure the supply of EC2 services, the team needs to shut down idle instances and set a goal to complete the reduction of computing power by the end of the first half of the year. This move aims to address the resource constraints of CPUs, memory, cabinets, and other resources caused by the surge in AI demand

According to a participant familiar with the meeting, the management of Amazon Web Services (AWS) held a meeting with engineers in May, delivering a stern directive: to ensure that the popular Elastic Compute Cloud (EC2) service can provide sufficient computing power for all customers in the future, engineers must find every possible way to conserve resources.
The insider stated that the computing power that needs to be saved includes both general-purpose Central Processing Unit (CPU) server resources and the long-sought-after AI-specific chip computing power. For decades, the internet industry has relied on CPU chips to operate; however, many engineers within AWS have reported that the waiting time for CPU server computing power applications for research and development has significantly increased.
Some engineers mentioned that server computing power that used to be approved in just a few hours now takes several days, which can easily lead to project delays. This engineer, who has worked at AWS for many years, has never encountered such a long resource scheduling period.
Additionally, individuals familiar with the internal planning revealed that AWS has set a deadline for all teams to complete the reduction of computing power by the end of this year. Engineers are shutting down idle EC2 virtual servers (i.e., cloud instances) after software development is completed to free up this computing power for external customers.
The computing power shortage that has swept the entire technology industry this year has now spread to the traditional general-purpose computing power sector. It is well known that the skyrocketing demand for AI chips, such as NVIDIA graphics processing units (GPUs), has directly caused a shortage of AI-specific computing power; meanwhile, driven by the AI industry, CPU servers are also facing supply constraints. The insufficient production capacity of memory chips used with CPUs and the limited physical cabinet space in data centers have further exacerbated the computing power gap.
Xie Jing (phonetic), co-founder and director of financial AI service provider Elendil Labs, analyzed that more and more employees are using AI agents to develop software, leading to a surge in corporate CPU consumption. The client companies she has encountered have seen their average IT spending per employee double—much of the work is handled by AI agents, necessitating the purchase of more cloud computing power, which generally relies on CPUs.
Even in the AI model development process, CPUs play a critical role: for example, when AI companies preprocess data before training models, they rely on CPUs to read documents, images, videos, and other raw materials.
Xie Jing stated, "Currently, the demand for CPUs in the vast majority of model development, generation, and operation tasks is far higher than before."
Intel CEO Pat Gelsinger revealed during the April earnings call that the ratio of CPU to GPU computing power in AI inference (model online operation) scenarios was 1:4 at that time; by July, Intel CFO David Zinsner stated that the ratio had approached 1:1, with executives from AMD and Arm providing similar assessments AWS denied that its operational strategy has changed and stated in an official announcement: "Even with sustained high demand, we can meet the computing power needs of the vast majority of internal and external customers. We always collaborate with internal teams to ensure their computing power supply while urging teams to use EC2 resources efficiently; this management policy has never changed."
Amazon stated that the guidelines for internal engineers optimizing resources are completely consistent with the recommendations for external customers: for example, shutting down idle cloud instances and switching to instance specifications that match their business load. If a customer's current instance's computing power is not fully utilized, they can switch to a lower configuration model.
AWS provided clear optimization standards: if a certain EC2 server has an average CPU and memory usage rate of less than 40% for four consecutive weeks, it is recommended that customers downgrade to a smaller specification instance.
AWS also offers an automated scanning tool that can automatically identify servers with low utilization. Amazon emphasized that even with the current shortage of memory chips, the resource management standards for internal employees have not been adjusted.
In the past two years, balancing internal R&D computing power with external customer computing power supply has become a common challenge for major tech giants. Last year, Google established an executive special committee to coordinate the allocation of computing power resources among Google Cloud, DeepMind AI Lab, and C-end business sectors, but conflicts over computing power remain prominent. Reports indicate that this summer, top AI researcher Noam Shazeer left due to limited computing resources and hindered R&D.
If companies can finely manage computing power resources, they can more quickly convert idle computing power into revenue. Microsoft CFO Amy Hood stated in last month's latest earnings call that Azure cloud service performance growth is partly due to the team's efficient scheduling and optimization of CPU and GPU computing clusters.
Hood stated: "This quarter, the engineering team has made significant progress in revitalizing computing power; due to the impact of the supply-demand imbalance, as long as we improve computing power utilization efficiency, the newly available resources can quickly be monetized."
An AWS industry consultant stated that there has not yet been a supply gap for reserved CPU computing power locked in by contracts, but in recent months, the difficulty of applying for AWS bidding instances (vendors' idle surplus computing power sold at low prices, with resources recoverable two minutes in advance) has significantly increased
