Post actualizado el día September 27, 2026 by DeiviSanzPlay
A data center of simultaneous.
What the official documentation does not specify is that the peak of augmented retrieval (RAG) over an 8-GPU H100 cluster triggers a 10 transient for the thermal average; reliability engineers size the distribution bars for that 2.3x step that ASHRAE TC 9.9 thermal design manuals do not explicitly cover.
The threshold value that separates a good result from a bad one is the instantaneous Power Usage Effectiveness (PUE), not the annualized one published in sustainability reports. An annual PUE of 1.15 can hide hourly peaks of 1.48 during massive inference windows. Measure with a DC head-end meter with 1 Hz sampling per PDU. If the instantaneous PUE exceeds 1.35 for more than 180 consecutive seconds, the cooling system is operating outside its design envelope and the marginal cost of that inference burst triples compared to the planned value.
Run a load characterization in three scenarios: base inference without context, inference with 80k tokens of context, and RAG with real-time embedding. Instrument the PDU DC bus with a power quality analyzer contracted with the utility is 18% monthly.
In my experience with multi-tenant inference clusters on DGX H100, the most common mistake when applying this is treating the load of an inference data center as stable. It is not. It is a pulsating load with an irregular duty cycle that resembles more the consumption profile of an arc furnace; machine learning is responsible for most of the energy expenditure associated with modern AI. Deep learning models require enormous amounts of compute both for their training and for their continuous operation (inference). However, there is a key distinction: the training phase, although intensive, is a one-time event, while inference—of language models
Despite the increase in total consumption, energy efficiency per individual task has improved at an unprecedented rate. The energy required per AI query has decreased by at least an order of magnitude per year in recent years thanks to advances in software and hardware. A simple text query now typically consumes less electricity than a television turned on for the same period.
However, this efficiency progress is offset by a phenomenon known as the Jevons paradox or “rebound effect”: by being more efficient and cheaper, AI applications are used far more, canceling out the gains. A simple chatbot query can consume 200 times more energy than a text classification; generating an image consumes 1,450 times more; and generating a short video can consume as much as 200,000 spam classifications.
Inference efficiency is measured in tokens per watt (tok/W). Greater efficiency in token processing allows generating more revenue with the same energy consumption. Curiously, tok/W performance can vary by up to 40 times on the same hardware depending on context length. The larger the context window, the fewer tokens can be processed per watt, following an approximate inverse relationship.
Emerging sustainable technologies
Faced with the energy challenge posed by AI, technologies with the potential to transform data center consumption are emerging:
Superconductors: Companies such as VEIR and Snowcap Compute are applying superconductivity to reduce electrical losses and increase power density. Microsoft is already evaluating this technology for its future AI infrastructure.
Photonics: Companies such as Lightmatter, Ayar Labs, and Celestial AI are replacing electrical interconnects with optical ones to reduce electrical losses and heat generation, one of the largest energy costs in moving data between processors.
Edge AI: Shifting workloads from centralized data centers to mobile devices, vehicles, and industrial infrastructure reduces the demand for centralized compute. Companies such as Qualcomm, Nvidia, and Apple are advancing in this direction.
Data centers coupled with energy recovery: Recent research proposes integrating data centers with waste-to-energy (WtE) plants, using low-grade waste heat to power absorption cooling systems, thereby displacing the electricity needed for cooling.
Environmental impact of generative AI
The environmental impact of generative AI goes beyond electricity consumption and spans several critical dimensions:
Carbon footprint: Generative AI generates CO₂ emissions both directly (from data center electricity consumption) and indirectly (from hardware manufacturing). Offsetting ChatGPT’s annual emissions would require 2.6 million tree seedlings for ten years, an area equivalent to Manhattan.
Water footprint: Water consumption for data center cooling and for electricity generation is massive. It is projected that by 2030, AI water consumption will equal the basic needs of 1.3 billion people, the total population of sub-Saharan Africa. ChatGPT’s water footprint already equals the annual water needs of half a million people in that region.
Electronic waste: The rapid obsolescence of AI hardware could generate up to 2.5 million tons of electronic waste per year by 2030, most of it processed in low-income countries with scarce environmental safeguards.
Land consumption: The land footprint of AI infrastructure will exceed 14,500 square kilometers by 2030, more than double the metropolitan area of Jakarta.
Energy use in the cloud
The cloud computing model, although more efficient in aggregate terms than local data centers, concentrates energy consumption. Data center energy efficiency is traditionally measured using Power Usage Effectiveness (PUE), the ratio between the total energy consumed by the facility and that consumed by the computing equipment. A PUE of 1.0 would indicate perfect efficiency; modern hyperscale data centers achieve PUEs as low as 1.05.
However, the pulsating load of LLM inference means that instantaneous PUE can be much higher than the annualized average. It has been observed that an instantaneous PUE above 1.35 for more than 180 consecutive seconds indicates that the cooling system is operating outside its design, tripling the marginal cost of the inference burst.
Storage batteries are emerging as a critical solution for managing these peaks. It is estimated that by 2030, data centers could install between 20-25 GW of battery storage, potentially turning them into an asset for the power grid.
AI and climate change
The relationship between AI and climate change is paradoxical: on the one hand, AI is a major contributor to the problem through its energy consumption; on the other, it is presented as a key tool for mitigation.
The AI sector is contributing to global warming through:
- Direct carbon emissions from data center electricity consumption (often powered by fossil fuels).
- Intensive water use in regions with water stress, such as Querétaro (Mexico) or Uruguay, exacerbating droughts.
- Dependence on critical minerals extracted in countries with little environmental oversight.
However, AI also offers mitigation opportunities:
- Optimization of power grids to integrate renewable energy and reduce losses.
- Monitoring and predictive maintenance of energy infrastructure to prevent failures and blackouts.
- The “rebound effect” of AI on the economy: AI-driven growth could increase GDP, but this growth does not translate one-to-one into greater energy demand, since it is concentrated in knowledge-intensive services where energy-GDP elasticity is lower.
- Potential energy savings in energy-intensive industries of 3-10 percentage points in energy costs through AI optimization.
Future of efficient computing
The future of efficient computing involves multiple complementary strategies:
1. Algorithmic optimization: Research into smaller, specialized models can drastically reduce consumption. Context routing (FleetOpt) can improve energy efficiency (tok/W) by up to 2.5 times compared to a homogeneous fleet.
2. Power Capping Policies: Studies with H100, H200, and MI300X GPUs show that there is no universal optimal power limit: it depends on the workload and the architecture. Configuring at 650W on H100 (compared to the 700W TDP) can offer savings on the contracted peak demand of 18% monthly, with a latency penalty of less than 3%.
3. Standardized measurement: The new UNE 0086:2025 Specification, published in Spain, establishes a pioneering framework for measuring energy consumption, carbon footprint, water consumption, and the performance of AI systems, aligning with international standards such as ISO/IEC TR 20226.
4. Transparency and governance: The United Nations University report on the “Environmental Cost of AI Energy Use” proposes an ecosystem based on transparency, efficiency by design, environmental equity, responsibility across the entire life cycle, global cooperation, and sustainable use.
The efficiency of the future of AI will depend not only on technological advances, but also on global governance and shared responsibility among governments, industry, investors, and users to ensure that the AI revolution develops within the limits of the planet.