Benchmarking the power demand of AI workloads within data centers.

Overview

Artificial intelligence workloads are computationally dense, requiring specialized hardware and consuming large amounts of power. These workloads can also strain power infrastructure, as they exhibit rapid fluctuations in power demand that can impact local power quality. Addressing these issues and the overall energy footprint of AI requires careful measurement of power use of the relevant hardware. 

We are empirically benchmarking the power demand of AI training and inference workloads. This is across multiple hardware configurations, power capping regimes, and cooling systems. These will then enable the development of predictive models of total energy across a variety of deployment scenarios. 


Our Work

Researchers from Lawrence Berkeley National Laboratory, Brookhaven National Lab, Florida Atlantic University, and Kansas State University partnered with Johnson Controls,  Nvidia, and SuperMicro to compare the performance of NVIDIA AI nodes when air and liquid cooled. This work explored how the cooling system affects both energy and computational performance. 

Impacts

  • Direct-to-chip liquid cooling generally reduces total cooling related energy use (PUE) in heat dense systems. LBNL researchers are working to quantify this performance gap on contemporary hardware configurations. 
  • Computational Throughput depends on job, processor, and the thermal performance of the system. LBNL researchers evaluate how the improved cooling performance of direct to chip liquid cooling keeps chips outside of thermal throttling envelopes, allowing them to complete jobs faster.
  • Total-Power-Usage-Effectiveness (TUE) can differ from measured PUE due to embedded server fan modules. LBNL works with hardware manufacturers to develop the TUE metric, and evaluate how it can differ from measured PUE in direct to chip liquid cooled systems.
Project Highlights
Predictive Models

Using empirical measurements of AI workloads, researchers are developing predictive models of both energy use and average power draw.  Characterizing training power draw data allows utilities to better understand power quality challenges these facilities can pose, and the energy model allows facilities to better predict their operating costs.

Benchmarking Liquid Cooling Hardware

AI hardware changes performance dynamically in response to thermal status. The same job on the same chips can differ in performance based on the cooling systems. Liquid cooling provided improved thermal performance that can complete jobs faster by maintaining temperature below throttling thresholds. The improvements to throughput mean facilities that adopt liquid cooling unlock additional computational performance over and above the improved cooling economics

Computational Throughput

GPUs can be adjusted in terms of their maximum power, with implications for both computational throughput and energy use. Researchers in ETA and Computing Sciences are empirically characterizing how adjustments to maximum GPU power can improve energy performance, and what affects this has on runtime, throughput, and overall operational performance.


Team Members
Policy Researcher
Energy Technology Researcher