Understand the system
How it works
The rack is emerging as the unit of AI system design, coupling silicon performance to power, networking, mechanics, operations, and thermal limits.
What makes a rack operate
Select a capability to explore its role and connections.
Selected: AI accelerator
AI accelerator
Packaged compute and memory device optimized for training and inference workloads.
Explore AI accelerator in AtlasDocumented connections
- AI accelerator → AI rack system
Accelerator modules are integrated with power, cooling, network, and service systems.
Relationship evidence · 2
- [1] NVIDIA's DGX GB200 documentation describes a rack-scale system with 72 GPUs, 36 Grace CPUs, and an approximately 120 kW power envelope, with liquid cooling alongside air-cooled components.
- [2] A DGX H200 system contains eight H200 GPUs and has a documented maximum input power of 10.2 kW; the system manual specifies front-to-back airflow.
Selected documented dependencies. Arrows retain their Atlas direction; they do not represent quantities or a complete engineering process.
An accelerator headline hides the rest of the computer. Host CPUs feed work, DPUs manage data, local memory holds state, scale-up links coordinate devices, NICs reach the cluster, storage supplies datasets, and firmware keeps the whole assembly observable. As densities rise, rack busbars, power shelves, coolant manifolds, cable routing, floor loading, and maintenance clearance become architectural constraints.
The heterogeneous node
AI servers combine accelerators, host processors, memory, storage, NICs, management controllers, and power conversion. Bottlenecks emerge at the boundaries between those components.
- Expose data paths through the whole server
- Separate accelerator power from platform power
Scale-up inside the machine
High-bandwidth, low-latency fabrics let accelerators exchange tensors and memory traffic as one system. Topology and collective behavior determine how close scaling comes to linear.
- Collectives follow the physical topology
- Synchronization consumes usable bandwidth
Rack-scale power delivery
At high density, upstream transformers, switchgear, busways, power shelves, backup strategy, and monitoring must be designed as one conversion chain from grid to package.
- Every voltage conversion introduces losses
- Rated and measured demand must remain distinct
Supporting evidence · 2
- [1] NVIDIA's DGX GB200 documentation describes a rack-scale system with 72 GPUs, 36 Grace CPUs, and an approximately 120 kW power envelope, with liquid cooling alongside air-cooled components.
- [2] Open Compute Project workstreams are addressing future AI rack power envelopes from 250 kW toward 1 MW alongside liquid-cooling standards.
Serviceability is availability
Cables, blind-mate connectors, coolant lines, firmware, spares, and safe maintenance procedures affect mean time to repair. Dense compute is valuable only while the cluster can schedule it.
- Replaceable units need safe maintenance paths
- Repair time reduces effective utilization
Supporting evidence · 3
- [1] NVIDIA's DGX GB200 documentation describes a rack-scale system with 72 GPUs, 36 Grace CPUs, and an approximately 120 kW power envelope, with liquid cooling alongside air-cooled components.
- [2] CUDA and ROCm each combine programming models, runtimes, compilers, libraries, debugging tools, and hardware interfaces into an accelerator software stack.
- [3] NVIDIA documents a maximum of 64 framebuffer page retirements for legacy GPUs and up to 512 framebuffer row-remapping entries for Ampere and later generations.
Featured evidence
Evidence in context
Cumulative 100%-nameplate rack energy scenario
A derived upper-bound-style scenario using the documented approximate NVL72 rack power continuously through one year.View chart values
| Category | Cumulative nameplate scenario |
|---|---|
| Month 0 | 0 GWh |
| Month 3 | 0.26 GWh |
| Month 6 | 0.53 GWh |
| Month 9 | 0.79 GWh |
| Month 12 | 1.05 GWh |
Key indicators
Key measures & constraints
- 01
NVIDIA's DGX GB200 documentation describes a rack-scale system with 72 GPUs, 36 Grace CPUs, and an approximately 120 kW power envelope, with liquid cooling alongside air-cooled components.
- Class
- vendor spec
- Geography
- Global
- Period
- DGX GB200 platform
- Confidence
- high
- 02
A DGX H200 system contains eight H200 GPUs and has a documented maximum input power of 10.2 kW; the system manual specifies front-to-back airflow.
- Class
- vendor spec
- Geography
- Global
- Period
- DGX H200 system specification
- Confidence
- high
- 03
Holding NVIDIA's approximately 120 kW DGX GB rack specification at 100 percent for 8,760 hours yields a transparent upper-use scenario of 1.0512 GWh per year; it is not a measured average and excludes facility overhead.
- Class
- estimate
- Geography
- Global scenario
- Period
- One 8,760-hour year using March 2026 specification
- Confidence
- high
- 04
NVIDIA documents a maximum of 64 framebuffer page retirements for legacy GPUs and up to 512 framebuffer row-remapping entries for Ampere and later generations.
- Class
- vendor spec
- Geography
- Global
- Period
- Documentation reviewed August 2026
- Confidence
- high
System map
Connections across the system
Explore inputs, outputs, and shared capabilities in the Atlas. Open the register for every documented relationship.
Enters this system 8
- Advanced packaging→ AI acceleratorTrace relationship
- Memory fabrication→ High-bandwidth memoryTrace relationship
Leaves this system 6
- Advanced packaging← High-bandwidth memoryTrace relationship
- Accelerator architecture← Runtime & frameworksTrace relationship
Shared capabilities can belong to more than one System. These links carry materials, energy, information, capital, or permissions; they describe dependencies, not quantities or a complete process model.
Explore 10 capabilities and operating boundaries
All documented relationships · 26
- criticaloutput
Qualified HBM stacks are bonded beside logic in advanced packages.
Trace relationship - criticalinput
Assembly and test turn logic and HBM into a qualified accelerator module.
Trace relationship - criticalinternal
Memory capacity and bandwidth bound accelerator workload performance.
Trace relationship - criticalinternal
Accelerator modules are integrated with power, cooling, network, and service systems.
Trace relationshipSupporting evidence · 2
- [1] NVIDIA's DGX GB200 documentation describes a rack-scale system with 72 GPUs, 36 Grace CPUs, and an approximately 120 kW power envelope, with liquid cooling alongside air-cooled components.
- [2] A DGX H200 system contains eight H200 GPUs and has a documented maximum input power of 10.2 kW; the system manual specifies front-to-back airflow.
- criticalinternal
Drivers, kernels, frameworks, and scheduling convert hardware into useful work.
Trace relationshipSupporting evidence · 2
- criticalinput
Qualified DRAM dies from memory fabs become the active layers of HBM stacks.
Trace relationshipSupporting evidence · 2
- criticalinput
Final qualification determines whether integrated accelerator packages can enter systems.
Trace relationshipSupporting evidence · 2
- [1] NVIDIA documents a maximum of 64 framebuffer page retirements for legacy GPUs and up to 512 framebuffer row-remapping entries for Ampere and later generations.
- [2] SIA's high-level semiconductor value chain runs from research and development through chip design, front-end wafer fabrication, back-end assembly, test and packaging, and finally circuit-board integration into products.
- criticalinternal
Server and tray integration turns qualified components into serviceable rack-scale compute.
Trace relationshipSupporting evidence · 2
- criticalinternal
Rack power conversion and distribution deliver usable electrical capacity to compute trays.
Trace relationshipSupporting evidence · 2
- [1] NVIDIA's DGX GB200 documentation describes a rack-scale system with 72 GPUs, 36 Grace CPUs, and an approximately 120 kW power envelope, with liquid cooling alongside air-cooled components.
- [2] Open Compute Project workstreams are addressing future AI rack power envelopes from 250 kW toward 1 MW alongside liquid-cooling standards.
- importantoutput
Software capabilities and workload behavior feed back into hardware design.
Trace relationship - importantinternal
Host processors dispatch work and manage data around accelerator execution.
Trace relationship - importantinternal
Host and infrastructure processors complete the rack compute platform.
Trace relationship - importantinput
Scale-up links couple accelerators within the rack-scale computer.
Trace relationship - importantoutput
Rack endpoints attach to the cluster's scale-out switching fabric.
Trace relationshipSupporting evidence · 2
- importantoutput
Parallelism and collective libraries shape network traffic and topology needs.
Trace relationship - importantinput
Datasets and checkpoints feed training and recovery workflows.
Trace relationship - importantoutput
Rack power, mass, cooling, and service envelopes constrain facility design.
Trace relationshipSupporting evidence · 2
- [1] NVIDIA's DGX GB200 documentation describes a rack-scale system with 72 GPUs, 36 Grace CPUs, and an approximately 120 kW power envelope, with liquid cooling alongside air-cooled components.
- [2] Open Compute Project workstreams are addressing future AI rack power envelopes from 250 kW toward 1 MW alongside liquid-cooling standards.
- importantinput
Liquid loops capture and transport heat from high-density rack components.
Trace relationshipSupporting evidence · 2
- [1] NVIDIA's DGX GB200 documentation describes a rack-scale system with 72 GPUs, 36 Grace CPUs, and an approximately 120 kW power envelope, with liquid cooling alongside air-cooled components.
- [2] Open Compute Project workstreams are addressing future AI rack power envelopes from 250 kW toward 1 MW alongside liquid-cooling standards.
- importantinput
Export rules can restrict advanced memory products and related technology.
Trace relationship - importantinternal
Refresh and retirement send systems into reuse, parts recovery, or recycling.
Trace relationship - importantinternal
Server & ODM integration supplies a distinct operating capability within its broader Atlas system.
Trace relationshipSupporting evidence · 2
- importantinternal
Rack power delivery supplies a distinct operating capability within its broader Atlas system.
Trace relationshipSupporting evidence · 2
- [1] NVIDIA's DGX GB200 documentation describes a rack-scale system with 72 GPUs, 36 Grace CPUs, and an approximately 120 kW power envelope, with liquid cooling alongside air-cooled components.
- [2] Open Compute Project workstreams are addressing future AI rack power envelopes from 250 kW toward 1 MW alongside liquid-cooling standards.
- importantinternal
System memory hierarchy supplies a distinct operating capability within its broader Atlas system.
Trace relationship - importantinternal
Firmware & rack management supplies a distinct operating capability within its broader Atlas system.
Trace relationshipSupporting evidence · 2
- [1] NVIDIA documents a maximum of 64 framebuffer page retirements for legacy GPUs and up to 512 framebuffer row-remapping entries for Ampere and later generations.
- [2] CUDA and ROCm each combine programming models, runtimes, compilers, libraries, debugging tools, and hardware interfaces into an accelerator software stack.
- importantinput
Repair, reuse & recycling supplies a distinct operating capability within its broader Atlas system.
Trace relationshipSupporting evidence · 3
- [1] U.S. EPA guidance places reuse and redeployment of functional electronics ahead of recycling, with refurbishment, repair, component harvesting, and secure data handling used to retain equipment value where practicable.
- [2] The world generated 62 million tonnes of electronic waste in 2022, while 22.3 percent was documented as formally collected and recycled.
- [3] UNITAR projects global electronic waste to reach 82 million tonnes in 2030 under current trends.
- enablingoutput
Recovery can return selected metals to material supply and reduce virgin demand.
Trace relationship
Named companies + institutions
Organizations & their roles
- principal · consortium
CXL Consortium
Compute Express Link ConsortiumIndustry consortium maintaining the Compute Express Link cache-coherent interconnect standard.
Roles, locations & evidence
- Role
- standards setter
- Products / service
- coherent memory interconnect standards
- Documented activity
- World
- Headquarters
- Not cataloged
Why this organization belongs in the System- System memory hierarchy
Primary-source evidence connects this institution to coherent memory interconnect standards in the specified Atlas capability.
- principal · nonprofit
Open Compute Project
Open Compute Project FoundationIndustry foundation developing open specifications and practices for scalable computing infrastructure.
Roles, locations & evidence
- Role
- standards setter
- Products / service
- open rack and AI-cluster infrastructure specifications
- Documented activity
- World
- Headquarters
- United States
Why this organization belongs in the System- AI rack system
Primary-source evidence connects this institution to open rack and ai-cluster infrastructure specifications in the specified Atlas capability.
- principal · standards body
PCI-SIG
PCI Special Interest GroupIndustry standards body maintaining the PCI Express interconnect specification.
Roles, locations & evidence
- Role
- standards setter
- Products / service
- PCI Express component interconnect standards
- Documented activity
- World
- Headquarters
- Not cataloged
Why this organization belongs in the System- Host CPU & DPU
Primary-source evidence connects this institution to pci express component interconnect standards in the specified Atlas capability.
- material · public company
Arista Networks
Arista Networks, Inc.Cloud-networking company supplying high-speed switching systems, routing platforms, and network software.
Roles, locations & evidence
- Role
- supplier
- Products / service
- Network operating software
- Documented activity
- World
- Headquarters
- United States
Why this organization belongs in the System- Runtime & frameworks
Arista Networks's filing-backed operating portfolio includes network operating software; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
All other organizations · 12
- material · public company
Cisco
Cisco Systems, Inc.Networking and security company supplying switching, routing, optical, observability, and data-center infrastructure.
Roles, locations & evidence
- Role
- supplier
- Products / service
- Network operating and observability software
- Documented activity
- World
- Headquarters
- United States
Why this organization belongs in the System- Runtime & frameworks
Cisco's filing-backed operating portfolio includes network operating and observability software; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
- material · public company
Eaton
Eaton Corporation plcPower-management company supplying electrical distribution, switchgear, UPS, and power-quality systems.
Roles, locations & evidence
- Role
- supplier
- Products / service
- Power distribution and conversion
- Documented activity
- World
- Headquarters
- Ireland
Why this organization belongs in the System- Rack power delivery
Eaton's filing-backed operating portfolio includes power distribution and conversion; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
- material · public company
HPE
Hewlett Packard Enterprise CompanyEnterprise infrastructure company supplying servers, supercomputers, networking, storage, and services.
Roles, locations & evidence
- Role
- supplier · manufacturer
- Products / service
- Integrated rack-scale systems · AI and high-performance computing systems
- Documented activity
- World
- Headquarters
- United States
Why this organization belongs in the System- AI rack system
HPE's filing-backed operating portfolio includes integrated rack-scale systems; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
- Server & ODM integration
HPE's filing-backed operating portfolio includes ai and high-performance computing systems; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
- material · nonprofit
PyTorch Foundation
Open-source foundation supporting the PyTorch machine-learning framework and distributed-computing ecosystem.
Roles, locations & evidence
- Role
- standards setter
- Products / service
- distributed machine-learning software
- Documented activity
- World
- Headquarters
- Not cataloged
Why this organization belongs in the System- Runtime & frameworks
Primary-source evidence connects this institution to distributed machine-learning software in the specified Atlas capability.
- material · research institution
UNITAR
United Nations Institute for Training and ResearchUnited Nations institute publishing the Global E-waste Monitor with international partners.
Roles, locations & evidence
- Role
- community party
- Products / service
- global electronics-lifecycle evidence
- Documented activity
- World
- Headquarters
- Not cataloged
Why this organization belongs in the System- Hardware lifecycle
Primary-source evidence connects this institution to global electronics-lifecycle evidence in the specified Atlas capability.
- representative · public company
AMD
Advanced Micro Devices, Inc.Semiconductor company that designs CPUs, GPUs, adaptive computing products, and associated AI software.
Roles, locations & evidence
- Role
- designer · supplier
- Products / service
- Instinct data-center accelerators · EPYC server processors · ROCm software platform
- Documented activity
- World
- Headquarters
- United States
Why this organization belongs in the System- AI accelerator
AMD's filing-backed operating portfolio includes instinct data-center accelerators; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
- Host CPU & DPU
AMD's filing-backed operating portfolio includes epyc server processors; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
- Runtime & frameworks
AMD's filing-backed operating portfolio includes rocm software platform; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
- representative · public company
Dell Technologies
Dell Technologies Inc.Enterprise technology company supplying servers, storage, networking, and integrated AI infrastructure.
Roles, locations & evidence
- Role
- supplier · manufacturer
- Products / service
- Rack-scale AI infrastructure · AI servers and integrated systems
- Documented activity
- World
- Headquarters
- United States
Why this organization belongs in the System- AI rack system
Dell Technologies's filing-backed operating portfolio includes rack-scale ai infrastructure; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
- Server & ODM integration
Dell Technologies's filing-backed operating portfolio includes ai servers and integrated systems; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
- representative · public company
Intel
Intel CorporationIntegrated semiconductor company spanning processor design, fabrication, packaging, systems, and foundry services.
Roles, locations & evidence
- Role
- designer
- Products / service
- Xeon server processors
- Documented activity
- World
- Headquarters
- United States
Why this organization belongs in the System- Host CPU & DPU
Intel's filing-backed operating portfolio includes xeon server processors; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
- representative · public company
Micron
Micron Technology, Inc.Memory and storage manufacturer supplying DRAM, HBM, NAND, and related products for AI systems.
Roles, locations & evidence
- Role
- manufacturer
- Products / service
- High-bandwidth memory · HBM advanced-packaging facility under construction
- Documented activity
- World · Singapore
- Headquarters
- Boise
Why this organization belongs in the System- High-bandwidth memory
Micron's filing-backed operating portfolio includes high-bandwidth memory; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
- High-bandwidth memory
Micron's project updates directly locate the HBM packaging buildout in Singapore and retain its contribution as future, not operating, capacity.
- representative · public company
NVIDIA
NVIDIA CorporationComputing-platform company that designs AI accelerators, interconnects, systems, and the software stack that operates them.
Roles, locations & evidence
- Role
- designer · supplier
- Products / service
- Data-center GPU accelerators · Rack-scale AI systems · CUDA software platform
- Documented activity
- World
- Headquarters
- United States
Why this organization belongs in the System- AI accelerator
NVIDIA's filing-backed operating portfolio includes data-center gpu accelerators; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
- AI rack system
NVIDIA's filing-backed operating portfolio includes rack-scale ai systems; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
- Runtime & frameworks
NVIDIA's filing-backed operating portfolio includes cuda software platform; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
- representative · public company
Supermicro
Super Micro Computer, Inc.Server and systems manufacturer supplying modular, rack-scale, and liquid-cooled AI infrastructure.
Roles, locations & evidence
- Role
- manufacturer
- Products / service
- Rack-scale AI systems · AI servers and rack integration
- Documented activity
- World
- Headquarters
- United States
Why this organization belongs in the System- AI rack system
Supermicro's filing-backed operating portfolio includes rack-scale ai systems; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
- Server & ODM integration
Supermicro's filing-backed operating portfolio includes ai servers and rack integration; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
- representative · public company
Vertiv
Vertiv Holdings CoCritical-digital-infrastructure company supplying power, thermal management, and service systems for data centers.
Roles, locations & evidence
- Role
- supplier
- Products / service
- Rack and facility power systems
- Documented activity
- World
- Headquarters
- United States
Why this organization belongs in the System- Rack power delivery
Vertiv's filing-backed operating portfolio includes rack and facility power systems; no subregional operating place is asserted, and headquarters is not used as a proxy. Representative status does not imply market rank.
Affected policies
Relevant policies
Legal status is separated from policy objective. Every record keeps its jurisdiction, mechanism, affected nodes, and verification date.
- effective
Advanced-semiconductor export controls, as amended
Restrict specified capabilities used to produce advanced semiconductors and advanced computing systems in China.
Mechanism, scope & sources
Combines controlled equipment, software, HBM, entity, end-use, and license requirements with later transaction-specific exceptions and case-by-case review policies in the current EAR framework.
United States export administrationOpen full policy record
Read this System in the field guide
Continue in the Volumes
Start with three selected readings, or explore the complete reading list.
Complete reading list
Machine
13 spreads- IV · 01From accelerator to nodeOpen spread
- IV · 02Node anatomyOpen spread
- IV · 03Host interconnect and memory expansionOpen spread
- IV · 04Firmware, management, and securityOpen spread
- IV · 05Tray, chassis, and rackOpen spread
- IV · 06Scale-up fabricOpen spread
- IV · 07Rack power deliveryOpen spread
- IV · 08Cabling, materials, and mechanicsOpen spread
- IV · 09Heat and sustained performanceOpen spread
- IV · 10The rack liquid loopOpen spread
- IV · 11Reliability and serviceabilityOpen spread
- IV · 12Build, ship, install, acceptOpen spread
- IV · 13Refresh, reuse, and retirementOpen spread
Evidence & review
currentEvidence health35 active claims · next review Oct 13, 2026+
- Active claims
- 35
- Sources
- 48
- High volatility
- 7
- Next review
- Oct 13, 2026
- Due soon
- 0
- Overdue
- 0
- Superseded
- 0