Explainer
The memory wall: why moving data limits AI hardware
The memory wall is the gap between how fast processors can calculate and how fast data can reach them. One 2024 analysis found server compute growing about 3.0x every two years against 1.6x for DRAM bandwidth, making memory the main bottleneck for AI, especially when serving models.[1][2]
The problem in one picture
Imagine a chef who can chop vegetables incredibly fast, but whose ingredients arrive through a narrow hatch. The chef spends most of the time waiting. Modern AI chips are that chef: they can do arithmetic much faster than memory can deliver the numbers. Researchers measured that peak server computing power grew about three times every two years, while memory bandwidth grew only about 1.6 times.[1]
Gholami and colleagues, writing in IEEE Micro, report peak server FLOPS scaling at 3.0x per two years, against 1.6x for DRAM bandwidth and 1.4x for interconnect bandwidth.[1] For decoder-style language models, whose generation step reads every weight for each token, they argue memory rather than compute is now the primary bottleneck, particularly in serving.[2]
Why moving data is expensive
Every time a model’s weights travel from memory to the processor, it costs time and energy. IBM researchers put it plainly: inference on traditional architectures pays that cost every time weights move.[3] The same applies between chips: in AI data centers, the pluggable optical transceivers that link servers were reported to use around 10% of total GPU compute power.[4]
Four ways hardware fights the wall
- Put memory closer. Advanced packaging places stacks of high-bandwidth memory right next to the processor. NVIDIA’s Blackwell GPU, for example, sits in one package with eight memory stacks.[5]
- Stack chips directly. Hybrid bonding joins chips with very dense copper connections, so data travels microns rather than millimetres.[6]
- Use light. Co-packaged optics moves data between chips as light, with less power than electrical links.[7]
- Compute inside memory. Analog in-memory computing and neuromorphic chips keep data where it is stored and compute there.[3][8]
- 2.5D and 3D integration. TSMC’s CoWoS put more than three reticles of silicon and eight HBM stacks in one Blackwell package as of 2024, and TSMC projected wafer-scale packages with room for more than 60 HBM stacks by 2027.[5][9]
- Denser vertical links. Production hybrid bonding reached about 9 micrometre pitch by 2024, with imec demonstrating 400 nanometres wafer-to-wafer.[10][11]
- Optical I/O. Optical chiplets such as Ayar Labs’ 8 Tbps TeraPHY target chip-to-chip bandwidth beyond what copper can deliver at acceptable power.[12][13]
- Near- and in-memory compute. IBM’s NorthPole keeps all memory on-chip and reported 72.7x the energy efficiency of a comparable GPU on a 3-billion-parameter model.[14][15]
The limits of each fix
None of these fixes is free. Advanced packaging capacity has been a supply bottleneck, with TSMC’s CoWoS demand reported to exceed supply through 2026.[16] Hybrid bonding requires nanometre-level flatness.[17] Analog compute has to cope with device non-idealities and needs models adapted to the hardware.[18] The following pages look at each approach in turn.[2]
Questions readers ask
What is the memory wall?
The widening gap between processor compute speed and the bandwidth of memory and interconnects that feed it data.[1]
Why does it matter most for AI inference?
Analysis of AI workloads finds memory bandwidth, not compute, is the primary bottleneck, particularly when serving models.[2]
How does in-memory computing help?
It computes where the weights are stored, avoiding the time and energy spent moving weights to a processor.[3]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
A 2024 analysis found peak server hardware FLOPS grew about 3.0x every two years, while DRAM bandwidth grew about 1.6x and interconnect bandwidth about 1.4x over the same period. confirmedas of 2024-03-21
- AI and Memory Wall · arXiv (published in IEEE Micro) · 2024-03-21 · Abstract (retrieved 2026-10-10)
- [2]
The same analysis argues that memory bandwidth, rather than compute, has become the primary bottleneck for AI workloads, especially when serving models. confirmedas of 2024-03-21
- AI and Memory Wall · arXiv (published in IEEE Micro) · 2024-03-21 · Abstract (retrieved 2026-10-10)
- [3]
IBM researchers describe the cost of conventional AI inference as the time and energy spent moving model weights between memory and processors. confirmedas of 2025-01-29
- Analog in-memory computing could power tomorrow's AI models · IBM Research · 2025-01-29 (retrieved 2026-10-10)
- [4]
IEEE Spectrum reported that pluggable optical transceivers consume about 10% of total GPU compute power in AI data centers, more than half of it to drive lasers. reportedas of 2025-03-25
- Nvidia Unveils Game-Changing Optical Network Switch · IEEE Spectrum · 2025-03-25 (retrieved 2026-10-10)
- [5]
As of 2024, NVIDIA's Blackwell GPU used TSMC's CoWoS packaging to combine more than three reticles' worth of silicon with eight HBM memory stacks. confirmedas of 2024-04-30
- Expect a Wave of Wafer-Scale Computers · IEEE Spectrum · 2024-04-30 (retrieved 2026-10-10)
- [6]
Hybrid bonding joins two chips with dense, direct copper-to-copper connections instead of solder bumps; the copper pads are surrounded by insulating oxide and slightly recessed from its surface. confirmedas of 2024-08-11
- Hybrid Bonding: 3D Chip Tech to Save Moore's Law · IEEE Spectrum · 2024-08-11 (retrieved 2026-10-10)
- [7]
In co-packaged optics, silicon optical-transceiver chiplets sit beside the switch chip inside one package, while the lasers, made from non-silicon materials, stay outside, shortening electrical paths and reducing components. confirmedas of 2025-03-25
- Nvidia Unveils Game-Changing Optical Network Switch · IEEE Spectrum · 2025-03-25 (retrieved 2026-10-10)
- [8]
Hala Point runs asynchronous, event-based spiking neural networks with memory and computing integrated. confirmedas of 2024-04-17
- Intel Builds World's Largest Neuromorphic System to Enable More Sustainable AI · Intel Newsroom · 2024-04-17 (retrieved 2026-10-10)
- [9]
In 2024 TSMC projected a 2027 wafer-scale System-on-Wafer package with more than 40 reticles' worth of silicon and room for more than 60 HBM stacks. reportedas of 2024-04-30· forecast
- Expect a Wave of Wafer-Scale Computers · IEEE Spectrum · 2024-04-30 (retrieved 2026-10-10)
- [10]
As of 2024, hybrid-bonded 3D chips in production had connections about 9 micrometres apart, versus tens of micrometres for the solder microbumps they replace. confirmedas of 2024-08-11
- Hybrid Bonding: 3D Chip Tech to Save Moore's Law · IEEE Spectrum · 2024-08-11 (retrieved 2026-10-10)
- [11]
imec has demonstrated wafer-to-wafer hybrid bonding with a 400-nanometre pitch, and 2-micrometre pitch for chip-on-wafer bonding. confirmedas of 2024-08-11
- Hybrid Bonding: 3D Chip Tech to Save Moore's Law · IEEE Spectrum · 2024-08-11 (retrieved 2026-10-10)
- [12]
Ayar Labs unveiled on March 31, 2025 an 8 Tbps TeraPHY optical I/O chiplet with a UCIe electrical interface, powered by a 16-wavelength light source. confirmedas of 2025-03-31
- Ayar Labs Unveils World's First UCIe Optical Chiplet for AI Scale-Up Architectures · Ayar Labs · 2025-03-31 (retrieved 2026-10-10)
- [13]
Ayar Labs' chief executive argues optical interconnects are needed to solve power-density limits of copper in large AI systems. confirmedas of 2025-03-31
- Ayar Labs Unveils World's First UCIe Optical Chiplet for AI Scale-Up Architectures · Ayar Labs · 2025-03-31 (retrieved 2026-10-10)
- [14]
IBM's NorthPole inference chip is made on a 12nm process with 22 billion transistors in 795 square millimetres and keeps its memory on-chip. confirmedas of 2024-09-26
- Breakthrough low-latency, high-energy-efficiency LLM inference performance using NorthPole · IBM Research · 2024-09-26 (retrieved 2026-10-10)
- [15]
IBM reported in September 2024 that NorthPole ran a 3-billion-parameter language model at under 1 millisecond per token with 72.7 times the energy efficiency of the next-lowest-latency GPU. confirmedas of 2024-09-26
- Breakthrough low-latency, high-energy-efficiency LLM inference performance using NorthPole · IBM Research · 2024-09-26 (retrieved 2026-10-10)
- [16]
Institutional investors cited by Taiwan's Economic Daily News expected the CoWoS supply-demand gap to narrow from about 20% to about 10% by the end of 2026. reportedas of 2026-06-15· forecast
- TSMC CoWoS Supply-Demand Gap Reportedly Seen Narrowing from 20% to 10% by End-2026 as Capacity Expands · TrendForce · 2026-06-15 (retrieved 2026-10-10)
- [17]
Hybrid bonding requires extreme surface flatness; engineers polish away the last few nanometres of oxide because slight bulges or warping can break dense connections. confirmedas of 2024-08-11
- Hybrid Bonding: 3D Chip Tech to Save Moore's Law · IEEE Spectrum · 2024-08-11 (retrieved 2026-10-10)
- [18]
The analog attention design could not run pre-trained models directly because of gain-cell non-idealities, so it needed an adaptation algorithm, and its gain-cell arrays were limited to 64x64 to contain voltage (IR) drop. confirmedas of 2025-09-08
- Analog in-memory computing attention mechanism for fast and energy-efficient large language models · Nature Computational Science (via PubMed Central) · 2025-09-08 · Abstract (retrieved 2026-10-10)
- Analog in-memory computing attention mechanism for fast and energy-efficient large language models · Nature Computational Science (via PubMed Central) · 2025-09-08 · Sub-tiling to scale attention dimensions (continues: × 64) (retrieved 2026-10-10)
Revision history (2)
- Page created.
- Corrected the description of analog compute limits to match the source.
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"The memory wall: why moving data limits AI hardware." ContentLora, updated Oct 10, 2026. https://contentlora.com/explain/memory-wall
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerNext-gen computing hardware in 2026: a crash courseA crash course on computing beyond conventional chip scaling: GAA transistors, backside power, chiplets, optical links, neuromorphic and analog compute.
- ExplainerWhy chip scaling slowed, and what replaced itHow the end of Dennard scaling around 2005 created a power wall, why multi-core and accelerators followed, and why new hardware ideas now matter.
- ExplainerHow optical interconnects and silicon photonics workWhy data centers are replacing copper and pluggable transceivers with light, and how silicon photonics, optical chiplets and co-packaged optics work.
- WikiAnalog in-memory computingAnalog in-memory computing does neural-network maths inside memory arrays so weights never move. How it works, IBM's phase-change work, 2025 results, limits.
- WikiCo-packaged optics (CPO)Co-packaged optics puts optical engines inside the switch or processor package. Products from NVIDIA, Broadcom and startups, claimed savings, open questions.
- WikiNeuromorphic computingNeuromorphic computing builds chips inspired by the brain's spiking neurons. Intel's Hala Point, IBM's NorthPole, the 2025 Nature roadmap and open questions.