Skip to content
ContentLora

    Tip: press / anywhere to search.

    Explainer

    The memory wall: why moving data limits AI hardware

    The memory wall is the gap between how fast processors can calculate and how fast data can reach them. One 2024 analysis found server compute growing about 3.0x every two years against 1.6x for DRAM bandwidth, making memory the main bottleneck for AI, especially when serving models.[1][2]

    Editor reviewedUpdated Next-gen computing hardwareSemiconductorsComputing

    The problem in one picture

    Imagine a chef who can chop vegetables incredibly fast, but whose ingredients arrive through a narrow hatch. The chef spends most of the time waiting. Modern AI chips are that chef: they can do arithmetic much faster than memory can deliver the numbers. Researchers measured that peak server computing power grew about three times every two years, while memory bandwidth grew only about 1.6 times.[1]

    Gholami and colleagues, writing in IEEE Micro, report peak server FLOPS scaling at 3.0x per two years, against 1.6x for DRAM bandwidth and 1.4x for interconnect bandwidth.[1] For decoder-style language models, whose generation step reads every weight for each token, they argue memory rather than compute is now the primary bottleneck, particularly in serving.[2]

    Why moving data is expensive

    Every time a model’s weights travel from memory to the processor, it costs time and energy. IBM researchers put it plainly: inference on traditional architectures pays that cost every time weights move.[3] The same applies between chips: in AI data centers, the pluggable optical transceivers that link servers were reported to use around 10% of total GPU compute power.[4]

    Four ways hardware fights the wall

    1. Put memory closer. Advanced packaging places stacks of high-bandwidth memory right next to the processor. NVIDIA’s Blackwell GPU, for example, sits in one package with eight memory stacks.[5]
    2. Stack chips directly. Hybrid bonding joins chips with very dense copper connections, so data travels microns rather than millimetres.[6]
    3. Use light. Co-packaged optics moves data between chips as light, with less power than electrical links.[7]
    4. Compute inside memory. Analog in-memory computing and neuromorphic chips keep data where it is stored and compute there.[3][8]
    • 2.5D and 3D integration. TSMC’s CoWoS put more than three reticles of silicon and eight HBM stacks in one Blackwell package as of 2024, and TSMC projected wafer-scale packages with room for more than 60 HBM stacks by 2027.[5][9]
    • Denser vertical links. Production hybrid bonding reached about 9 micrometre pitch by 2024, with imec demonstrating 400 nanometres wafer-to-wafer.[10][11]
    • Optical I/O. Optical chiplets such as Ayar Labs’ 8 Tbps TeraPHY target chip-to-chip bandwidth beyond what copper can deliver at acceptable power.[12][13]
    • Near- and in-memory compute. IBM’s NorthPole keeps all memory on-chip and reported 72.7x the energy efficiency of a comparable GPU on a 3-billion-parameter model.[14][15]

    The limits of each fix

    None of these fixes is free. Advanced packaging capacity has been a supply bottleneck, with TSMC’s CoWoS demand reported to exceed supply through 2026.[16] Hybrid bonding requires nanometre-level flatness.[17] Analog compute has to cope with device non-idealities and needs models adapted to the hardware.[18] The following pages look at each approach in turn.[2]

    Questions readers ask

    What is the memory wall?

    The widening gap between processor compute speed and the bandwidth of memory and interconnects that feed it data.[1]

    Why does it matter most for AI inference?

    Analysis of AI workloads finds memory bandwidth, not compute, is the primary bottleneck, particularly when serving models.[2]

    How does in-memory computing help?

    It computes where the weights are stored, avoiding the time and energy spent moving weights to a processor.[3]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      A 2024 analysis found peak server hardware FLOPS grew about 3.0x every two years, while DRAM bandwidth grew about 1.6x and interconnect bandwidth about 1.4x over the same period. confirmedas of 2024-03-21

      • AI and Memory Wall · arXiv (published in IEEE Micro) · 2024-03-21 · Abstract (retrieved 2026-10-10)
    2. [2]

      The same analysis argues that memory bandwidth, rather than compute, has become the primary bottleneck for AI workloads, especially when serving models. confirmedas of 2024-03-21

      • AI and Memory Wall · arXiv (published in IEEE Micro) · 2024-03-21 · Abstract (retrieved 2026-10-10)
    3. [3]

      IBM researchers describe the cost of conventional AI inference as the time and energy spent moving model weights between memory and processors. confirmedas of 2025-01-29

    4. [4]

      IEEE Spectrum reported that pluggable optical transceivers consume about 10% of total GPU compute power in AI data centers, more than half of it to drive lasers. reportedas of 2025-03-25

    5. [5]

      As of 2024, NVIDIA's Blackwell GPU used TSMC's CoWoS packaging to combine more than three reticles' worth of silicon with eight HBM memory stacks. confirmedas of 2024-04-30

    6. [6]

      Hybrid bonding joins two chips with dense, direct copper-to-copper connections instead of solder bumps; the copper pads are surrounded by insulating oxide and slightly recessed from its surface. confirmedas of 2024-08-11

    7. [7]

      In co-packaged optics, silicon optical-transceiver chiplets sit beside the switch chip inside one package, while the lasers, made from non-silicon materials, stay outside, shortening electrical paths and reducing components. confirmedas of 2025-03-25

    8. [8]

      Hala Point runs asynchronous, event-based spiking neural networks with memory and computing integrated. confirmedas of 2024-04-17

    9. [9]

      In 2024 TSMC projected a 2027 wafer-scale System-on-Wafer package with more than 40 reticles' worth of silicon and room for more than 60 HBM stacks. reportedas of 2024-04-30· forecast

    10. [10]

      As of 2024, hybrid-bonded 3D chips in production had connections about 9 micrometres apart, versus tens of micrometres for the solder microbumps they replace. confirmedas of 2024-08-11

    11. [11]

      imec has demonstrated wafer-to-wafer hybrid bonding with a 400-nanometre pitch, and 2-micrometre pitch for chip-on-wafer bonding. confirmedas of 2024-08-11

    12. [12]

      Ayar Labs unveiled on March 31, 2025 an 8 Tbps TeraPHY optical I/O chiplet with a UCIe electrical interface, powered by a 16-wavelength light source. confirmedas of 2025-03-31

    13. [13]

      Ayar Labs' chief executive argues optical interconnects are needed to solve power-density limits of copper in large AI systems. confirmedas of 2025-03-31

    14. [14]

      IBM's NorthPole inference chip is made on a 12nm process with 22 billion transistors in 795 square millimetres and keeps its memory on-chip. confirmedas of 2024-09-26

    15. [15]

      IBM reported in September 2024 that NorthPole ran a 3-billion-parameter language model at under 1 millisecond per token with 72.7 times the energy efficiency of the next-lowest-latency GPU. confirmedas of 2024-09-26

    16. [16]

      Institutional investors cited by Taiwan's Economic Daily News expected the CoWoS supply-demand gap to narrow from about 20% to about 10% by the end of 2026. reportedas of 2026-06-15· forecast

    17. [17]

      Hybrid bonding requires extreme surface flatness; engineers polish away the last few nanometres of oxide because slight bulges or warping can break dense connections. confirmedas of 2024-08-11

    18. [18]

      The analog attention design could not run pre-trained models directly because of gain-cell non-idealities, so it needed an adaptation algorithm, and its gain-cell arrays were limited to 64x64 to contain voltage (IR) drop. confirmedas of 2025-09-08

    Revision history (2)
    1. Page created.
    2. Corrected the description of analog compute limits to match the source.

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "The memory wall: why moving data limits AI hardware." ContentLora, updated Oct 10, 2026. https://contentlora.com/explain/memory-wall

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.