Skip to content
ContentLora

    Tip: press / anywhere to search.

    technology

    Analog in-memory computing

    Also known as AIMC, compute-in-memory, CIM, in-memory computing, analog AI

    Analog in-memory computing performs the multiply-and-add maths of neural networks inside memory arrays, avoiding the time and energy of moving weights to a processor.[1] A 2025 Nature Computational Science design reported up to four orders of magnitude lower attention energy than GPUs, but at low precision and with hardware-specific model adaptation.[2][3]

    Editor reviewedUpdated Next-gen computing hardwareComputingSemiconductors
    Key facts

    Analog in-memory computing (AIMC) is a way of running neural networks in which the memory that stores a model’s weights also does the arithmetic. It targets the memory wall: in conventional hardware, inference costs time and energy every time weights move to the processor.[1][4]

    How it works

    A neural network is mostly a huge number of multiplications and additions using stored numbers called weights. In an analog in-memory chip, each weight is stored as a physical property, such as how easily a tiny memory cell conducts electricity. Send a voltage into a grid of these cells and the currents that come out are already the answer to many multiplications at once, so the weights never have to travel.[5][1]

    Weights are stored in the memory array itself and dot products are computed in the analog domain, so a matrix-vector multiply happens where the data sits.[6] IBM’s work uses phase-change memory, where chalcogenide glass switches between crystalline and amorphous states to set conductance.[5] Converting results back to digital costs energy, and the 2025 gain-cell attention design used charge-to-pulse circuits to avoid power-hungry analog-to-digital converters.[6]

    Key results

    IBM (January 2025). IBM Research described three studies applying AIMC to transformer models: a 3D mixture-of-experts architecture featured in Nature Computational Science, a phase-change-memory edge accelerator presented at IEDM, and a transformer run on an analog chip, published in Nature Machine Intelligence.[7]

    • Mixture of experts. In simulation, mapping each expert of a mixture-of-experts model onto its own tier of a 3D analog chip gave higher throughput, area efficiency and energy efficiency than commercially available GPUs running the same model.[8]
    • Edge devices. The edge study proposed a neural processing unit mixing analog phase-change-memory accelerators with digital nodes; on MobileBERT it beat an existing low-cost accelerator on IBM’s own throughput benchmark.[9]
    • Transformers on a real chip. Running every static-weight matrix operation of a transformer on an analog chip stayed within 2% of floating-point accuracy on the Long Range Arena benchmark.[10]

    Analog attention (September 2025). A Nature Computational Science paper described an in-memory design for transformer attention that stores the key-value cache in charge-based gain cells.[6] It reported attention latency and energy up to two and four orders of magnitude lower than GPUs, respectively.[2] With an adaptation algorithm, the model reached accuracy similar to a pre-trained GPT-2 without training from scratch.[11]

    Limits

    The same 2025 paper illustrates the trade-offs. Gain-cell non-idealities meant pre-trained models could not be mapped directly, so the design needed a hardware-aware adaptation algorithm, and its arrays were limited to 64x64 to contain voltage (IR) drop.[3] IBM’s January 2025 results were partly simulations and benchmarks of IBM’s own choosing.[8][9] The large energy gains so far are for specific operations measured in research settings rather than whole products.[2]

    Relation to other approaches

    AIMC overlaps with neuromorphic computing, which also co-locates memory and compute, and with digital near-memory designs such as IBM’s NorthPole, which keeps all memory on-chip.[12][13] NorthPole’s designers make the same diagnosis, noting that processor efficiency has outpaced memory bandwidth.[14] Whether analog approaches can reach the accuracy, programmability and manufacturing scale needed to compete with GPUs such as NVIDIA‘s is one of the open questions in the post-Moore hardware debate.[3]

    Questions readers ask

    Why compute inside memory?

    Conventional inference spends time and energy every time weights move from memory to the processor; in-memory computing avoids that.[1]

    What memory does IBM use?

    Phase-change memory, which stores values as the conductivity of chalcogenide glass switched between crystalline and amorphous states.[5]

    What are the main drawbacks?

    In a 2025 attention design, precision was limited to 3-5 bits, arrays to 64x64 because of voltage drop, and models had to be adapted to the hardware.[3]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      IBM researchers describe the cost of conventional AI inference as the time and energy spent moving model weights between memory and processors. confirmedas of 2025-01-29

    2. [2]

      A September 2025 Nature Computational Science paper reported an analog in-memory attention design that cut attention latency and energy by up to two and four orders of magnitude, respectively, compared with GPUs. confirmedas of 2025-09-08

    3. [3]

      The analog attention design could not run pre-trained models directly because of gain-cell non-idealities, so it needed an adaptation algorithm, and its gain-cell arrays were limited to 64x64 to contain voltage (IR) drop. confirmedas of 2025-09-08

    4. [4]

      The same analysis argues that memory bandwidth, rather than compute, has become the primary bottleneck for AI workloads, especially when serving models. confirmedas of 2024-03-21

      • AI and Memory Wall · arXiv (published in IEEE Micro) · 2024-03-21 · Abstract (retrieved 2026-10-10)
    5. [5]

      IBM's analog in-memory computing work stores neural-network weights in phase-change memory devices, whose conductivity changes as chalcogenide glass switches between crystalline and amorphous states. confirmedas of 2025-01-29

    6. [6]

      The design stores the attention key-value cache in charge-based "gain cell" memories and computes dot products in the analog domain. confirmedas of 2025-09-08

    7. [7]

      In January 2025 IBM Research described three studies on analog in-memory computing for transformer models - a 3D mixture-of-experts architecture featured in Nature Computational Science, a phase-change-memory edge accelerator presented at IEDM, and a transformer run on an analog chip, published in Nature Machine Intelligence. confirmedas of 2025-01-29

    8. [8]

      IBM's simulations found that mapping each expert of a mixture-of-experts model onto a tier of a 3D analog in-memory chip gave higher throughput, area efficiency and energy efficiency than commercially available GPUs running the same model. confirmedas of 2025-01-29

    9. [9]

      IBM's edge study proposed a neural processing unit mixing phase-change-memory analog accelerators with digital nodes; on the MobileBERT model it beat an existing low-cost accelerator on IBM's throughput benchmark. confirmedas of 2025-01-29

    10. [10]

      IBM researchers ran every static-weight matrix operation of a transformer on an analog in-memory chip and stayed within 2% of floating-point accuracy on the Long Range Arena benchmark. confirmedas of 2025-01-29

    11. [11]

      With its adaptation algorithm, the analog attention design reached accuracy similar to a pre-trained GPT-2 model without training from scratch. confirmedas of 2025-09-08

    12. [12]

      Hala Point runs asynchronous, event-based spiking neural networks with memory and computing integrated. confirmedas of 2024-04-17

    13. [13]

      IBM's NorthPole inference chip is made on a 12nm process with 22 billion transistors in 795 square millimetres and keeps its memory on-chip. confirmedas of 2024-09-26

    14. [14]

      IBM says NorthPole avoids the von Neumann bottleneck by placing memory and processing together on the chip, noting that processor efficiency has tripled every two years while memory bandwidth grows about half as fast. confirmedas of 2024-09-26

    Revision history (2)
    1. Page created.
    2. Corrected the description of IBM's three studies and the analog attention limits; added MoE, edge and transformer results and the GPT-2 adaptation result.

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "Analog in-memory computing." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/analog-in-memory-computing

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.