technology
Analog in-memory computing
Also known as AIMC, compute-in-memory, CIM, in-memory computing, analog AI
Analog in-memory computing performs the multiply-and-add maths of neural networks inside memory arrays, avoiding the time and energy of moving weights to a processor.[1] A 2025 Nature Computational Science design reported up to four orders of magnitude lower attention energy than GPUs, but at low precision and with hardware-specific model adaptation.[2][3]
Key facts
Analog in-memory computing (AIMC) is a way of running neural networks in which the memory that stores a model’s weights also does the arithmetic. It targets the memory wall: in conventional hardware, inference costs time and energy every time weights move to the processor.[1][4]
How it works
A neural network is mostly a huge number of multiplications and additions using stored numbers called weights. In an analog in-memory chip, each weight is stored as a physical property, such as how easily a tiny memory cell conducts electricity. Send a voltage into a grid of these cells and the currents that come out are already the answer to many multiplications at once, so the weights never have to travel.[5][1]
Weights are stored in the memory array itself and dot products are computed in the analog domain, so a matrix-vector multiply happens where the data sits.[6] IBM’s work uses phase-change memory, where chalcogenide glass switches between crystalline and amorphous states to set conductance.[5] Converting results back to digital costs energy, and the 2025 gain-cell attention design used charge-to-pulse circuits to avoid power-hungry analog-to-digital converters.[6]
Key results
IBM (January 2025). IBM Research described three studies applying AIMC to transformer models: a 3D mixture-of-experts architecture featured in Nature Computational Science, a phase-change-memory edge accelerator presented at IEDM, and a transformer run on an analog chip, published in Nature Machine Intelligence.[7]
- Mixture of experts. In simulation, mapping each expert of a mixture-of-experts model onto its own tier of a 3D analog chip gave higher throughput, area efficiency and energy efficiency than commercially available GPUs running the same model.[8]
- Edge devices. The edge study proposed a neural processing unit mixing analog phase-change-memory accelerators with digital nodes; on MobileBERT it beat an existing low-cost accelerator on IBM’s own throughput benchmark.[9]
- Transformers on a real chip. Running every static-weight matrix operation of a transformer on an analog chip stayed within 2% of floating-point accuracy on the Long Range Arena benchmark.[10]
Analog attention (September 2025). A Nature Computational Science paper described an in-memory design for transformer attention that stores the key-value cache in charge-based gain cells.[6] It reported attention latency and energy up to two and four orders of magnitude lower than GPUs, respectively.[2] With an adaptation algorithm, the model reached accuracy similar to a pre-trained GPT-2 without training from scratch.[11]
Limits
The same 2025 paper illustrates the trade-offs. Gain-cell non-idealities meant pre-trained models could not be mapped directly, so the design needed a hardware-aware adaptation algorithm, and its arrays were limited to 64x64 to contain voltage (IR) drop.[3] IBM’s January 2025 results were partly simulations and benchmarks of IBM’s own choosing.[8][9] The large energy gains so far are for specific operations measured in research settings rather than whole products.[2]
Relation to other approaches
AIMC overlaps with neuromorphic computing, which also co-locates memory and compute, and with digital near-memory designs such as IBM’s NorthPole, which keeps all memory on-chip.[12][13] NorthPole’s designers make the same diagnosis, noting that processor efficiency has outpaced memory bandwidth.[14] Whether analog approaches can reach the accuracy, programmability and manufacturing scale needed to compete with GPUs such as NVIDIA‘s is one of the open questions in the post-Moore hardware debate.[3]
Questions readers ask
Why compute inside memory?
Conventional inference spends time and energy every time weights move from memory to the processor; in-memory computing avoids that.[1]
What memory does IBM use?
Phase-change memory, which stores values as the conductivity of chalcogenide glass switched between crystalline and amorphous states.[5]
What are the main drawbacks?
In a 2025 attention design, precision was limited to 3-5 bits, arrays to 64x64 because of voltage drop, and models had to be adapted to the hardware.[3]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
IBM researchers describe the cost of conventional AI inference as the time and energy spent moving model weights between memory and processors. confirmedas of 2025-01-29
- Analog in-memory computing could power tomorrow's AI models · IBM Research · 2025-01-29 (retrieved 2026-10-10)
- [2]
A September 2025 Nature Computational Science paper reported an analog in-memory attention design that cut attention latency and energy by up to two and four orders of magnitude, respectively, compared with GPUs. confirmedas of 2025-09-08
- Analog in-memory computing attention mechanism for fast and energy-efficient large language models · Nature Computational Science (via PubMed Central) · 2025-09-08 · Abstract (retrieved 2026-10-10)
- [3]
The analog attention design could not run pre-trained models directly because of gain-cell non-idealities, so it needed an adaptation algorithm, and its gain-cell arrays were limited to 64x64 to contain voltage (IR) drop. confirmedas of 2025-09-08
- Analog in-memory computing attention mechanism for fast and energy-efficient large language models · Nature Computational Science (via PubMed Central) · 2025-09-08 · Abstract (retrieved 2026-10-10)
- Analog in-memory computing attention mechanism for fast and energy-efficient large language models · Nature Computational Science (via PubMed Central) · 2025-09-08 · Sub-tiling to scale attention dimensions (continues: × 64) (retrieved 2026-10-10)
- [4]
The same analysis argues that memory bandwidth, rather than compute, has become the primary bottleneck for AI workloads, especially when serving models. confirmedas of 2024-03-21
- AI and Memory Wall · arXiv (published in IEEE Micro) · 2024-03-21 · Abstract (retrieved 2026-10-10)
- [5]
IBM's analog in-memory computing work stores neural-network weights in phase-change memory devices, whose conductivity changes as chalcogenide glass switches between crystalline and amorphous states. confirmedas of 2025-01-29
- Analog in-memory computing could power tomorrow's AI models · IBM Research · 2025-01-29 (retrieved 2026-10-10)
- [6]
The design stores the attention key-value cache in charge-based "gain cell" memories and computes dot products in the analog domain. confirmedas of 2025-09-08
- Analog in-memory computing attention mechanism for fast and energy-efficient large language models · Nature Computational Science (via PubMed Central) · 2025-09-08 (retrieved 2026-10-10)
- [7]
In January 2025 IBM Research described three studies on analog in-memory computing for transformer models - a 3D mixture-of-experts architecture featured in Nature Computational Science, a phase-change-memory edge accelerator presented at IEDM, and a transformer run on an analog chip, published in Nature Machine Intelligence. confirmedas of 2025-01-29
- Analog in-memory computing could power tomorrow's AI models · IBM Research · 2025-01-29 (retrieved 2026-10-10)
- [8]
IBM's simulations found that mapping each expert of a mixture-of-experts model onto a tier of a 3D analog in-memory chip gave higher throughput, area efficiency and energy efficiency than commercially available GPUs running the same model. confirmedas of 2025-01-29
- Analog in-memory computing could power tomorrow's AI models · IBM Research · 2025-01-29 (retrieved 2026-10-10)
- [9]
IBM's edge study proposed a neural processing unit mixing phase-change-memory analog accelerators with digital nodes; on the MobileBERT model it beat an existing low-cost accelerator on IBM's throughput benchmark. confirmedas of 2025-01-29
- Analog in-memory computing could power tomorrow's AI models · IBM Research · 2025-01-29 (retrieved 2026-10-10)
- Analog in-memory computing could power tomorrow's AI models · IBM Research · 2025-01-29 (retrieved 2026-10-10)
- [10]
IBM researchers ran every static-weight matrix operation of a transformer on an analog in-memory chip and stayed within 2% of floating-point accuracy on the Long Range Arena benchmark. confirmedas of 2025-01-29
- Analog in-memory computing could power tomorrow's AI models · IBM Research · 2025-01-29 (retrieved 2026-10-10)
- [11]
With its adaptation algorithm, the analog attention design reached accuracy similar to a pre-trained GPT-2 model without training from scratch. confirmedas of 2025-09-08
- Analog in-memory computing attention mechanism for fast and energy-efficient large language models · Nature Computational Science (via PubMed Central) · 2025-09-08 (retrieved 2026-10-10)
- [12]
Hala Point runs asynchronous, event-based spiking neural networks with memory and computing integrated. confirmedas of 2024-04-17
- Intel Builds World's Largest Neuromorphic System to Enable More Sustainable AI · Intel Newsroom · 2024-04-17 (retrieved 2026-10-10)
- [13]
IBM's NorthPole inference chip is made on a 12nm process with 22 billion transistors in 795 square millimetres and keeps its memory on-chip. confirmedas of 2024-09-26
- Breakthrough low-latency, high-energy-efficiency LLM inference performance using NorthPole · IBM Research · 2024-09-26 (retrieved 2026-10-10)
- [14]
IBM says NorthPole avoids the von Neumann bottleneck by placing memory and processing together on the chip, noting that processor efficiency has tripled every two years while memory bandwidth grows about half as fast. confirmedas of 2024-09-26
- Breakthrough low-latency, high-energy-efficiency LLM inference performance using NorthPole · IBM Research · 2024-09-26 (retrieved 2026-10-10)
- Breakthrough low-latency, high-energy-efficiency LLM inference performance using NorthPole · IBM Research · 2024-09-26 (retrieved 2026-10-10)
Revision history (2)
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"Analog in-memory computing." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/analog-in-memory-computing
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- WikiCo-packaged optics (CPO)Co-packaged optics puts optical engines inside the switch or processor package. Products from NVIDIA, Broadcom and startups, claimed savings, open questions.
- WikiNeuromorphic computingNeuromorphic computing builds chips inspired by the brain's spiking neurons. Intel's Hala Point, IBM's NorthPole, the 2025 Nature roadmap and open questions.
- DevelopingNext-gen computing hardware tracker: 2024-2026 milestonesA live timeline of milestones in GAA transistors, backside power, chiplets, co-packaged optics, neuromorphic and analog in-memory computing.
- AnalysisBeyond Moore's law: which hardware bets will pay off?The open debates in next-gen computing hardware: incremental scaling vs radical new compute, how fast optics replaces copper, and vendor claims vs evidence.
- ExplainerNext-gen computing hardware in 2026: a crash courseA crash course on computing beyond conventional chip scaling: GAA transistors, backside power, chiplets, optical links, neuromorphic and analog compute.
- ExplainerHow optical interconnects and silicon photonics workWhy data centers are replacing copper and pluggable transceivers with light, and how silicon photonics, optical chiplets and co-packaged optics work.