Positron challenges HBM: how chips built on ordinary LPDDR5X memory could change the economics of AI inference

Edited by: Svitlana Velhush

On 10 September 2026, Positron AI announced it had raised 875 million dollars in a Series C/C-1 round at a post-money valuation of 5 billion dollars. The Reno-based company is developing chips and systems for inference that, instead of the usual high-bandwidth HBM memory, use ordinary LPDDR5X from smartphones.

Positron's key technical thesis is that traditional accelerators use less than 30 % of HBM's bandwidth, whereas their architecture can harness more than 90 % of LPDDR5X bandwidth. According to the company, this compensates for the difference in peak performance and reduces dependence on scarce HBM supply chains and advanced CoWoS packaging.

The next chip, Asimov, is planned to move to TSMC N3P at the end of 2026, with production in the second half of 2027. Each die will be able to carry from 288 GB to 2,3 TB of LPDDR5X. The Titan system will combine four to eight such chips and targets models of over 16 trillion parameters with a context of more than 10 million tokens in a single node. Already more than 50 racks of the first Atlas platform are deployed in Oracle Cloud Infrastructure.

Architecturally, Positron is betting on a systolic array with tightly integrated memory, which makes it possible to minimise data movement. Unlike classic GPUs, where memory and computation are separated, here the emphasis shifts to capacity and real bandwidth utilisation. This approach is especially relevant for long-context and large models, where the bottleneck is often memory rather than peak FLOPS.

Methodologically, Positron's claims so far rest on internal measurements and comparisons with typical figures for HBM systems. There are as yet no independent benchmarks or detailed descriptions of testing protocols in open sources. The company emphasises advantages in tokens per dollar and per watt, but without public third-party tests it is hard to assess the real gain in real workloads.

In the inference-chip landscape, Positron sets itself against the dominant solutions from Nvidia and AMD, which are tightly tied to HBM. In parallel, other startups and hyperscalers are experimenting with alternative approaches — from optical interconnect to custom ASICs. The choice of LPDDR5X brings Positron closer to the ideas of some Chinese developers seeking to circumvent sanctions restrictions on advanced memory, but in this case the emphasis is on utilisation and cost, not only on availability.

If the architecture really delivers the claimed bandwidth utilisation, it could substantially lower the TCO of inference for ultra-large models and long contexts. It would become possible to scale clusters faster without fighting for limited volumes of HBM and CoWoS. At the same time, success depends on how effectively the software and compiler can load the systolic array in real-world scenarios.

The question remains open whether independent tests will confirm Positron's advantages after the Asimov tape-out and how the system will behave when scaled to thousands of nodes. The next interesting publications in this field will probably compare the real energy efficiency and latency of Titan with current HBM solutions on identical models.

Positron shows that rethinking priorities — capacity and memory utilisation instead of peak bandwidth — could become a real alternative in the race for inference hardware.

1 Views

Sources

  • Positron Raises $875M for Memory-First Inference Chips

Comments

Read more articles on this topic:

Did you find an error or inaccuracy?We will consider your comments as soon as possible.