14 SEP 2026 — Positron AI has raised $875 million at a $5 billion valuation on a proposition that sounds like a downgrade: build inference silicon around ordinary laptop memory instead of the exotic stuff everyone is fighting over.
The company's Asimov chip uses LPDDR5X rather than high-bandwidth memory. HBM is the component the entire accelerator market is short of, and the one whose price has been climbing all year. Positron's argument is that for inference specifically, you do not need it.
The round
The financing announced on 10 September splits into a $375 million Series C and a $500 million Series C-1, closing at a $5 billion post-money valuation against a $3.5 billion pre-money. The company was valued at just over $1 billion in February.
It is co-led by NEA, Atreides Management, Valor Equity Partners, Andra Capital and SemiAnalysis Capital, with Forest Baskett, Gavin Baker, Thomas Jermoluk and Dylan Patel joining the board. Other backers named include DFJ Growth, the Qatar Investment Authority, Hudson River Trading, Cisco Investments and Naver Ventures.
Hudson River Trading and Jump Trading stand out on that list. Both are high-frequency trading firms, and Jump is also named as a production customer, which says something about where latency-sensitive inference is already paying for itself.
The bet on commodity memory
Inference is memory-bound in a way training is not. Generating a token requires reading the model's weights, and for a large model that read is the work; the arithmetic between reads is comparatively cheap. What a chip needs is the ability to move weights, and ideally to hold enough of them that it does not have to move them across a network at all.
HBM solves that by stacking memory dies beside the compute on an interposer, which delivers enormous bandwidth and brings four problems with it. The parts are expensive. Supply is short. Power draw is high. And the advanced packaging the whole arrangement depends on is itself one of the industry's tightest bottlenecks.
Positron's claim is that a GPU wastes most of what HBM provides. The company says Asimov achieves more than 90 per cent of available memory bandwidth utilisation on inference workloads, against under 30 per cent for a typical GPU. If that holds, a chip with less raw bandwidth and far better utilisation lands in the same place for a fraction of the cost and power.
That is the whole thesis and it is a claim, not a measurement anyone outside the company has verified.
What the money buys and when
Asimov tapes out on TSMC's N3P at the end of 2026 with production in the second half of 2027. Memory per chip runs from 288 GB to 2,304 GB. The Titan system pairs four to eight of them, targeting models beyond 16 trillion parameters and context windows above 10 million tokens.
Read those dates against the valuation. The product this round funds is roughly a year from production and about two from meaningful revenue, and it is being priced at five times what the company was worth in February. That is a bet on the memory constraint persisting, not on the current generation.
There is existing revenue to point at. More than 50 racks of the first-generation Atlas system are deployed at Oracle Cloud Infrastructure, with Parasail, Jump Trading and i3d.net named as production customers. Fifty racks is a real deployment and a small one against what a hyperscaler operates.
What has to hold for this to work
Start with the utilisation claim, which has to survive contact with somebody else's workload. Ninety per cent of available bandwidth on a benchmark the vendor chose is a different assertion from ninety per cent on a customer's actual serving mix, with its irregular batch sizes and its long tail of small requests.
Then there is the timing. A chip whose advantage is avoiding a constrained component loses part of its case if the constraint eases, and capacity is being added across the industry. Our reporting on Chinese DRAM has made the mirror-image point, which is that commodity DDR5 and LPDDR5 capacity is expanding fastest precisely because the incumbents pivoted away from it towards HBM.
The software is the oldest of the three obstacles and the one that has decided this contest before. Every challenger to Nvidia in the past decade has been beaten by CUDA rather than by silicon, and nothing in this announcement addresses what it costs a customer to move a serving stack.
What it says about the market
The interesting signal is not whether Positron wins. It is that $875 million was available for an argument that the industry's most expensive component is the wrong one for more than half of what accelerators are now used for.
Inference is where the volume has gone. Training runs are concentrated in a few organisations; serving happens everywhere, continuously, and its economics are measured in tokens per dollar and tokens per watt rather than in time to convergence. A chip designed for the second problem rather than adapted from the first is a reasonable thing to fund.
Whether it is a $5 billion thing to fund a year before tape-out is a separate question, and the February-to-September markup suggests the answer is being set by scarcity of exposure rather than by evidence.
The signals worth tracking
Tape-out at the end of 2026 is the first checkpoint nobody can argue with. Silicon either comes back working or it does not.
A named hyperscaler beyond Oracle would be the second. Fifty racks establishes that the architecture runs, while a second large buyer would establish that the economics convince somebody with alternatives.
Memory pricing decides the rest. If HBM supply loosens materially through 2027, the case for this chip narrows without anybody at Positron doing anything wrong.