newsroom
AI Hardware & Chip Architecture · 1:1 matrix-vector ratio architecture purpose-built for inference attention bottlenecks
now playing · AI Hardware & Chip Architecture
1:1 matrix-vector ratio architecture purpose-built for inference attention bottlenecks
Positron's systolic array treats matrix-vector math as a first-class citizen, achieving a 1:1 ratio of matrix-matrix to matrix-vector throughput versus 32:1 on Blackwell — directly targetin…
Inference cost reduction via 93% memory bandwidth utilization and high memory capacity
Positron's architecture sustains 93% of theoretical peak memory bandwidth on FPGAs, and the Azimov chip extends this with massive LPDDR capacity — enabling high token throughput at low batc…
LPDDR enables 2.3TB/chip for inference, bypassing HBM capacity and supply constraints
Positron's Azimov chip uses commodity LPDDR5X to achieve 2.3TB memory per chip at 400W — 6-8x the capacity of HBM-based GPUs like B200 — allowing single-server deployment of 16T parameter m…
Organic substrates avoid advanced packaging supply chain choke point for AI compute scale
By using regular organic substrates instead of CoWoS/advanced packaging, Positron sidesteps the same constrained supply chain that Nvidia, AMD, and Google compete for — a strategic advantag…