Why Leo X Matters In The AI Memory Stack
Astera Labs announced its expanded Leo smart memory controller family on September 15, 2026, and as of September 30, 2026 the interesting part is not only bigger CXL memory for CPUs. The standout is Leo X-Series, a fabric-attached memory controller aimed at AI inference systems where accelerators need more working memory for long-context and multi-turn workloads. Instead of treating expanded DRAM as something only a host CPU can reach, Leo X is designed to connect memory into accelerator fabrics, giving GPUs or other AI accelerators another memory tier for KV cache and agent context. Astera says the Leo X-Series works with its Scorpio X-Series Fabric Switches and can use PCIe as well as platform-specific protocols, which is important because not every accelerator fabric is built around standard CXL.mem support. (asteralabs.gcs-web.com)
Article contains affiliate links, commission may be earned.
Fabric-Attached DRAM, Not Just CPU-Attached Expansion
The core idea is simple: AI inference is increasingly limited by how much context can be kept close to the accelerator. Large language model serving, especially with longer prompts and agent-style workflows, can build large KV caches that compete with model weights and active batches for expensive HBM capacity. Leo X does not replace HBM, and Astera is not positioning it as a universal magic fix. Instead, it adds a DRAM-based tier that can sit on the scale-up fabric, potentially reducing the need to spill KV cache to slower storage or route access through CPU-attached memory. ServeTheHome reported that Astera demonstrated a Leo X memory expander with four onboard DIMM slots, and that the controller supports both DDR4 and DDR5, though Astera has not published a full Leo X capacity and speed table in the same way it has for Leo 2. (servethehome.com)
View NVIDIA DGX Spark Personal AI Desktop Supercomputer on partner website
Leo 2 Fills Out The More Traditional CXL Side
Leo X gets the attention because it reaches into accelerator fabrics, but Leo 2 E-Series and Leo 2 P-Series are the more familiar CXL memory expansion products. Astera lists Leo 2 E-Series as a CPU-attached memory expansion controller with CXL 3.2 and PCIe 6.0 x16 host connectivity, while Leo 2 P-Series adds dual-port PCIe 6.0 2×8 connectivity for pooled and shared memory designs. The published product table lists the newer Leo E-Series and P-Series controllers with 4-channel DDR5 at up to 6400 MT/s or 4-channel DDR4 at up to 3200 MT/s, using 2 DIMMs per channel for DDR5 or 3 DIMMs per channel for DDR4 in the listed configurations. Astera also lists A2000 CXL add-in card designs with 8 DDR5 RDIMM slots or 12 DDR4 RDIMM slots. (asteralabs.com)
See Intel Core Ultra 7 265K price
The DDR4 Reuse Angle Is A Big Part Of The Story
A major reason this product family is timely is cost. New AI servers are hungry for DDR5, but many data centers still have large amounts of DDR4 coming out of retired systems. Leo 2 is built to make that older memory more useful in newer infrastructure, while also supporting DDR5 where higher bandwidth and newer deployments make sense. Astera says the family includes RAS features, memory-health management, test engines, automated repair, scrubbing, error reporting, telemetry, and COSMOS software integration. That matters because reusing older DIMMs at cloud scale is not only about plugging them in; operators need a way to screen modules, monitor faults, and keep memory pools serviceable over time. (asteralabs.gcs-web.com)
Buy Crucial T705 13 GB/s NVMe SSD here
Performance Claims Need The Right Context
Astera’s headline inference numbers are attention-grabbing: the company claims Leo X can deliver up to 62% lower time to first token and up to 22% more tokens per second for KV-cache-heavy workloads. Those are vendor-provided figures, not independent review results, and ServeTheHome noted that the cited testing was against a PCIe-based NVIDIA H200 setup rather than the newest rack-scale systems. So the safer takeaway is not that every AI rack will see those gains, but that fabric-attached DRAM is becoming a serious design option for inference platforms trying to balance HBM capacity, DRAM cost, latency, and token throughput. For infrastructure teams, Leo X is worth watching because it shifts expanded memory from a CPU-side feature into the AI fabric discussion itself. (asteralabs.gcs-web.com)
See ASUS GeForce RTX 5090 price
Comments
No comments yet. Be the first to share your thoughts.