Four Walls Constrain AI Today. We’re Breaking Them All.

Jitendra Mohan, Chief Executive Officer, Co-Founder
breaking through four walls to scaling AI

Just 16 months ago, I wrote about AI outgrowing the server. While this thesis holds, the workloads have changed: The challenges then were defined by the size of training clusters, but AI now has entered the inference era in earnest.

The hottest topic among data center operators is no longer the “AI model” but the “AI factory” – and the winning approach must deliver not just peak performance and reliability, but economic efficiency that becomes competitive advantage.

Today, four walls constrain tokens per second and, more importantly, tokens per dollar:1

  • The Wafer Wall, Compute’s Hidden Ceiling: This is the real ceiling on how much compute AI can get. A gigawatt of AI compute needs roughly 55,000 logic wafers and 170,000 memory wafers.2 By 2030, AI wafer demand is expected to have a 1-4 million advanced node wafer shortfall.3
  • The Power Wall, a Long-Term Challenge with Immediate Implications: The U.S. electrical grid is expanding linearly while AI demand expands exponentially. Data center developers are facing a 34% net power shortfall through 2028 (roughly 32 GW),4 with 70% of standard grid connection requests getting withdrawn due to years-long delays.5
  • The Memory Wall, the Industry’s Most Talked About Bottleneck: This bottleneck has driven conversations at every industry show this year. Last month at AI Infra Summit 2026, analyst Jim Handy reported a 250% Y/Y jump in DRAM prices as inference and agentic workloads push KV cache demand ever higher.
  • The Connectivity Wall, Where Utilization Is Won or Lost: In a large GPU cluster, the interconnect determines how much of the compute actually does useful work. Each new generation of GPUs piles on PFLOPS and memory bandwidth, which only raises the bar for the fabric that feeds them: Move more data, and move it faster, or leave million-dollar accelerators stalled and waiting on data instead of generating tokens.

Three of these walls are physically fixed: No one can manufacture, generate, or buy past wafer supply, power, or memory scarcity on any timeline that matters this cycle. The fourth – the connectivity wall – is the lever. Improving connectivity translates directly to improving utilization of the AI infrastructure that already exists, keeping scarce GPUs saturated instead of idle and pushing back on all three fixed walls at once.

Breaking Through Starts With Utilization

With every other front constrained, each point of efficiency here is effectively new capacity: It extracts more useful work from the same silicon, power, and budget.

Astera Labs’ scale-up switching and memory-expansion products create just these efficiencies, increasing token throughput possible from existing infrastructure without treating additional accelerators as the only path to scale.

At OCP 2026, Astera Labs will showcase multiple live demos featuring the Scorpio™ X-Series 320L Smart Fabric Switch, the industry’s highest-radix PCIe® scale-up switch. Scorpio’s Hypercast™ and In-Network Compute features deliver measurable gains in GPU utilization, boosting collective operations by up to 2x by creating optimized pathways and cutting communication overhead across the scale-up network.6

The recently released Leo X-Series Smart Memory Controller pairs with Scorpio Smart Fabric Switches to create a fabric-attached memory tier for KV cache offload. Early data shows a 22% increase in GPU utilization – 22% more tokens per second – from this dedicated offload tier alone. A live demo of Leo X-Series will run in the Astera Labs booth.

The second-generation Leo E-Series and P-Series Smart Memory Controllers, released alongside the X-Series, directly ease the memory wall with CPU-attached CXL® memory expansion, memory sharing and pooling, and even reuse of recycled DDR4 memory while DDR5 supply remains extremely limited.

Because these gains keep GPUs generating tokens instead of idling, they reach beyond the power, wafer and memory walls. More tokens from every saturated GPU also means more tokens per watt, and more tokens per dollar.

Running AI Infrastructure Smoothly With COSMOS

Hardware gains only translate into sustained throughput if the infrastructure runs smoothly at scale. COSMOS, Astera Labs’ software framework, lets customers customize Astera products, optimize system performance, and bring up complex systems through deep debug capabilities. It also keeps all that infrastructure running smoothly with rich, fleet-scale telemetry.

In scale-up fabrics featuring Scorpio X-Series Smart Fabric Switches, COSMOS helps maximize accelerator utilization by configuring and managing Scorpio’s Hypercast and In-Network Compute capabilities to accelerate GPU data distribution and offload collective communication operations.

COSMOS unifies the entire Astera Labs portfolio and has been adopted by multiple customers. It will feature prominently in every demonstration exhibited at OCP 2026.

Cutting-Edge AI Silicon Means Lower Power

Signal integrity has always been the backbone of the Astera Labs portfolio, and we are committed to staying at the leading edge as protocols, data rates, and architectures evolve. In a power-constrained environment, maintaining that cutting-edge position means offering lower-power alternatives within the connectivity layer.

The recently released Taurus Smart Signal Conditioners provide footprint-compatible optionality between Smart Retimers and Smart Redrivers for Ethernet, ESUN, and UALink applications. Where a channel can be compensated with signal amplification rather than full signal regeneration, system designers can choose the lower-power redriver option.

On the physical layer, Astera Labs will show the latest advances along both our copper and optical roadmaps. Copper supports dense links within the rack, while optical connectivity extends PCIe fabrics beyond the box and rack, enabling larger, more modular AI clusters. The result is a set of building blocks for the mix of copper and optics that each system requires rather than a single prescribed technology.

Come to our booth for our expert perspective on where copper and optics can contribute to your scale-up and scale-out architectures.

Partnering to Deliver Efficient, Performant, and Profitable AI Your Way

Astera Labs treats the OCP Global Summit as our flagship annual event because open ecosystems accelerate innovation through parallel progress – a conviction that deepens with every year spent on the boards of standards bodies like UALink and CXL. Thanks go to our partners across the ecosystem who have contributed to this year’s technology demonstrations.

Open standards are also core to our AI Your Way philosophy, our commitment to meeting customers’ needs with interconnect purpose-built for their workloads and architectures regardless of accelerator, protocol, or physical media. Come by the Astera Labs booth at OCP 2026 to see how we’re breaking down the walls standing between our customers and high-performance AI at scale.

References:

  1. Adapted from a framework developed by Dan Rabinovitsj.
  2. Sunil Venkataram, “EUV Lithography as the Binding Constraint on AI Scaling,” accessed October 2026, https://sunilvenkataram.com/garden/euv-lithography-as-the-binding-constraint-on-ai-scaling/.
  3. Ondrej Burkacky et al, “Generative AI: The next S-curve for the semiconductor industry?” accessed October 2026, https://www.mckinsey.com/industries/semiconductors/our-insights/generative-ai-the-next-s-curve-for-the-semiconductor-industry.
  4. Regan Caraher, “Here’s how scarce electricity could hamper the AI investing boom,” accessed October 2026, “https://privatebank.jpmorgan.com/nam/en/insights/markets-and-investing/heres-how-scarce-electricity-could-hamper-the-ai-investing-boom.
  5. Michael Kern, “Why the AI Boom Is About to Break the U.S. Power Grid,” accessed October 2026, https://finance.yahoo.com/energy/articles/why-ai-boom-break-u-000000959.html?guccounter=1.
  6. Based on Astera Labs internal analysis. “Up to 2x” reflects at least 50% reduction in AllReduce collective operation latency compared to traditional Ring AllReduce, achieved by offloading ReduceScatter and AllGather operations to Scorpio’s In-Network Compute and Hypercast™ engines, reducing per-GPU transmit and receive operations from N−1 to 1 per phase. Actual results may vary based on workload, system configuration, and cluster size.

About Jitendra Mohan, Chief Executive Officer, Co-Founder

Jitendra co-founded Astera Labs in 2017 with a vision to remove performance bottlenecks in data-centric systems. Jitendra has more than two decades of engineering and general management experience in identifying and solving complex technical problems in datacenter and server markets. Prior to Astera Labs, he worked as the General Manager for Texas Instruments’ High Speed Interface Business and Clocking Business. Earlier at National Semiconductor Corp, Jitendra led engineering teams in various technical leadership roles. Jitendra holds a BSEE from IIT-Bombay, an MSEE from Stanford University and over 35 granted patents. In addition to work, Jitendra enjoys outdoor activities and reading about the origins of the Universe.

Share:

Related Articles