Specialized Hardware for LLM Inference

Etched has reached a $10.3 billion valuation following a $300 million Series C round, signaling growing investor confidence in hardware optimized specifically for AI inference. While initial skepticism centered on the viability of chips designed for transformer architectures, the company has pivoted to a broader approach, claiming its systems can run various architectures, including Mixture of Experts (MoE) and state-space models like Mamba.

Optimizing the Inference Pipeline

Etched’s technical strategy focuses on decoupling the two primary phases of LLM inference to maximize speed and efficiency:

  • Prefill Phase (Compute-Intensive): Etched developed a chip that operates at a significantly lower voltage than standard AI hardware. This reduction in voltage lowers heat generation, allowing for a higher density of transistors, which accelerates the processing of prompts and context.
  • Decode Phase (Memory-Intensive): To handle the token generation phase, the company engineered a "cluster scale memory" interconnect technology. This allows multiple chips to share a memory pool with low latency, addressing the memory-bandwidth bottlenecks common in traditional GPU-based inference.

By selling these as full rack systems rather than standalone chips, Etched aims to provide a turnkey solution for high-speed, cost-effective inference. The company reports having already booked $1 billion in orders, with systems currently undergoing testing by major AI companies.