femtoAI Scales Inference Tech as Dual Sparsity Gains Traction
With 200,000 chips already in the field, San Bruno-based femtoAI has reported five-fold growth in the first half of 2026. The company is capitalizing on a shift toward on-device intelligence, moving beyond raw compute to optimize how AI models consume memory and power at the silicon level.

The company’s expansion is anchored by strategic partnerships with industry incumbents, including long-term collaborations with Samsung and ABOV Semiconductor. These alliances span a wide range of hardware, from home appliances to premium audio equipment with Marshall. Beyond consumer electronics, the platform is gaining a foothold in specialized sectors such as robotics, smart glass, and industrial fault monitoring.
At the core of this growth is the Sparse Processing Unit (SPU), which utilizes a concept the company calls dual sparsity. By applying this technique simultaneously to the hardware architecture and software stack, femtoAI claims it can achieve a 100-fold increase in energy efficiency and a 10-fold reduction in memory requirements. This approach has attracted attention from analysts like Swetha Srinivasan, who noted that the company’s focus on full-stack alignment—sparsifying both the model and the silicon—represents a departure from traditional industry reliance on mere compute scaling. As the company prepares its next-generation SPU, it continues to focus on developer adoption, with nearly half of its customer base actively building custom models through the firm's dedicated portal.
Comments (0)
No comments yet. Be the first!