RELEReleases

femtoAI Scales Inference Tech as Dual Sparsity Gains Traction

With 200,000 chips already in the field, San Bruno-based femtoAI has reported five-fold growth in the first half of 2026. The company is capitalizing on a shift toward on-device intelligence, moving beyond raw compute to optimize how AI models consume memory and power at the silicon level.

Bio & NewsJuly 22, 20261,408 reads0

The company’s expansion is anchored by strategic partnerships with industry incumbents, including long-term collaborations with Samsung and ABOV Semiconductor. These alliances span a wide range of hardware, from home appliances to premium audio equipment with Marshall. Beyond consumer electronics, the platform is gaining a foothold in specialized sectors such as robotics, smart glass, and industrial fault monitoring.

At the core of this growth is the Sparse Processing Unit (SPU), which utilizes a concept the company calls dual sparsity. By applying this technique simultaneously to the hardware architecture and software stack, femtoAI claims it can achieve a 100-fold increase in energy efficiency and a 10-fold reduction in memory requirements. This approach has attracted attention from analysts like Swetha Srinivasan, who noted that the company’s focus on full-stack alignment—sparsifying both the model and the silicon—represents a departure from traditional industry reliance on mere compute scaling. As the company prepares its next-generation SPU, it continues to focus on developer adoption, with nearly half of its customer base actively building custom models through the firm's dedicated portal.

Comments (0)

Leave a comment

No comments yet. Be the first!