VentureBeat
Follow
Stop adding more GPUs: Weka's new storage platform reduces load by caching 100% of an AI model's pre-calculated tokens
GPU memory is the most expensive and rapidly depleted resource in AI production. Longer context windows and multi-turn conversations cause inefficient recomputation of previously processed information. Weka, with its NeuralMesh 6 platform and Wekapod 3 hardware, aims to extend GPU memory using affordable flash storage. Their Augmented Memory Grid aggregates NAND flash to mimic GPU memory at a lower cost. This innovation enters a competitive market with established players like Dell and NetApp also focusing on AI infrastructure. Weka emphasizes its AI-native design, addressing customer needs for immediate compute availability. The core benefit is improved GPU utilization, reduced inference costs, and faster AI workload deployment. This technology is particularly valuable for large-scale AI operations and those experiencing rapid growth. Key NeuralMesh 6 features include composable and virtual multi-tenancy for efficient resource sharing. It also offers unified file and object storage, eliminating data duplication. Metadata-first replication speeds up data availability in destination environments. The Augmented Memory Grid specifically tackles wasted compute by caching pre-calculated tokens, preventing redundant processing in extended conversations. This approach allows for significantly more NAND storage than traditional GPU memory, enabling complete caching of pre-calculated tokens.