WEKA has announced impressive performance metrics for its NeuralMesh platform, achieving a tenfold increase in AI inference throughput on Oracle Cloud Infrastructure (OCI). This development is important for organizations aiming to optimize GPU resources while handling large-scale AI workloads.
The benchmarks indicate that WEKA’s solution, when combined with the Augmented Memory Grid, can support over 5,000 concurrent users on a nine-node bare-metal H100 cluster. In comparison, traditional DRAM-only configurations typically manage around 600 users. This significant change effectively eradicates performance bottlenecks related to memory saturation, increasing the active cache working set from 8.64 TiB of DRAM to an astonishing 287 TiB of usable NVMe storage.
Transforming AI Workloads
The implications of this enhancement are substantial. Serving ten times the number of concurrent users without needing additional infrastructure allows for significantly higher returns on investment. Furthermore, the throughput measurement indicates that NeuralMesh can handle about two million tokens per second, far surpassing the under 200,000 tokens achieved through DRAM-only systems. This boost in token output is vital for real-time AI applications, including search functions, summarization, and multi-turn conversational agents, where speed and efficiency directly influence user experience and revenue potential.
Liran Zvibel, CEO of WEKA, pointed out that the bottleneck in AI inference often arises from the limitations of effective memory available to GPUs. "Inference is bottlenecked by how much effective memory is available to GPUs," he stated. He added that the advancements presented by NeuralMesh are not just a result of improved hardware but also a strategic solution to the memory wall that has historically limited AI performance.
Cost Efficiency and Scalability
The benchmarks also reveal a sevenfold increase in the volume of tokens served, with NeuralMesh managing five billion tokens in a single hour during a test involving 2,400 users. This accomplishment signifies a substantial reduction in cost per token, improving the return on investment for organizations managing complex AI workflows. In contrast, traditional DRAM configurations handled only 700 million tokens in the same timeframe.
Pablo Selem, senior director of software development at Oracle Cloud Infrastructure, commented on the significant potential of WEKA’s technology. "These benchmarks show how WEKA’s NeuralMesh platform with Augmented Memory Grid on OCI helps remove memory bottlenecks so customers can support larger, more demanding inference workloads without simply adding more GPUs."
A New Era for AI Economics
As enterprises increasingly depend on AI for various applications, WEKA's findings establish a new benchmark for what can be achieved in AI infrastructure. The capability to process more tokens at lower costs while accommodating a greater number of concurrent users fundamentally reshapes the economic landscape of AI deployment. This shift may encourage more organizations to invest in AI capabilities that were previously limited by hardware constraints.
WEKA's advancements make a strong case for rethinking how AI resources are allocated and optimized. By tackling memory constraints, organizations can improve their operational efficiency and promote growth in AI-driven services. The future of AI workloads may hinge on solutions that emphasize memory optimization alongside traditional hardware improvements.
The stories that move AI & crypto markets — before the market reacts.
Free. 7am ET. Five stories. 62,400 readers.

