AI INFRASTRUCTURE

WEKA and Oracle’s Benchmark Demonstrates 10x Throughput Gains in AI Inference

WEKA's NeuralMesh platform on Oracle Cloud achieves unprecedented gains in AI inference, serving 10x more users and tokens without additional GPUs, enhancing the economics of long-context applications.

WEKA and Oracle’s Benchmark Demonstrates 10x Throughput Gains in AI Inference
CoinSynaptic Desk
AI INFRASTRUCTURE · Correspondent
· PUBLISHED JUN 9, 2026 · 3 MIN READ

A recent benchmark test has revealed that WEKA's NeuralMesh platform, when paired with Oracle Cloud Infrastructure (OCI), greatly enhances the performance of AI inference workloads. The results show a tenfold increase in concurrent users and token throughput compared to traditional DRAM-only setups, signaling a shift in how organizations can handle long-context AI tasks.

The benchmarks, validated on a nine-node OCI H100 cluster, demonstrated that WEKA's Augmented Memory Grid enables 10 times more concurrent users, reaching over 5,000 compared to around 600 in standard configurations. This increased capacity comes without requiring additional GPUs, effectively expanding the active cache working set from 8.64 TiB of DRAM to an impressive 287 TiB of usable NVMe. These enhancements reduce the bottleneck caused by memory limitations, allowing businesses to maximize their investments.

In terms of throughput, the NeuralMesh platform achieved approximately two million tokens per second, a significant increase from under 200,000 tokens in DRAM-only environments. This capability is crucial for teams engaged in real-time AI applications, where the speed and efficiency of token processing directly impact user experience and revenue potential.

The impact goes beyond just speed; the benchmarks showed that WEKA's system handled five billion tokens in a single hour during a test with 2,400 users, compared to 700 million tokens from the baseline. This ability to reduce the cost per token at scale is essential for organizations relying on agentic workflows, where memory saturation can lead to inefficient GPU usage and higher operational costs.

Pablo Selem, senior director of software development at OCI, emphasized the importance of these advancements: "Enterprise AI workloads are pushing context windows and GPU utilization to new limits. These benchmarks show how WEKA's NeuralMesh platform with Augmented Memory Grid on OCI helps remove memory bottlenecks so customers can support larger, more demanding inference workloads without simply adding more GPUs." This highlights the need for innovative solutions that tackle the architectural limitations of current infrastructure.

See also  Amazon's Data Centers Consumed 2.5 Billion Gallons of Water in 2025

Liran Zvibel, CEO of WEKA, also underscored the importance of effective memory management in AI inference, stating, "Inference is bottlenecked by how much effective memory is available to GPUs. These results prove that AI token economics aren't solved by hardware alone; they're solved by eliminating the memory wall that has been the real ceiling on what existing hardware can do."

The Augmented Memory Grid decouples key-value cache from local GPU memory, storing it in a high-performance token warehouse that is accessible throughout the cluster. This architectural change allows any host to serve any session while maintaining cache integrity, enhancing performance and enabling better load balancing as concurrency demands increase. The outcome is a more efficient and scalable approach to persistent context memory for AI agents, ensuring that the economics of long-context inference can be managed effectively at scale.

As the demand for AI capabilities continues to rise, the inefficiencies in current infrastructures become more evident. Each cache eviction not only affects GPU cycles and latency but also impacts the overall user experience and the cost of tokens served. WEKA's advancements illustrate a clear path to overcoming these challenges, establishing a new standard for performance in the AI inference space.

In the rapidly evolving AI sector, these innovations represent more than just incremental improvements; they signify a foundational shift in how organizations can utilize their AI infrastructures to enhance both performance and cost efficiency. The implications for the AI token economy are significant, as companies adapt to these new capabilities to support increasingly complex and demanding applications.

See also  AI Agents Redefine Checkout Efficiency in Digital Commerce

CoinSynaptic Desk

AI Infrastructure · 2,404 stories

CoinSynaptic Desk covers the intersection of artificial intelligence and decentralized networks — frontier AI infrastructure, crypto-native AI agents, Bittensor subnets, DePIN economies, and tokenized compute.

THE DAILY SIGNAL

The stories that move AI & crypto markets — before the market reacts.

Free. 7am ET. Five stories. 62,400 readers.