Multiverse Computing: CompactifAI nearly doubles Llama 3.3 performance on Intel Xeon 6
Multiverse Computing
Meta
Intel Corporation
Multiverse Computing announced that its CompactifAI technology can nearly double the performance of the Llama 3.3 model when running on Intel Xeon 6 processors. The solution uses neural network compression techniques to accelerate inference without significant loss of accuracy.
Multiverse Computing announced that its neural network compression technology CompactifAI can nearly double the performance of the large language model Llama 3.3 when running on Intel Xeon 6 server processors. The company claims that CompactifAI uses quantum-inspired and tensor compression methods to reduce model size and accelerate inference while maintaining high accuracy. Testing showed that on the Intel Xeon 6 platform with 288 cores, task execution time is reduced by almost 50%, making the model more efficient for deployment in data centers. Intel and Multiverse Computing emphasized that the solution requires no specialized hardware and runs on standard CPUs.
Source: Meta AI (GNews) —
original
