Multiverse Computing's AI Model Compression Speeds Up Inference on Standard Hardware
Multiverse Computing
Multiverse Computing has developed a method to compress AI models, enabling them to run faster on standard hardware without specialized accelerators. The technique, which uses singular value decomposition, reportedly achieves up to 10x speedup on CPUs.
Multiverse Computing, a quantum-inspired computing company, announced a new approach to compress AI models using matrix factorization techniques like singular value decomposition. The method reduces model size and computational requirements, allowing large models to run efficiently on conventional CPUs and GPUs. In tests, the technique achieved up to 10x inference speedup on CPUs without significant loss in accuracy. The company claims this could democratize AI by reducing dependency on expensive, specialized hardware.
- Сокращения
- CPU = Central Processing Unit
- GPU = Graphics Processing Unit
Source: Meta AI (GNews) —
original
