Hardware & InferenceBusiness & Market 🇨🇳 04.08.2026 07:02

Billion-DAU App's Compute Crisis: Inference Costs Inverted, Cross-Cloud Architecture Cuts GPU Cluster by 75%

NVIDIANVIDIA AkamaiAkamai
A cross-border AI fashion and shopping app with over 100 million daily active users was losing $1 per user due to high inference costs and low ARPU. Switching to Akamai's cloud with NVIDIA RTX PRO 6000 GPUs and FP4 quantization reduced costs dramatically, cutting GPU cluster size by 75%.
An AI-powered fashion and shopping recommendation app, with over 100 million daily active users, faced a severe cost crisis. The app generated realistic scene images on users' lock screens based on selfies, but the revenue per user (ARPU) was only $2, while cloud compute and data transfer costs were $3 per user, resulting in a net loss of $1 per user. Previously, the company used NVIDIA L4 GPUs from a top cloud provider, where generating one image took 12 seconds, and high egress fees and latency further inflated costs. The company then migrated its AI inference layer to Akamai's cloud, switching to NVIDIA RTX PRO 6000 GPUs. With faster generation (3-5 seconds per image) and FP4 quantization support, the total GPU cluster size was reduced by 75%, and the cost per image became lower than with L4. Akamai also offered egress fees of $0.005/GB, which is less than 1/20th of traditional clouds, and deployed 19 GPU data centers and over 4,400 edge nodes globally, reducing latency to under 10ms for 95% of internet users. The app adopted a hybrid multi-cloud approach, keeping its database and main program on the old cloud and migrating only the inference layer to Akamai, using the open-source MultiKueue scheduler for load balancing. Akamai also provided migration subsidies up to $5,000 and 24/7 support. Additionally, the Korean gaming company DevSisters also adopted RTX PRO 6000 for real-time NPC dialogue generation, using RTX 4000 Ada for offline asset creation.
Abbreviations
ARPU = Average Revenue Per User — средняя выручка на пользователя
GPU = Graphics Processing Unit — графический процессор
CDN = Content Delivery Network — сеть доставки контента
KV Cache = Key-Value Cache — кэш ключ-значение
FP4 = Floating Point 4-bit — 4-битное число с плавающей точкой
FP8 = Floating Point 8-bit — 8-битное число с плавающей точкой
GDPR = General Data Protection Regulation — Общий регламент по защите данных
Source: QbitAI 量子位 — original
Our earlier posts on this topic ↓
Fresh news