Hardware & Inference RSS

Mini-PC on Strix Halo under parallel load: 236 tok/s on 32 concurrent requests and three errors

A Beelink GTR9 Pro with Ryzen AI Max+ 395 achieved 236 tok/s aggregate on 32 concurrent requests in short runs, sustaining 226 tok/s average over 30 minutes without throttling. The author discovered a reproducible throughput drop between 8 and 10 concurrent clients across all tested models, and found that speculative decoding (MTP) accelerated single-client performance by 42% but became a penalty under concurrency.

Google/DeepMindGoogle/DeepMind
Habr — хаб NLP24.07 · 10:04
Fresh news