Ray Serve

Latest AI news, models and releases from Ray Serve.

LLM Inference Degradation in Production: KV-cache, OOM, and the p99 Tail

Production LLM inference often degrades gradually due to KV-cache fragmentation, OOM on long contexts, head-of-line blocking, and p99 latency spikes. VK Cloud evangelist Stas Pogorzhelsky explains four mechanisms causing performance decay and offers reproduction scripts and metrics to catch issues before incidents.

vLLMvLLM MetaMeta NVIDIANVIDIA Hugging FaceHugging Face AnyscaleAnyscale Ray ServeRay Serve Text Generation Inference (TGI)Text Generation Inference (TGI) FasterTransformerFasterTransformer
Habr — хаб ИИ05.08 · 04:04
Fresh news