LLM Inference Degradation in Production: KV-cache, OOM, and the p99 Tail
Production LLM inference often degrades gradually due to KV-cache fragmentation, OOM on long contexts, head-of-line blocking, and p99 latency spikes. VK Cloud evangelist Stas Pogorzhelsky explains four mechanisms causing performance decay and offers reproduction scripts and metrics to catch issues before incidents.
vLLM
Meta
NVIDIA
Hugging Face
Anyscale
Ray Serve
Text Generation Inference (TGI)
FasterTransformer
Habr — хаб ИИ05.08 · 04:04
