Latest innovations in LLM architectures: shared KV caches, per-layer embeddings, and compressed attention
New open LLMs increasingly focus on efficient long-context processing. Gemma 4 (Google) uses KV tensor sharing between layers to reduce cache size, as well as per-layer embeddings to improve parametric efficiency. Laguna XS.2 (Poolside) employs per-layer attention budgeting, and DeepSeek V4 introduces mHC and compressed attention mechanisms.
Google/DeepMind
Poolside
DeepSeek
Technology Innovation Institute






