Models RSS

Models 🇨🇳

GPT-5.6 is optimizing itself, a new bootstrap approach has emerged

OpenAI's technical report reveals that GPT-5.6 is now applied in production to optimize its own runtime environment, including analyzing traffic, adjusting request routing, rewriting low-level kernels, and optimizing speculative decoding. These improvements reduced end-to-end service cost by 20% and increased token generation efficiency by over 15%. Meanwhile, human oversight remains, with optimization goals and deployment decisions controlled by people.

OpenAIOpenAI
QbitAI 量子位30.07 · 12:04
Models 🇨🇳

PDD Architecture Breaks Cross-Cluster Heterogeneous Inference Barrier: 51.5% Lower First-Token Latency, 37.5% Cost Reduction

InfiniteG (Infinigence) unveiled PDD, a three-layer cross-cluster heterogeneous inference architecture that reduces first-token latency by 51.5% and per-token cost by 37.5%. By inserting a local RelayDecode (RLD) instance to mask WAN KV Cache transfer delays, PDD enables efficient utilization of dispersed homogeneous clusters for LLM inference.

InfinigenceInfinigence
QbitAI 量子位30.07 · 11:02
Fresh news