Transformers Are Showing Their Age: Startups Chase the Next Big Thing in LLMs, and AI Academic Research Faces New Realities
NVIDIA
Meta
MIT Technology Review explores the limitations of transformer models in large language models and highlights four new ideas to overcome them. Additionally, the newsletter reports on the shifting landscape of AI academic research, Nvidia's $500 billion infrastructure deals, and other top technology stories.
The newsletter's lead story discusses how transformers, introduced by Google nine years ago, are becoming a bottleneck in LLMs due to their expensive dense attention mechanism and inability to handle large amounts of information. Four new innovative approaches are highlighted that could make LLMs faster, more efficient, and smarter. Another article reports on the AI2050 program's convening, where AI professors are negotiating new realities of academic research, as most are university researchers. Nvidia has secured $500 billion from Wall Street, with deals from BlackRock, Goldman Sachs, and four others, for AI infrastructure. Mark Zuckerberg's new manifesto advocates for open-source AI to save the US, and Bernie Sanders has called for a pause on AI development. A US court will allow thousands of social media lawsuits to proceed against addictive mechanisms. Unitree's IPO is massively oversubscribed, and Flock's car-tracking cameras face backlash. China introduces rules for emotionally interactive AI, an AI tool claims to pick the best 1% of scientific papers, and an 82-year-old rejected $26 million for a data center. The newsletter also includes a feature on the ethical mess of genetic embryo testing.
- Abbreviations
- LLM = Large Language Model — большая языковая модель
Source: MIT Technology Review —
original
