Startups Chase the Next Big Thing in LLMs
Liquid AI
A wave of startups is trying to overcome the limitations of transformers, the technology underpinning today's LLMs. They are proposing new architectures and techniques to make models faster, more efficient, and potentially smarter, aiming to dethrone the AI giants.
Transformers, introduced in a 2017 Google paper, power every major large language model, but their dense attention mechanism is computationally expensive and struggles with long contexts. Startups are tackling these flaws: Subquadratic claims its sparse attention model SubQ rivals mainstream LLMs; Manifest AI replaces attention with 'power retention' and has released PowerCoder and Brumby; Liquid AI pairs transformers with liquid neural networks to create hybrid LFMs (liquid foundation models) that are smaller and more energy-efficient; Inception uses diffusion to generate entire text blocks at once, with its Mercury 2 model claimed to be 10 times faster than GPT-4; and Pathway's 'Dragon Hatchling' model excels at sudoku puzzles, suggesting a move beyond language. These startups see an opportunity to change how LLMs are built, with the potential for significant gains in speed, efficiency, and capability.
- Abbreviations
- LLM = Large Language Model — Большая языковая модель
- LFM = Liquid Foundation Model — Жидкая фундаментальная модель
Source: MIT Technology Review —
original
