These startups are chasing the next big thing in LLMs

Share

Multiple startups are developing alternative neural network architectures to replace or supplement transformers, including sparse attention mechanisms, retention-based models, liquid neural networks, diffusion-based text generation, and state space models. These approaches aim to address transformers' computational inefficiency, high energy costs, and limitations in handling large context windows and certain reasoning tasks.


Source: MIT Technology Review