
A new generation of startups is working to push the boundaries of large language models (LLMs), a technology that has become a cornerstone of the AI industry.
The current dominant technology, called transformers, has been the driving force behind LLMs, but it’s starting to show its age, with a mechanism called dense attention that encodes the meaning of a block of text in a series of numbers.
However, as the length of the text grows, the number of computations needed to process it adds up fast, making it a power-hungry and expensive technology, with OpenAI set to spend $50 billion on computing this year.
The International Energy Agency predicts that the total amount of electricity consumed by AI models will double by 2030.
Transformers also struggle with keeping track of a lot of information at once, which is a major limitation for LLMs that need to carry out tasks that involve processing huge amounts of data, and they are working on new ideas to solve this problem.
Several startups are working on new ideas, including rethinking attention, making models smaller and more flexible, generating text all at once, and moving beyond words, to solve the transformer problem.
Subquadratic, a startup based in Miami, claims it has invented the first sparse attention mechanism that rivals top mainstream LLMs on a handful of tasks, including search and coding, and it is making progress.
Manifest AI, a startup based in San Francisco, is working on a mechanism called power retention, which stores only the most relevant information for a given task and ensures that the amount of data an LLM has to keep track of doesn’t blow up, and they are seeing results.
Liquid AI, an MIT spinout based in Cambridge, Massachusetts, is pairing transformers with its own tech, liquid neural networks, to build smaller and more flexible models, and it has proved popular.
Liquid AI’s models are far smaller and use less energy than most LLMs, and they have racked up almost 34 million downloads, and the company’s models are available for free to any organization with an annual revenue less than $10 million.
Inception, a startup based in Palo Alto, California, is building LLMs using a technique called diffusion, which is better known as the technology that drives most image and video generation models, and it is showing promise.
Inception’s LLMs can generate text all at once, spitting out whole sentences or paragraphs in one shot, making them faster and more cost-efficient than traditional LLMs, and the company’s latest model, Mercury 2, performs as well as some of OpenAI’s GPT-4 models, but is 10 times faster.
Pathway, another startup based in Palo Alto, is working on a type of LLM that can move beyond words and process information in a more abstract way, using a mathematical structure called a state space to compress information.
It is likely that we’ll see more startups working on new ideas to solve the transformer problem and push the boundaries of LLMs, as the AI industry continues to evolve.
The future of AI is looking bright.
With the potential for LLMs to become even faster, more efficient, and more intelligent, it will be exciting to see what these startups and others can achieve in the coming years, including the potential to reshape asteroid detection and other areas.
