Summary of the Key Points
This report is like a cold shower on the currently overheated large model industry: All the mainstream AI products today (from GPT to domestic products like Tongyi Qianwen) are based on the Transformer technology introduced by Google in 2017. This technology has propelled the AI industry to unprecedented heights over the past seven years, but it has now reached its inherent performance ceiling. The new features such as “long context processing” and “accelerated reasoning” that people have recently seen are essentially just patches for the inherent flaws of the old technology, which do not address the core issues of skyrocketing computing costs, explosive energy consumption, and the inability to handle complex tasks.
A group of AI startups without historical burdens have developed four completely different technical approaches that can significantly increase the speed of large models while reducing costs by a factor of several. These approaches can even perform complex reasoning tasks that current Transformer models cannot handle, and they could potentially become the next wave of innovation in the AI industry, potentially reshaping the competitive landscape.
---
Detailed and Easy-to-Understand Explanation
1. Why has the once-great foundation of AI become a hindrance now?
The core advantage of Transformer is also its biggest weakness: its approach to processing text is overly rigid. For example, if you provide 10,000 words of input, the model has to compare each word with every other 9,999 words, which is like reading a 10,000-word article and considering every possible combination of words before understanding the meaning. While this approach offers high precision, it comes at a huge cost in terms of computational power. OpenAI spends billions of dollars annually on purchasing computing resources, and the International Energy Agency predicts that data center energy consumption will double by 2030, equivalent to the electricity consumption of most of the European Union.
The problems are compounded by the increasing complexity of the tasks AI needs to handle. For instance, asking an AI to read an entire codebase, run experiments independently for weeks, or solve tens of thousands of math problems is beyond the capabilities of the rigid Transformer model. The long-context processing features that have been developed are essentially makeshift solutions, similar to trying to run the latest 3D games on a 20-year-old computer by adding external hard drives.
2. “Slimming down” the attention mechanism: Cutting 90% of unnecessary calculations without sacrificing performance
One practical approach is to focus on modifying the most energy-consuming part of the Transformer’s algorithm—namely, the word-by-word comparison process—and eliminate all the unnecessary calculations. Similar “sparse attention” techniques have been tried before, but they resulted in a loss of semantic understanding. Two startups have solved this problem: Subquadratic has developed a dynamic filtering mechanism that identifies key information and ignores irrelevant words, significantly improving performance. Manifest AI has replaced the traditional all-examine approach with a “rolling summary” system, which is like taking meeting notes during a three-hour discussion—only the relevant content is recorded, and old information is automatically deleted. This not only reduces computational load but also allows existing open-source models to use this feature with minimal modification.
3. The counterintuitive approach: Smaller models are smarter and more energy-efficient
The industry’s conventional belief is that larger model parameters lead to better performance, but Liquid AI, a company founded by MIT, takes a different approach. They combined “liquid neural networks” inspired by worm brains with Transformer technology to create models that are several times more efficient than traditional models. These models are small enough to run on a Raspberry Pi and can be powered by low-end chips in vehicles. Small companies with annual revenues of less than $10 million can use them for free, lowering the barriers to using AI. This means AI can run locally on mobile devices without worrying about privacy concerns or slow responses.
4. Overcoming the word-by-word limitation: Enabling AI to output entire segments of text in one go
AI chat tools often display output one word at a time, which is due to the inherent limitation of Transformer models. A startup called Inception has adopted the mature diffusion technology from image generation to improve text processing. Their latest model performs as well as the 2023 version of GPT-4, but at a fraction of the cost. Google is also exploring this approach. With this technology, you can get an immediate response without waiting for the model to process each word individually.
5. The radical idea: Letting AI think using abstract logic instead of text
All current Transformer models have an inherent performance limit because they must convert all their thoughts into text sequences. This makes them inefficient and unable to handle abstract problems. Pathway has replaced the traditional attention mechanism with a mathematical “state space” approach, allowing the model to process information directly in an abstract format. Their model can solve 97% of extremely difficult Sudoku puzzles. They believe that future AI systems should learn from existing text data but also be capable of original research, such as tackling cancer, which requires unconventional problem-solving methods.
In summary, these new technologies represent a shift in the AI industry, potentially revolutionizing how we use and interact with AI systems.