第一财经

**AMD Acquires Talas, an AI推理 Chip Company: The AI Chip Giant Moves Towards Heterogeneous Computing** This headline accurately captures the key points of the news: AMD's acquisition of Talas, a company specializing in AI inference chips, and the strategic move towards heterogeneous computing. It uses idiomatic expressions such as "AI chip giant" and "heterogeneous computing" to convey the significance of the announcement to the financial news audience.

原文:AMD收购AI推理芯片公司Taalas,AI芯片巨头走向异构计算

Summary of Key Points

Global AI chip giants such as AMD and NVIDIA are shifting away from the approach of relying solely on GPUs to handle all tasks, adopting a strategy of using “heterogeneous chips.” This involves acquiring or collaborating with chip companies (for example, AMD’s acquisition of Taalas and its partnership with Cerebras) to customize chips that combine different types of components, thereby optimizing performance for specific AI tasks like inference. This change is driven by the increasingly diverse demands in the AI industry, as a single GPU can no longer meet the high-efficiency requirements of all scenarios. Heterogeneous systems allow various chips to perform their respective functions efficiently, enhancing overall performance.

Detailed Analysis

Why GPUs Are No Longer “Universal”? —— Diverse AI Requirements Mean General-Purpose Chips Are Not Enough

In the past, GPUs were the “all-rounders” in the AI field, suitable for both training and inference tasks. However, AI applications have become more diverse: intelligent assistants require real-time responses (low latency), large models need to generate text quickly, and enterprises need efficient batch inference at low costs. While GPUs are versatile, they are not optimal for specific tasks—similar to a Swiss army knife that is less effective than a dedicated kitchen knife for certain tasks. For instance, Taalas’ chips outperform traditional GPUs by 1000 times when running Llama3.1, demonstrating the advantage of specialized chips. As a result, these giants are combining such specialized chips with GPUs to create heterogeneous systems.

AMD’s Heterogeneous Strategy: Acquisitions and Partnerships to Enhance Inference Capabilities

AMD has recently taken two significant steps:

  • Acquisition of Taalas: Founded in 2023, Taalas develops chips customized for a single AI model, utilizing high-speed SRAM memory to overcome the computational and memory bottlenecks associated with general-purpose GPUs.
  • Partnership with Cerebras: Cerebras’ chips are “wafer-level” solutions that integrate hundreds of thousands of computing cores on a single chip, reducing communication delays during model execution. Both companies produce ASICs (Application-Specific Integrated Circuits) designed for specific tasks. AMD has integrated these chips into its comprehensive AI platform, allowing customers to choose the appropriate option based on their needs—Taalas for fast inference or Cerebras for low latency.

NVIDIA’s Efforts to Strengthen Inference Capabilities

As the leader in GPUs, NVIDIA is also working to address its weaknesses:

  • Earlier this year, NVIDIA obtained intellectual property rights from Groq and hired some of its talent. Groq’s LPU (Language Processing Unit) is designed for language-related AI tasks, featuring fast performance using SRAM memory and low costs.
  • NVIDIA has incorporated Groq’s technology into its Rubin platform to meet the requirements of intelligent agent systems that require low latency and long context processing. This indicates that NVIDIA is also moving away from relying solely on GPUs and using customized chips to optimize inference.

The Advantages of Heterogeneous Chips

The core advantage of heterogeneous chips is their focus on specialized tasks:

  • Although ASICs cannot handle all scenarios like GPUs, they can excel in specific tasks—Taalas offers exceptional speed, Cerebras provides low latency, and Groq offers cost-effectiveness.
  • By combining these chips, the overall efficiency of AI systems is significantly improved. For example, the combined solution from AMD and Cerebras will be available on the Cerebras Cloud later this year, and NVIDIA’s Vera Rubin platform will also be released in the same period, with tangible results expected soon.

Future Trends: The Era of Divided AI Infrastructure

AMD CEO Lisa Su stated, “No single chip can do everything well; this is the world of heterogeneity.” In the future, AI infrastructure will not be dominated by a single type of chip. Instead, there will be a division of labor among different types of chips:

  • Companies will select combinations of chips based on their business needs—GPUs for training, ASICs for inference, and wafer-level chips for low latency applications.
  • The use of customized ASICs will increase, leading to more specialized markets and improved efficiency and cost control in AI systems.

This shift towards heterogeneity represents a development from a “generalized” to a “precision-oriented” approach in the AI industry. It’s akin to a kitchen equipped with various tools (kitchen knives, fruit knives, scissors) each performing its specific task more efficiently.