虎嗅

From DeepSeek and Kimi to the "Huang Renxun Alliance": What exactly is being "opened up" with the open-sourcing of AI models?

原文:从DeepSeek、Kimi到“黄仁勋联盟”:AI模型开源,到底“开”了什么?

Summary of Key Points

Recently, Chinese open-source AI models (such as DeepSeek V4 Pro and Kimi K3) have made a strong emergence. Their level of technical openness (e.g., by making toolchains and infrastructure available publicly) and their commercial viability (through lenient licensing agreements) far surpass previous efforts, directly challenging the American AI ecosystem. American AI giants (NVIDIA, Microsoft, Google, etc.) have rarely joined forces to form the "Open Security AI Alliance," with Elon Musk personally advocating for the promotion of open-source initiatives; however, Anthropic has refused to join, citing security concerns and disputes over the practice of "distillation" (the process of using existing models as training data). This article explores topics such as what open-source models actually entail, why Chinese models pose a threat to the United States, how open-source companies generate revenue, the disagreements between the alliance and Anthropic, and the pros and cons of open-source models, highlighting the critical trend of the AI industry shifting from closed-source monopolies to open-source competition.

I. "Open-source models" are not just "open code"; they represent "open-trained brains"

Many people assume that open-source models simply mean making the code available for everyone to use. In reality, there are three types of open-source models in AI:

  • Closed-source: Models like ChatGPT and Claude do not share anything; users can only access their APIs or products.
  • Fully open-source: These models make all training data, code, and toolchains public (e.g., OLMo), but this is relatively rare.
  • Open weights: The current mainstream form of open-source involves making the trained model weights (which represent the knowledge and capabilities learned by the model, represented as billions or trillions of parameters in digital files) and the inference code (which shows how to use these weights) available. Users can download these to their own servers for use or fine-tuning, but the training data and core training processes remain confidential.

For example, Meta's Llama is open-source in terms of its model weights, but the training data and internal tools are not publicly available. DeepSeek V4 Pro goes a step further by also making the engineering toolchain (such as DeepEP) open-source, providing users with everything needed to easily reproduce or improve the model.

II. Why do Chinese open-source models concern the United States?

Chinese models combine high levels of openness and practicality:

1. Transparency of technical details: Models like DeepSeek and Kimi make their research and development insights (e.g., how to optimize MoE models and reduce memory usage) available in technical reports, allowing the industry to learn quickly.

2. Low barriers to commercial use: Licensing agreements are becoming more lenient. DeepSeek uses the MIT license (which is relatively permissive), while Alibaba's Qwen uses Apache 2.0 (which allows for free commercial use). Kimi K3 is free for small and medium-sized companies but has additional restrictions for those with over 100 million monthly active users or annual revenues of over $20 million (e.g., requiring the addition of a logo).

3. Performance on par with closed-source models: Chinese open-source models have performance levels comparable to ChatGPT and Claude. Companies can deploy these models locally, saving costs and avoiding reliance on single vendors (e.g., in case of API bans by the U.S.).

These factors have directly competed with American closed-source models, leading to concerns about the loss of technological advantages, which is why Musk initiated the alliance in response.

III. Open-source model companies are not naive; they have various ways to make money

Making the model weights available does not mean offering them for free. Instead, they have adopted different business models:

1. Selling cloud services: Companies like Alibaba offer Qwen’s model weights, but users must use Alibaba Cloud’s GPUs and storage; the model itself is free, with charges for cloud services.

2. Providing customized services: Enterprises in finance, healthcare, etc., need to deploy models on their own servers (to prevent data breaches). Model companies can help with adaptation, fine-tuning, and maintenance, charging for these services.

3. MaaS (Machine as a Service) revenue sharing: Companies like Together AI deploy open-source models on the cloud and sell APIs, paying royalties to the model owners according to licenses similar to Intel's "Intel Inside" certification programs to ensure quality.

4. Selling tool platforms: They provide tools for fine-tuning and evaluation, charging based on usage or subscription.

For instance, Moonshot AI (the parent company of Kimi) generates annual revenue of $300 million through APIs, enterprise services, and MaaS revenue sharing.

IV. Disagreements between the American alliance and Anthropic: Security vs. business interests

While the alliance aims to promote open-source models, Anthropic has refused to join, due to two main concerns:

1. Security concerns: Anthropy argues that open-source models could be misused by malicious actors (e.g., for generating malicious code), whereas closed-source models can at least be monitored.

2. Disputes over the use of training data: They claim that Chinese models have gained an advantage by using ChatGPT’s outputs as training data, which saves significant computing resources. However, this is a common practice in the industry.

The underlying issue is more about business interests: As an open-source company, Anthropic fears that open-source models will compete with its API-based services. If companies deploy open-source models locally, they may no longer need to purchase Anthropy’s APIs, weakening its market position.

V. The benefits and risks of open-source models: A double-edged sword

Benefits:

  • Cost savings: Companies can avoid paying high API fees and use open-source models for routine tasks.
  • Greater control: They are not at the mercy of single vendors (e.g., in case of API bans).
  • Accelerated innovation: Developers can create AI agents and robots based on open-source models, lowering the barriers to entry.

Risks:

  • Security vulnerabilities: Open-source models, being deployed widely, may be more vulnerable to security breaches and could be used for malicious purposes.
  • Management challenges: With thousands of companies using open-source models, it becomes difficult to manage and update security patches uniformly.

The industry is still debating these issues, but open-source has become a major trend in AI development. It allows more participants to enter the market, breaking the monopoly of a few companies. The key is to find a balance between security and innovation.

In summary, the rise of Chinese open-source models is changing the rules of the AI game, shifting from closed-source monopolies to open-source competition. In the future, those who successfully combine openness with commercial viability will emerge as winners. The American response also indicates that they regard Chinese open-source models as a real threat.