第一财经

The Second Half of Global AI Regulation: As Data Becomes a Strategic Resource, Who Will Define the Standards?

原文:全球AI监管的下半场:数据成为战略资源后,谁来定义标准

Summary of Key Points

This news article focuses on data governance in the era of AI, with the central argument being that data has become a strategic resource. The key to distributing its value lies not in “who owns the data” but in “who establishes the rules and standards for its use.” Currently, there are three distinct approaches being pursued globally: Europe (which emphasizes rules and standards), the United States (which focuses on markets and investment), and certain regions in Asia (which prioritize infrastructure and the ambition to set standards). Developing countries face the risk of merely providing raw data without any say in the formulation of these rules. AI companies often engage in unsophisticated data collection practices and lack transparency about data sources. Incompatibility between standards can hinder scientific progress, which highlights the importance of ensuring that standards are compatible rather than uniform.

I. The Core of Data Governance: It's About Who Sets the Rules, Not Who Owns It

In the past, it might have seemed that those with data had the upper hand, but now the situation has changed. The value of data depends on who establishes the rules for its use. For example, if Country A sets standards for “AI-ready data” (data that is already organized and suitable for AI training), other countries would have to follow these rules if they want to use that data for AI purposes, resulting in more benefits flowing to Country A. UNCTAD points out that the focus of data governance has shifted from “who owns the data” to “who defines what data can be used and how it can be used,” similar to how the person who sets the rules in a game always has the advantage.

II. Three Global Approaches: Europe, the United States, and Asia

The competition for global data standards follows three main paths:

  • Europe: A combination of mandatory rules and prioritizing standards. Europe has incorporated “open data” into its research funding initiatives; for instance, the Horizon Europe program requires that data from funded studies be made available publicly. It also promotes the FAIR principles (Findable, Accessible, Reusable, Interoperable) and the PlanS initiative to make research papers freely accessible. Additionally, regulations like GDPR (General Data Protection Regulation) and AI-specific laws impose clear requirements on both companies and researchers.
  • The United States: A market-driven approach with fluctuating policies. The National Institutes of Health in the U.S. require research projects to submit data management plans, but government policies are not consistent. For example, the White House ordered the free release of federal-funded research data in 2022, which was later overturned by President Trump. This has attracted more capital to AI data initiatives in the U.S., but the rules are less stable compared to Europe.
  • Certain regions in Asia (such as China): Rapid development of infrastructure with the ambition to set their own standards. Asia is making significant progress in data infrastructure and does not wish to follow the rules established by Europe and the U.S. passively.

Each approach has its advantages and disadvantages: the U.S. has substantial resources but lacks consistent rules; Europe has a solid foundation, though it may stifle innovation; Asia has strong infrastructure and is striving for greater influence in setting standards.

III. The Dilemma of Developing Countries

The real concern is not the differences in approaches among Europe, the U.S., and Asia, but the gap in AI data infrastructure. For instance, DNA samples from Africa have often been collected and analyzed in Europe and the U.S., with the results and tools remaining there. Africa only provides the raw material, similar to a farmer selling their crops to others who then process them into valuable products, leaving the farmer with a small share of the profits. UNCTAD warns that without their own AI data infrastructure, developing countries will be limited to acting as mere suppliers, unable to participate in rule-making or reap any benefits.

IV. The Cost of Incompatible Standards: A Barrier to Scientific Progress

Incompatible data standards are like different electrical plugs; you need a converter to use your equipment in another country, which is time-consuming and costly. This applies to scientific research as well. Rare diseases may have only a few cases in each country, making it difficult to identify the cause alone. However, by combining data from Europe and Asia, researchers could make significant breakthroughs. Incompatible standards can prevent such collaboration, delaying or even preventing discoveries. Hil describes this as a direct tax on science: researchers must repeat their work, slowing progress and delaying treatments for the public. Standards do not need to be identical, but they must be compatible.

V. The Chaotic Data Practices of AI Companies

Many AI companies collect data in a careless manner, often without considering its source or providing any traceability. Hil argues that AI companies should treat data just like researchers: they should know where the data comes from, credit the providers, and follow the rules for its use. When packaging data for AI use, the source information should be included and not removed. Data quality should be verified by independent experts using clear standards. Only with these measures can data be trusted, and the distribution of benefits be fair.

In summary, data governance in the AI era is essentially a battle for rules and standards. Those who establish compatible and reasonable standards will control the distribution of data value, preventing obstacles to scientific progress and ensuring that developing countries are not marginalized.