The New Gold Mine for Young Entrepreneurs: The Glorious Transformation of the AI Data Business and the Capital Boom
Summary of Key Points
This report reveals a significant industry shift: The business of AI training data, once considered "low-end outsourcing," is being redefined by a group of young entrepreneurs (born in the 95s and 00s) with backgrounds from major tech companies, and it has quickly become highly sought after by venture capitalists (VCs).
In the past, data annotation involved finding cheap labor to label data, with low barriers to entry and thin profits, which did not attract investors. However, with the increasing demand for high-quality data from large models and the emergence of new requirements such as "reinforcement learning environments," data companies are no longer just selling "corpus" but rather "infrastructure for intelligent training."
The article highlights a startup valued at $2.5 billion that is about to receive investments from Sequoia, Alibaba, and Tencent as an example. Such companies have transformed from businesses focused on cash flow into high-valued tech companies by providing high-quality expert data, building verifiable training environments, and even entering niche areas like financial forecasting. Although there are limitations to simply selling data, these young entrepreneurs are trying to break through these barriers by offering SaaS services, integrating deeply into customer workflows, and exploring technologies like "Recursive Self-Improvement (RSI)," aiming to transform data companies into providers of essential resources for the AI era, potentially even into new model development centers.
---
In-Depth Analysis: Understanding the New Logic of the AI Data Business from Five Dimensions
1. The Reversal of Roles: From "Cheap Labor" to "Graduates from Top Universities"
Plain Language:
Data annotation used to involve people in third- and fourth-tier cities, working for low salaries to click on images on screens. It was a typical "labor-intensive" task, with the lowest bidder getting the job.
But now, the landscape has changed. Many of the people feeding data into AI systems are graduates from prestigious universities (Peking University's Chinese Department), medical doctors, or interns from top companies like Alibaba's Tongyi or Moon's Dark Side. They jokingly call themselves "data contractors," but their work is far from mundane.
Why the Change?
AI has become much more sophisticated and no longer needs massive amounts of low-quality data; it requires high-quality, precise information.
- Previously: AI could be trained with basic instructions like "This is a cat."
- Now: AI needs to write code, make decisions like a lawyer, or diagnose medical conditions. This requires expertise. The hourly rate for such expert-level annotation can range from $500 to $1,000 or more. These highly educated individuals not only improve data quality but also understand what the models need: how to design effective questions, clean data, and help AI learn through trial and error. This transformation turns what was once a dull task into a technical job with R&D elements, and it is this expertise that attracts capital.
2. Upgrading Requirements: From "Feeding Data" to "Creating Environments for Learning"
Plain Language:
Training AI used to be like feeding a child—showing it thousands of cat pictures to teach it to recognize cats. This is called "supervised learning," where data served as questions and answers.
Now, AI needs to be able to work independently, like a human agent. To achieve this, we need to provide it with an environment where it can learn through trial and error.
- Old Approach: You tell AI if a piece of code is good or bad.
- New Approach: You give it a real browser environment where it can write code, run it, and see the results. The environment provides feedback (errors or successes). This dynamic process is extremely expensive and scarce, and companies like OpenAI and Anthropic are investing heavily in it (OpenAI expects data costs to reach $8 billion by 2030). Those that can provide such environments control the direction of AI development.
3. The Entrepreneurial Profile: Defectors from Big Companies Becoming New Riches
Plain Language:
These entrepreneurs come from large model companies and have hands-on experience in model training. They know the shortcomings of outsourcing and the real needs of data production processes. They see the potential to turn this work into a lucrative business.
- Background: They often worked as interns or early employees at companies like DeepSeek, Kimi, or Alibaba.
- Motivation: They realized that this work could be more profitable if done in-house and decided to start their own businesses.
Why Do Investors Love Them?
1. Low Trust Risk: They understand the technology, so investors don't need to explain complex concepts like RLHF (Human Feedback Reinforcement Learning) or data cleaning.
2. High-Capacity Teams: Their teams are highly skilled, with members who are researchers, engineers, and business strategists.
3. Potential for Wealth: Overseas companies like Scale AI, Mercor, and AfterQuery have founders in their 20s who are already billionaires. Domestic investors see similar opportunities and are willing to invest heavily.
4. Business Challenges: The Limits of the Data Business and How to Overcome Them
Plain Language:
While the concept sounds promising, investors recognize the limitations: Simply selling data is not enough to support a high valuation.
- Low Profit Margins: Companies like Mercor have to pay 60%-70% of their revenue to experts.
- Unstable Revenue: Data deliveries are one-time; if the quality is poor, customers may cancel orders.
- Limited Exit Paths: In China, these companies may face limited opportunities for listing or acquisition if they remain just traditional outsourcing providers.
How to Overcome These Challenges?
Young entrepreneurs are shifting from selling data to providing services and infrastructure:
1. SaaS/Platformization: They offer systems rather than individual data sets, helping companies set up AI workflows.
2. Niche Focus: They target specific industries, such as financial forecasting, and charge for API usage.
3. Building Model Development Capabilities: They use their data and experience to train their own models or support other companies in model development.
5. The Ultimate Concept: Recursive Self-Improvement (RSI) and Its Infinite Potential
Plain Language:
RSI is the latest and most advanced concept, offering exponential potential for AI evolution.
- Current Process: Humans write code, AI learns, humans evaluate, and humans adjust the code.
- RSI: AI writes code, evaluates its own performance, adjusts it, and generates new tasks, creating a self-improving cycle.
Only giants like OpenAI and Anthropic can currently implement RSI.
How Data Companies Can Participate: They develop systems that automate parts of this process, providing essential resources for AI development.
- Investor Perspective: This shifts the focus from cash flow to technological dominance. A company that can accelerate AI evolution becomes a valuable asset, with unlimited potential as long as AI continues to advance.
In Summary:
For young entrepreneurs, the data business offers quick profits and the opportunity to build a strong foundation in talent and technology. As an investor noted, "If a team can generate revenue while building research and engineering capabilities, that’s very attractive." This is not just a quick way to make money but also a gateway to the core of the AI industry.