Summary of Key Points
As AI has evolved to its current level, the publicly available data on the internet has been largely exhausted. To become more intelligent—moving from chatbots to “digital employees” capable of performing actual tasks—AI companies have turned their attention to the internal data of startups that have gone bankrupt or been acquired, including code, chat records, emails, and meeting transcripts. The trade in such data has seen explosive growth this year, but it comes with significant issues such as privacy breaches, legal gaps, and ethical conflicts. Ironically, giants like Google and Meta have long been using user data to train AI by modifying their privacy policies in a “legal” manner, while companies like Mercor have merely pushed this practice into a more shady realm. The evolution of AI essentially relies on the consumption of human digital privacy, and this contradiction is unlikely to be resolved in the short term.
Why Do AI Companies Seek Data from Failed Startups?
The intelligence of AI does not come out of thin air; it requires data to feed its algorithms. In the past, research laboratories scoured the internet for web pages, papers, code, and forum posts, which constituted a vast source of “public data.” However, this era of abundance is over. For AI to progress from being able to engage in simple conversations to performing complex tasks, what’s needed are not just answers but the processes involved—how engineers identify problems, discuss solutions, modify code, and fix bugs. This type of information is found in the internal records of failed startups (such as Slack chats, Zoom recordings, and code change histories), which represent the most valuable “real-world business data” for AI training.
For example, GitHub has enabled significant advancements in AI capabilities because its commit logs (records of code changes) contain the complete process from identifying issues to implementing solutions. Other industries lack such readily available data sources, making the internal data of failed startups a crucial resource. This is why companies like Mercor are willing to pay millions of dollars for these records, driving this business trend.
Risks of Acquiring Such Data: Privacy Breaches Are Just the Tip of the Iceberg
The internal data contains highly sensitive information, and mishandling it can lead to serious problems:
1. Inability to Protect Privacy: Although Mercor claims it can obscure sensitive details, the industry knows that the most valuable information is about who said what and when, and what actions were taken as a result. Forging employee names (e.g., calling them “Engineer A”) does not remove associated information like job roles, projects, and customer relationships, which can easily lead to privacy breaches.
2. Lack of Employee Consent: Employees are often unaware that their chat records and meeting transcripts are being sold. Even if they have signed intellectual property agreements, it does not necessarily mean they agree to their work conversations being used for AI training.
3. Commercial Secrets and Biases: The data may contain customer information and company secrets. Worse still, biases present in human decision-making (such as discrimination against certain customers) can be replicated and amplified by AI, leading to “algorithmic biases.”
4. Lack of Clear Legal Guidance: There is no global consensus on whether companies can sell internal work data to AI companies, creating a legal vacuum that could either accelerate AI development or unleash a Pandora’s box of privacy breaches.
Giants Have Been Doing This for Years
Don’t think Mercor is the only company engaging in these shady practices. Giants like Google and Meta have long been using user data to train AI legally:
- Google quietly updated its privacy policy in 2023, stating that it uses public information (including Gmail emails, Google Docs documents, and search history) for AI training. In 2026, it automatically included all users in the AI training program, with users having to opt out.
- Meta also changed its policy in 2023, using public posts and photos from Facebook and Instagram for AI training. It temporarily stopped using EU user data in 2024 but resumed in 2026, giving users only the option to “object” rather than a formal “consent” process.
In essence, Mercor has just brought these practices into the spotlight. The difference is that while giants use legal means (such as modifying service terms), Mercor takes a more overtly shady approach.
How to Address This Situation?
A solution requires efforts from multiple parties in the fields of law, technology, and industry self-regulation:
1. Legal Clarification: It is necessary to establish who owns the data (companies, employees, or customers) and who must give consent before it can be sold.
2. Technological Solutions for Privacy Protection: Technologies like “federated learning” (where AI is trained on the data without transferring the original data) and “differential privacy” (which adds noise to the data to obscure identities without affecting training outcomes) can help reduce privacy risks.
3. Industry and Corporate Self-Regulation: Data companies should establish traceability mechanisms and clearly inform employees before selling their data, obtaining their consent.
However, these solutions will take time to implement. AI companies need to grow, users want their privacy protected, and giants seek profits—this conflict will continue. In the short term, the evolution of AI will likely rely on the consumption of human digital privacy, and the race is far from over.
In Conclusion
Every advancement in AI consumes our digital traces: from public web pages to work records, from the “legal” collection by giants to shady acquisitions by intermediaries like Mercor. While we marvel at the power of AI, we must also accept that its foundation is built on the digital legacy of countless individuals whose privacy has not been properly protected. Future regulations may set boundaries, but for now, the game of trading privacy for intelligence continues.