虎嗅

"7 sampling points, not a single drop of real water: We live in a cocoon wrapped in false data"

原文:7个采样点,没有一滴真水:我们生活在一个被虚假数据包裹的茧房里

Summary of the Key Points

This article begins with a revelation by CCTV, which undercover investigated four environmental testing companies and exposed outrageous practices such as using tap water from villagers’ wells or nearby pipes at seven groundwater monitoring sites, playing white noise in offices to meet noise pollution standards, and sample collectors digging shallow pits to pretend to take deep samples. It highlights a truth that many people are unaware of: the vast majority of the core data upon which society relies—whether it’s environmental monitoring, credit reports, macroeconomic statistics, AI training data, medical records, educational certifications, or vehicle and bridge inspection data—lack an independent third-party verification mechanism. The cost of falsifying this data is extremely low, and as a result, ordinary people are constantly surrounded by information that appears legitimate but is actually incorrect. This phenomenon is often referred to as a “data cocoon.” The article also provides a practical “data survival guide” that doesn’t require specialized knowledge to help people avoid being deceived by these misleading figures.

---

Detailed and Easy-to-Understand Explanation

1. “None of the samples at seven monitoring sites contained real water”: Our data systems, which we’ve used for decades, inherently lack a “third-party quality inspection” mechanism.

When people first hear about environmental testing fraud, their reaction is often that the companies involved are extremely unethical. However, this is not just a matter of individual wrongdoing; the problem lies in the design of the entire data production system. No one independent of the data producers is responsible for verifying its accuracy. For example, sample collectors collect the samples and fill in the results themselves, with no one on-site to confirm whether the samples were actually taken from the designated locations. Banks enter their own credit data, and no third-party organization checks whether borrowers really owe the money. Hospitals create and store their own medical records without supervision to ensure accuracy. Even in the case of listed companies’ audits, the companies often provide the data to the auditors, which may be fabricated in advance. It’s like going to a restaurant where the chef issues their own “hygiene certificate” without any external verification.

There are over 10,000 third-party organizations specializing in environmental monitoring nationwide, and there are 370,000 companies required to conduct their own pollution monitoring. Since 2022, more than 2,400 cases of fraud have been identified, resulting in the arrest of over 300 people. It wasn’t until 2026 that the new “Ecological Environment Monitoring Regulations” introduced a “dual penalty” system, punishing both the companies and the individuals involved, effectively filling a gap that had existed for many years.

2. Your daily life is already affected by fake data, you just don’t realize it.

Many people think that data fraud only affects industries they don’t relate to. However, every decision you make in the morning could be influenced by misleading information. For instance, when applying for a mortgage, the credit report used by the bank may contain incorrect figures, leading to low-income individuals being charged for loans they don’t owe, or factory owners being labeled as in arrears and unable to obtain loans for months, forcing them to sell their goods at a discount. When investing in funds or stocks, the U.S. non-farm employment data, which is widely referenced, is often based on historical models and not actual on-site surveys. In 2025, 898,000 non-existent job positions were corrected retroactively, causing millions of people to make wrong investment decisions and suffer financial losses. When searching for the best robotic vacuum cleaner, the rankings provided by AI may be based on fabricated reviews paid for by advertisers. When buying a school district house, the school’s enrollment rates may be inflated, and a car that passed a safety inspection last year may be deemed substandard after the inspection station was bribed.

3. The most deceptive aspect of the “data cocoon” is that fake data often looks more “perfect” than real data.

The reason it’s so difficult for people to spot fake data is that it is designed to mimic the real data. Real environmental data fluctuates; for example, water quality may improve after rain or decline during droughts. However, fake data remains consistently within the acceptable range, with no anomalies over several years, making it impossible to detect issues from the reports. Real school enrollment rates vary, but fake data shows steady increases year after year, making parents believe the school offers excellent education. Real employment statistics also fluctuate, but fake data appears stable, giving investors a false sense of security. As the sample collector mentioned, “With just a photo, everything can be made to seem legitimate.” Fakers ensure that all required documents and reports comply with regulations, making it appear everything is correct until problems arise.

4. You don’t need to be a data expert; seven simple habits can help you avoid being deceived.

Ordinary people don’t need to learn complex statistical or auditing skills. By adopting a few simple habits, you can avoid 90% of the pitfalls associated with fake data:

  • Don’t rely on a single source of information: Don’t make important decisions based on just one source. For example, verify different sources when ordering food or consulting multiple AI systems. Cross-check information from different hospitals for serious illnesses.
  • Don’t blindly trust the word “qualified”: Just because something is labeled as qualified doesn’t mean it’s accurate. Always verify additional information from third parties, especially for matters involving safety or large amounts of money.
  • Don’t take newly released data at face value: Initial versions of economic data, company reports, or news statistics are often preliminary and may be revised later. Don’t rush into investments or major decisions based on them.
  • Don’t rely on AI as the ultimate authority: While AI can be useful for general information, don’t trust its conclusions in critical areas like healthcare, finance, or law, as its answers may be biased.
  • Be skeptical of overly perfect data: Data that appears flawless—such as continuously meeting pollution standards or consistently high school enrollment rates—is likely fabricated.
  • Keep original copies of important information: Back up medical records, contracts, and other important documents. In case the system alters the data, these physical copies can help you prove your case.
  • Always ask important questions: Who produced the data? Does it have any motive for fraud? Has it been independently verified by a third party? Asking these questions can help you avoid being misled by fake information.