虎嗅

AI Antibody Design Competition: Outperforming Traditional Experiments, Yet Far from Being Considered a Revolution

原文:AI设计抗体挑战赛:超越传统实验,但远未封神

Unveiling the “Big Test” of AI Antibody Design: Beyond the Hype, How Good Is It Really?

Hello everyone, I’m your financial journalist and economist. Today, we’re not talking about the ups and downs of the stock market, but about one of the hottest and most overhyped topics in the biopharmaceutical industry: Can AI directly create antibodies that can cure diseases?

Over the past two years, whenever AI in pharmaceuticals was mentioned, the terms “disruption,” “revolution,” and “instant fame” were often used. Especially with the emergence of AlphaFold, a tool for predicting protein structures, the entire industry was in an uproar. Many companies claimed that given an antigen (such as a virus), AI could design highly affinity antibodies in a computer, eliminating the time-consuming and painstaking screening process traditionally done in laboratories.

However, a recent article in *Nature Biotechnology*, a top journal in the field of biotechnology, conducted a “real-life simulation test” called the “AIntibody Challenge.” This test had no shortcuts or cheating; 29 top institutions submitted 511 antibody sequences designed by AI, all of which were synthesized for actual experiments.

So, what were the results? In a nutshell: AI does have its strengths and can be very helpful in certain aspects, but it’s still far from being omnipotent, and the differences between different AI models are much greater than you might think.

Below, I’ll break down the results of this “big test” into five key points to show you the true capabilities of AI in antibody design.

---

1. How was the test conducted? – Not open-book, but a “blind test”

First, we need to understand why this test is important. Many previous AI papers showed good results because they used “retrospective data.” It’s like a teacher giving you the answers before the exam; of course, you can guess correctly. But real drug development is “prospective”: you only have the raw materials (sequencing data) and don’t know the final outcome; you have to predict based on your skills.

The AIntibody Challenge mimicked the famous protein structure prediction competition (CASP) and used a blind, prospective design method:

  • Test topic: The most extensively studied SARS-CoV-2 receptor binding domain (RBD) was chosen as the target.
  • Three tasks:

1. Affinity optimization: You were given a set of existing antibody data and had to tweak it to make it bind more tightly. This corresponds to the later optimization steps in the industrial process.

2. Selection of the best candidate: You had to pick the best sequence from thousands of options that had the potential to become a drug. This corresponds to the process of selecting candidates after high-throughput screening.

3. De novo design: You had to design a completely new antibody sequence from scratch, without any existing databases. This is the most challenging part of从头 design.

  • Evaluation criteria: The evaluation wasn’t just about how tightly it bound; it also considered whether the antibody could potentially become a drug. For example, would the antibody aggregate on its own? Would it be unstable (thermally)? Would it bind to things it shouldn’t (cross-reactivity)? If any of these criteria were not met, even if the binding was strong, the antibody would be eliminated.

Core logic: This test eliminated the possibility of “ hindsight”; all the sequences submitted by AI were synthesized and tested uniformly by the organizers. This is the true test of AI’s capabilities.

---

2. Highlights: In the “optimization” task, AI can indeed save time and money

Among the three tasks, AI performed best in Task 1: Affinity Optimization.

It’s like having a decent draft (a parent antibody) and asking AI to refine it to make it perfect.

  • Result: The antibody designed by Aureka’s structure-aware model had its affinity increased by about 2000 times, reaching 95 pM (picoM; the lower the value, the stronger the binding). This result was almost on par with the best antibody identified through traditional experimental screening (113 pM), and the difference was not statistically significant.
  • Surprising discovery: A team that used a simple statistical method (ProBioGen) without complex deep learning also came in third place. This shows that with enough real data, simple statistical techniques can be very effective.
  • Industrial value: In this task, AI can indeed be helpful. Traditional experimental screening can take 2-3 weeks; if AI can identify a few high-potential candidates in advance, it can significantly shorten the development cycle and reduce the cost of trial and error.

In simple terms: If you ask AI to refine a good product, it can do a great job. It’s like asking AI to edit a well-written essay; it can spot mistakes and improve the language, with immediate results.

---

3. The embarrassing moment: In the “selection” task, AI isn’t even as good as random chance

This is the most surprising and embarrassing part of the article for the industry:

  • Cruel reality: Researchers found that if you randomly selected an antibody from thousands of sequences, there was a 39% chance of getting a better result than the “control group.”
  • AI’s failure: The performance of most AI models was below this random baseline. In other words, the “top performers” identified by AI were no better than a random choice.
  • The only bright spot: Among the 26 teams, only the University of Washington (WashU) exceeded the random baseline.
  • Limited generalization: The same AI algorithm performed well on one set of data but poorly on another. This shows that AI doesn’t truly understand the rules of antibody behavior; it just remembers the characteristics of specific data.

In simple terms: It’s like asking AI to grade exams. If the questions are similar to what it’s done before (optimization), it performs well; but if the questions are new and require it to find the correct answer from a set of options, it’s not as good as a random guesser. This exposes a major weakness in AI’s affinity prediction capabilities.

---

4. The tricky moment: In the “innovation” task, AI created “beautiful but deadly” molecules

  • The best performer: Xencor’s protein language model designed an antibody with extremely strong binding (2.9 pM), the strongest of all.
  • Fatal flaw: This “champion” antibody couldn’t be eluted from a hydrophobic chromatography column, indicating it would likely aggregate during production, which is a serious issue in drug development.
  • Evaluation loophole: Because the competition’s scoring system allowed other factors to offset failures, it won first place. But in real drug development, such a molecule would be discarded.
  • Limited innovation: Many so-called “new designs” were actually just minor modifications to existing sequences; true de novo innovations were rare.

In simple terms: AI is like a talented artist lacking engineering experience. It can create stunning works (high affinity), but the materials used may be toxic or the structure may not be suitable for practical use. It might win an art award, but it could damage a wall if used in real life. This shows that high affinity doesn’t necessarily mean a good drug; AI still doesn’t understand the biological and physical risks.

---

5. Industry implications: Don’t rely on “universal models”; AI is a tool, not a replacement

This test serves as a wake-up call for the entire biopharmaceutical industry and offers practical advice:

  • No one-size-fits-all model:
  • For optimization, a combination of structure-aware and language models works best.
  • For selection, traditional machine learning or even random selection can be effective.
  • For innovation, protein language models are relatively good, but they need to be rigorously validated.
  • Conclusion: Don’t expect one AI model to solve all problems. The industry should use a combinatorial approach, treating AI as a “highly efficient candidate generator” rather than a “final decision-maker.”
  • Laboratory experiments are essential:
  • Computational scoring can’t replace real experimental testing. High-affinity sequences identified by AI often come with risks such as aggregation and cross-reactivity.
  • The gold standard is always blind, prospective testing combined with comprehensive experimental validation.
  • Directions for future improvements:
  • Remove pre-assigned affinity labels to simulate earlier stages of discovery.
  • Introduce a “no-go” criterion for molecules that are not viable as drugs.
  • Expand the range of antigen targets to test AI’s cross-species and cross-target capabilities.

In summary:

AI antibody design is not a “completed revolution” but is at a turning point. It has already shown its value in partial optimization, significantly speeding up and reducing costs in drug development. However, the ultimate goal of “inputting an antigen and directly obtaining a clinical-grade antibody” is still a long way off due to biological realities.

For investors and industry professionals, defusing the hype is urgent. Don’t believe claims that AI can completely replace experiments; understand the real value of AI in specific tasks. The future winners will be those who know how to perfectly integrate AI with laboratory experiments, not those who blindly worship algorithms.