虎嗅

"Name-duplicated 'hack' goes viral: It allows DeepSeek V4-Pro to dominate Fable5, but no one can replicate the results; however, the token cost has doubled."

原文:撞名Anthropic的“外挂”刷屏:让“DeepSeek V4‑Pro碾压Fable5”但无人能复现,Token开销反而翻倍

Summary of Key Points

A recent community project called J-Space Cognition Suite has gained attention, claiming that by simply wrapping the DeepSeek V4-Pro model with an auxiliary tool called Harness, its performance can be significantly improved, even surpassing that of Fable 5. However, there are three major controversies surrounding this claim:

1. Lack of third-party validation: All the results were measured by the project itself;

2. Misleading similarity with Anthropic’s J-space: The two technologies are fundamentally different;

3. Double-token consumption: This leads to increased costs in practical use.

This incident highlights the current trends in AI model auxiliary tools (Harnesses): they aim to address specific shortcomings of models and there is a division in development approaches, with options being either general-purpose or tailored to specific models.

1. The Project’s Claims Are Promising, but No One Can Replicate the Results

The project claims that using Harness significantly boosts DeepSeek V4-Pro’s scores in several tests—e.g., NL2Repo increased from 61.5 to 73.4, and Terminal-Bench 2.1 from 87.9 to 90.1. However, these results were all obtained by the project team, and no independent team has been able to replicate them under the same conditions (model version, testing environment, parameters).

Why is replication so important? Because self-measured results could be biased due to hidden adjustments in testing conditions or random variations in data. Without third-party verification, these improvements are merely speculative and cannot be taken seriously. Some developers who tried to replicate the results found that performance did not improve; in some cases, it even declined.

2. Don’t Confuse the Two J-Spaces!

Many people mistakenly assume that J-Space is related to Anthropic, the company behind the Claude model. However, they are entirely different:

  • Anthropic’s J-space: Focuses on studying the internal mechanisms of their own Claude model (such as how information is transmitted and memory is stored);
  • The community project J-Space: Is an external auxiliary tool designed to manage the model’s task execution process (e.g., determining when to retry, tracking progress, and verifying task completion).

The similar names are either a coincidence or a strategic attempt to capitalize on the popularity of Anthropic. However, this has led to confusion among users, who mistakenly believe it involves Anthropic’s technology.

3. Does Using J-Space Actually Cost More?

Tokens represent the “cost units” for AI models; more tokens used mean higher costs. Developers have found that using J-Space can increase token consumption by 2 to 4 times (for example, increasing costs by 50% to 300% when testing the V4 Flash version). In some cases, performance did not improve, and in others, it even declined.

This suggests that even if the claimed improvements are true, the increased cost may make J-Space less viable for practical use, especially for businesses where cost is a critical factor.

4. Harnesses Are Designed to Address Common Issues in Long-Term Tasks

The problem J-Space aims to solve is common to all AI models:

  • Representational drift: The model may forget its goal mid-task (e.g., repeating already completed steps or deviating to unrelated content);
  • Premature termination: The task may be declared complete before it is actually finished.

Current Harnesses are working to address these issues: for instance, DeepSeek’s official Harness includes a “task progress dashboard” to prevent the model from going off track, while some approaches use context compression to reduce memory usage but may result in information loss. Other tools shift the responsibility for determining task completion from the model to an external tool. These solutions are trying to balance performance and cost, but no perfect solution has yet been found.

5. Two Directions for Harness Development: General-Purpose or Tailored?

There are two main approaches to developing Harnesses:

  • General-purpose: (e.g., OpenCode) Aim to work with multiple models by dividing tasks among multiple agents (e.g., one for planning and another for execution), but this can lead to complex permission management issues (agents may bypass restrictions).
  • Tailored: (e.g., DeepSeek’s DSH, OpenAI’s Agents SDK) Developed by model manufacturers to fit their specific model characteristics, offering better stability but limited versatility.

Neither approach is perfect; general-purpose Harnesses are more flexible but prone to issues, while tailored ones are more stable but less versatile.

Conclusion

The J-Space project seems more like a marketing effort, proposing ideas for solving long-task problems without providing solid evidence or demonstrating significant cost savings. It highlights that improving AI model performance relies not only on the model itself but also on external auxiliary tools. In the future, Harness development will focus on balancing performance enhancement with cost control, while continuing to explore both general-purpose and tailored approaches. For users and businesses, it’s important to verify claims independently and consider the costs before adopting such tools.