Summary of Key Points
Recently, the AI community witnessed a farce that transitioned from excitement to skepticism: The J-Space team claimed to have developed a Skill plugin that could significantly enhance the performance of DeepSeek V4 Pro without altering the model weights, even surpassing top models like Fable 5. However, when community developers retested the claims, they found not only that there was no actual improvement in performance but also that the plugin consumed more resources. The team’s author refused to share the evaluation details and deleted any posts questioning the results, revealing that the “miraculous” enhancement was actually a scam based on data manipulation. The deception was effective because it targeted a real weakness of DeepSeek V4 Pro—its sensitivity to the external environment—making many people believe in its claims.
1. J-Space’s “Magical Skill”: An Exaggerated Claim of Outperforming Fable5
J-Space’s Cognition Suite V3.6 is essentially a set of “runtime control mechanisms” for models. Instead of modifying the model itself, it adds a layer of management logic during task execution. They identified four common issues with AI assistants:
- Information Overload: Too many goals and tools overwhelm the system, obscuring important information;
- Memory Distortion: Repeated reasoning leads to changes in the meaning of the same names/values;
- Repetitive Failures: Tasks are attempted without analyzing the reasons for failures;
- Premature Completion: Answers are provided prematurely without verifying their accuracy.
Their solution turned the task execution process into a cycle of “brief judgment → execution → in-depth consideration → verification → retry with diagnostic information,” and they used an “external ledger” to ensure key information was not forgotten. They then presented exaggerated results: Terminal Bench scores increased from 87.9 to 90.1, NL2Repo from 61.5 to 73.4, and even claimed to surpass Fable5 and Opus4.8, which quickly went viral.
2. The Community’s Disbelief: Retests Revealed False Claims and Increased Resource Usage
Soon, skeptics emerged:
- Developer A’s Test: Two rounds of A/B testing with V4 Flash showed no difference in completion rates, but J-Space’s approach used more resources; a third-party blind evaluation yielded scores of 8.3 for the control group and only 7.87 for J-Space’s group.
- Developer B’s Verification: Using 8 NVIDIA H20 GPUs to solve 89 problems, only 69 were correctly solved (77.5%), but the report claimed a 87.1% success rate, a nearly 10% discrepancy; failed attempts were retried three times with no improvement.
- Post Deletion: Users who posted questions were blocked by the author, forcing them to re-post their findings; the author never made the evaluation environment, process, or sample data public, further fueling doubts.
3. The Deception Lies in Targeting DeepSeek V4 Pro’s Weaknesses
J-Space’s success was due to their identification of a genuine issue: DeepSeek V4 Pro is highly sensitive to its operating environment. For example, using similar prompts for 3D game testing yielded rough initial results but complete projects after several hours; the same model performed differently under different Harness modes (Standard/PTC/Minimal), with a score difference of 8 points. Since people already believed that adjusting the environment could improve performance, J-Space’s claims seemed plausible.
4. A Lesson for the AI Community: Claims of Performance Improvements Must Be Verifiable
This incident serves as a reminder that any claimed improvement in AI performance must be able to be replicated by others to be credible. Just like financial company results need to be audited, AI achievements should also include transparent details (evaluation environment, process, samples) for verification. No matter how impressive J-Space’s results seem, they are baseless without verifiable evidence. In the future, when you see exaggerated AI claims, ask: “Can this be replicated? Is there any public evidence?”
5. The Technical Issue: The Problems Are Real, but the Solution May Not Be Effective
It’s worth noting that the four issues identified by J-Space (information overload, memory distortion, etc.) are indeed real and being addressed by other teams. For instance, Routing Suite selects the appropriate reasoning mode for tasks, and Anchored Standard uses concise prompts to stabilize model behavior. However, J-Space’s approach either failed to deliver significant improvements or involved data fraud, making it a cautionary tale.
In summary, innovation in the AI community requires genuine technical breakthroughs, not merely relying on deception to attract attention. This case highlights the importance of verifying claims and using reproducibility as a standard for measuring AI progress.