J-Space Faces Fabrication Questions After Community Retest Contradicts Its Published AI Benchmark Claims

J-Space Faces Fabrication Questions After Community Retest Contradicts Its Published AI Benchmark Claims

N
News Editor
2026-08-18 15:31:29
J-Space Cognition Suite, an AI project that gained rapid traction on X, is facing community scrutiny after a retest challenged the benchmark results it had promoted for DeepSeek V4. According to MaxForAI, the project had claimed that pairing V4 Flash with J-Space could match GLM-5.3, while V4 Pro could outperform Fable 5 across multiple agent benchmarks. It also advertised a 2.53x speed improvement and a 2.21x gain in token efficiency. A GitHub user, GoForceX, said they reran the test with an 87-question subset from Terminal Bench 2.1 under high concurrency and confirmed that J-Space-related modules were loaded. The retest reportedly showed the opposite direction from J-Space’s claims: benchmark scores fell slightly after adding J-Space, while token usage and costs increased. The community has since called on the project to release its full evaluation setup, per-question results, execution logs, raw timing data, and token consumption records. So far, the project has mainly published aggregated results, with no complete raw experimental records available to verify the precise figures. The project author had previously said the data was 「indeed exaggerated」 and estimated actual gains at roughly 1.6x to 3x. As of now, there has been no formal response to the latest criticism, and some related issue threads have been deleted.

According to BlockBeats on Aug. 18, citing MaxForAI, J-Space Cognition Suite, an AI project that went viral on X today, has come under community criticism over claims that its published DeepSeek V4 benchmark results cannot be reproduced.

The project had previously said that V4 Flash combined with J-Space could match GLM-5.3, while V4 Pro could outperform Fable 5 across multiple agent benchmarks. It also claimed a 2.53x increase in speed and a 2.21x improvement in token efficiency.

GitHub user GoForceX carried out a high-concurrency retest using an 87-question subset from Terminal Bench 2.1 and said J-Space-related modules were confirmed to be loaded during the run. The result pointed in the opposite direction of the project’s published claims: after adding J-Space, benchmark performance slipped slightly, while token usage and cost both increased.

Following the retest, community members asked the project to disclose its full evaluation configuration, per-question results, run logs, raw timing data, and token consumption records. Based on what is currently available, the project has mainly published summary results and has not released complete raw experiment records sufficient to verify those precise figures.

The project author had previously responded to criticism by saying the figures were 「indeed exaggerated」 and added that the real improvement was roughly in the 1.6x to 3x range. As of now, the author has not issued a formal response to the community’s latest questions, and some related issue threads have been deleted.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
5700

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.