Magic, an AI coding company, said a new pretraining setup can achieve results close to DeepSeek V4 Pro Base using about $500,000 worth of NVIDIA GB200 compute, with total compute demand at roughly 1/50 of the comparison benchmark. The company said the comparison targets a base model that has only gone through pretraining and has not yet received reinforcement learning or other post-training work. That means the result is not presented as evidence that its chat, coding, or agent performance has caught up with the finished DeepSeek V4 Pro product.
Magic said it later scaled the training run by 10x, or about $4 million in GB200 compute. In the same internal evaluation, that larger version outperformed the latest public base models from DeepSeek, Kimi, and NVIDIA that it tested. The company also estimated that reaching the same level at DeepSeek V4 Pro’s training efficiency would cost more than $100 million in compute.
Magic did not release the full training recipe. It only said the efficiency gains came from dozens of changes across model architecture, optimizer design, training objectives, and data processing. Model weights have not been published either, leaving the claimed 50x efficiency gain without independent external reproduction for now.
AI coding company Magic has published a pretraining study saying a new training approach can reach results close to DeepSeek V4 Pro Base with about $500,000 worth of NVIDIA GB200 compute. According to the company, the required compute was roughly 1/50 of the comparison level.
Magic said the comparison is specifically about a base model that has completed pretraining but has not gone through reinforcement learning or other post-training stages. To evaluate the model, it used unseen code, math problems, and research papers, then measured which model predicted the next content more accurately. On that basis, the company said the result reflects base-model pretraining quality rather than finished chat, coding, or agent performance, and does not mean its system has matched the full DeepSeek V4 Pro product.
A 10x larger training run
Magic said it later increased the training scale by 10x, equivalent to about $4 million in GB200 compute. In the same internal test set, that larger version outperformed the latest public base models from DeepSeek, Kimi, and NVIDIA that it evaluated.
The company also estimated that if the same level were reached using DeepSeek V4 Pro’s training efficiency, compute cost would exceed $100 million.
Recipe and weights remain undisclosed
Magic did not publish the full training recipe. It said the efficiency gain came from dozens of changes involving model architecture, optimizer design, training objectives, and data processing.
The model weights have not been released either. As a result, the claimed 50x efficiency improvement cannot yet be independently reproduced by outside researchers.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.