Shanghai AI Lab and Shanghai Jiao Tong University have introduced NCP-ArchPreview, an 8.9 billion-parameter model that extends the usual next-token prediction setup by also forecasting internal concept signals for a short span of upcoming text. Those signals are then used to assist token generation, allowing the model to keep emitting output one token at a time while internally anticipating what comes next.
The same mechanism is also applied to speculative decoding. In systems such as DFlash and DSpark, a smaller draft model predicts upcoming text before a larger model verifies it. The NCP team feeds internal concept signals to that smaller model as an added hint about how the continuation may unfold, then lets it predict 16 tokens in parallel. According to the reported tests, this improved how much of the draft the larger model accepted.
Across four evaluations, the average number of tokens accepted per verification rose from 5.933 to 6.180, a 4.17% increase. On HumanEval, the gain reached 7.59%. Adding the signal increased the draft model by only about 40,000 parameters. At the same 8.9B scale, NCP also reduced training loss to a similar level using about 85% of the training compute required by a standard Transformer.
Shanghai AI Lab and Shanghai Jiao Tong University have released NCP-ArchPreview, an 8.9 billion-parameter model. Unlike standard large models trained mainly to predict the next token, NCP also predicts internal concept signals tied to a short stretch of upcoming text and uses them to support token generation.
The model still outputs text one token at a time. Inside the model, though, it is already making an early guess about what comes next.
Internal concept signals added to draft generation
The same signals can also be used to speed up speculative decoding. In approaches such as DFlash and DSpark, a smaller model drafts the continuation first, and a larger model then checks it in one pass. The NCP team passes internal concept signals to that smaller model as well, effectively giving it a hint about how the next part may unfold before asking it to predict 16 tokens in parallel.
According to the team, that leads the smaller model to guess more of the continuation correctly.
Test results and training efficiency
Across four tests, the average number of tokens accepted each time the larger model performed verification rose from 5.933 to 6.180, an increase of 4.17%. On HumanEval, the improvement reached 7.59%. Adding the signal increased the draft model by only about 40,000 parameters.
At the same 8.9 billion-parameter scale, NCP was also able to reduce training loss to a similar level while using about 85% of the training compute of a standard Transformer.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.