OpenAI chief scientist warns against race to unmonitorable models

OpenAI chief scientist warns against race to unmonitorable models

N
News Editor
2026-09-02 06:39:34
OpenAI chief scientist Jakub Pachocki says OpenAI wants to avoid a race toward “unmonitorable” models that misleading reporting could trigger, according to ChainCatcher. He said today’s frontier models, including Astra, have computational graph depth that differs from GPT-4 by no more than a factor of two. OpenAI has worked to preserve and exploit chain-of-thought monitoring since its earliest reasoning models; Pachocki argues the technique can show how well model alignment generalizes from the training distribution. He cautioned that chain-of-thought monitoring is fragile and the current trend is not encouraging, but said research can strengthen it and the approach has become a core goal of OpenAI’s research program. The statements come after The Information reported that Astra uses recurrent depth, letting the same set of Transformer layers compute repeatedly. Some reasoning can then happen inside the model rather than being written out as a chain of thought, which has raised worries about whether models will become harder to monitor.

According to ChainCatcher, OpenAI chief scientist Jakub Pachocki said he hopes to avoid a race toward “unmonitorable” model behavior that misleading reporting could trigger.

Pachocki said current frontier models, including Astra, run computational graphs whose depth differs from GPT-4’s by no more than a factor of two. Since its earliest reasoning models, OpenAI has focused on preserving and using chain-of-thought monitoring. In his view, the technique can show how well model alignment generalizes beyond the training distribution.

He also said chain-of-thought monitoring is fragile and the trend is not encouraging, but research could strengthen it. The approach has become a core target in OpenAI’s current research plans.

The remarks come after a report by The Information stating that Astra uses recurrent depth, letting the same group of Transformer layers compute repeatedly. Some reasoning can then happen inside the model instead of being written out as a chain of thought. That design has stirred concerns that models may become increasingly difficult to monitor.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
1300

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.