OpenAI chief scientist warns AI may move into self-improvement and says slowing down could become necessary

OpenAI chief scientist warns AI may move into self-improvement and says slowing down could become necessary

N
News Editor
2026-09-07 01:48:20
OpenAI chief scientist Jakub Pachocki used a Sept. 6 essay, “An Alien Mind,” to issue one of the clearest warnings yet from a senior lab executive about where advanced AI could be heading. He said AI has already begun to surpass humans in some capabilities in ways he described as transformative, and argued that current progress could continue into recursive self-improvement, or RSI, where AI systems take on a deeper role in building the next generation of AI. Pachocki traced that view back to OpenAI’s 2023 “RLSlow” research project, which he said helped convince the team that pretrained models could develop their own chain of thought through scaled training. Three years later, he wrote, reasoning models can operate computers and graphical interfaces, collaborate with humans and other AI systems, and carry out research tasks on their own, while also changing the balance in cybersecurity. He argued that alignment and monitoring remain unresolved. OpenAI says GPT-6 Astra has benefited from parts of its long-term alignment research and performs noticeably better on alignment than GPT-5.6 Sol, but Pachocki said progress is still far from enough and may not keep pace with broader capability gains. He added that if shared safety standards are not yet in place, voluntary slowing by AI companies could become more common and should.

OpenAI chief scientist Jakub Pachocki published a long essay on Sept. 6 titled An Alien Mind, warning that AI development over the next few years could push into recursive self-improvement, or RSI, with systems taking a growing role in the research and development of future AI models. He wrote that AI has already started to exceed humans in some capabilities in what he described as a transformative way, and said people may not be ready for what comes next.

A turning point, in his account, began in 2023

Pachocki said the story goes back to 2023, when OpenAI made key progress in a research project called RLSlow. Those results led the team to believe that pretrained models could form their own chain of thought through scaled training. He wrote that this was the point when he started to believe he might see machines that are “significantly smarter than humans” within his lifetime.

Three years later, he said, reasoning language models can operate computers and graphical interfaces, work with humans and other AI systems, and even carry out research tasks independently. At the same time, they are beginning to reshape both attack and defense in cybersecurity. If the current trajectory holds, Pachocki said, capability gains over the next few years may be no smaller than those seen in the past few years, with AI taking on a larger role in its own development.

AI is being “grown,” not fully designed

One of the essay’s central arguments is Pachocki’s description of what modern AI is. Rather than something humans fully design from top to bottom, he said, AI is better understood as something that is “grown.” Modern deep learning systems are built by repeating relatively simple optimization steps at massive scale, eventually producing systems of extreme complexity.

Researchers can inspect local mechanisms, he wrote, much as neuroscientists study the brain, but they still cannot fully describe how the whole system works. In that sense, large-scale AI training remains experimental. Teams propose algorithms and predictions, run large training jobs, and the results can still surprise the people building them.

That is also why he chose the title An Alien Mind. Pachocki argued that machine intelligence produced by deep learning should not be compared directly with human intelligence. AI does not have to outperform humans across every domain to have major effects. It only needs to exceed humans across enough important dimensions to become both extremely useful and extremely dangerous.

Alignment is not only about following instructions

Pachocki broke AI alignment into two layers. The first is goal alignment: whether an AI system actually carries out the task a user wants, including following instructions, understanding human aims, and cooperating with people. The second is value alignment, which he described as harder: whether the system still acts in line with higher-level human values in ambiguous, conflicting, previously unseen, or even hostile environments.

The real difficulty, he said, lies in generalization. A model may appear safe in the environment it was trained in, but once its capabilities rise, it faces new settings, or it begins interacting with other AI systems, humans cannot be sure it will keep interpreting those values in the same way.

He stressed one requirement in particular: future AI must continue to uphold the values humans want it to follow whether or not the system believes it is being supervised.

OpenAI says its latest GPT-6 Astra has benefited from some alignment advances from the company’s long-term research and performs noticeably better on alignment than GPT-5.6 Sol. Pachocki also acknowledged that current progress is still far from enough, especially because there is no basis to assume alignment improvements will keep pace with overall capability growth.

Humans are losing visibility into how AI works

OpenAI has long put weight on chain-of-thought monitoring, using a model’s written reasoning process to observe how it thinks. Pachocki said this was also one of the main reasons OpenAI did not publicly release the full chain of thought for o1-preview. If humans directly apply training pressure to the reasoning trace, he wrote, future models may learn to hide thoughts they do not want overseers to see.

OpenAI wanted to preserve chain of thought as a window into the model. By the Astra generation, Pachocki said, that window is becoming less dependable.

He gave three reasons. Reasoning in modern agents is now deeply entangled with interactions involving humans, other AI systems, and tools. AI is getting better at reasoning and even at manipulating its own reasoning process. And as pretraining improves, models can complete increasingly complex tasks without relying on a text-form chain of thought that humans can read.

This creates a central bottleneck. AI is getting smarter, but human ability to understand why it made a decision may not improve at the same pace. Pachocki said the key constraint in future AI progress may gradually shift from whether researchers can make stronger models to whether they can monitor stronger models with enough confidence.

Why keep building stronger systems if the risks are clear

Pachocki’s answer was the familiar paradox of an AI arms race: defending against stronger AI may itself require stronger AI. He pointed to cybersecurity as the clearest case. AI, he said, is reaching superhuman capability in breaking into and compromising computer systems. Future agents may not need physical robot bodies to affect the real world; information systems alone could be enough.

That is why OpenAI believes there is a limited defensive window right now, during which the strongest available models should be used to strengthen critical infrastructure. As agent autonomy rises, Pachocki said, the line between malicious human use of AI and AI taking actions that diverge from human intent may also begin to blur.

He went further, writing that some future agents may negotiate with people, deceive them, or even extort them in pursuit of their own goals. Even so, he rejected the conclusion that competition means companies should accelerate at any cost. Once the risks are understood, he argued, that response does not make sense.

RSI is the next escalation

What raises the stakes from here, in his view, is recursive self-improvement. Pachocki described RSI as a process in which machine intelligence begins to help improve itself: assisting with AI research, refining algorithms, and even improving the computing systems that support AI workloads, all of which feeds into the creation of stronger next-generation models.

If current technical progress continues, he said, machine intelligence participating in its own development is close to a natural outcome. Automated AI research could eventually become central to scientific discovery. OpenAI is also concentrating more of its research around RSI because, in the company’s view, staying at the frontier of AI research will ultimately require automated AI research.

OpenAI says slowing down should remain an option

Pachocki argued that AI scaling should be constrained by the level of confidence researchers have in safety. Existing mechanisms such as the Preparedness Framework and Responsible Scaling Policy should, he said, develop into industry-wide safety thresholds enforced by third-party auditing bodies, governments, or international organizations.

He listed three “north star” goals for OpenAI: building automated AI researchers and using them to solve alignment, turning the scientific and economic growth from highly intelligent machines into human benefit, and making personal AGI available to everyone. Of those, he said the first is the most urgent.

His reason was straightforward. If AI can do work that once required thousands of experts, then a small number of people with access to large-scale compute could gain unprecedented capability and power.

That means the problem is not only whether AI is safe. Pachocki said society will also have to deal with human autonomy, concentration of power, and how human value is preserved in a world where AI can do most work. His conclusion was direct: no AI lab has yet solved alignment and monitoring well enough to keep scaling at maximum speed over the long term. He said, and hopes, that voluntary slowdowns by AI companies will become more common before shared safety standards are in place, and that governments should make international coordination on future AI development one of their highest priorities.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.