OpenAI’s chief scientist, Jakub Pachocki, said that no artificial intelligence laboratory has solved alignment and monitoring problems to a sufficient level to sustain indefinitely the current pace of advancement in systems. In an essay published on Sunday (6), he advocated voluntary slowdowns in AI development when safety safeguards do not keep pace with the increase in capabilities.

Pachocki also called for international coordination on the development of advanced systems to become a priority for governments. In the researcher’s assessment, voluntary commitments currently adopted by labs should evolve into widely required safety limits, with oversight by independent auditors, governments, or international institutions.

The position comes from within OpenAI itself, which continues to invest in expanding the capabilities of its models and in systems capable of participating in the development of new AIs. Pachocki said the company may even unilaterally halt new scaling stages when it deems necessary, but acknowledged that decisions made by a single company would not be enough to reduce the risks of a competitive race.

Monitoring advanced models is getting harder

In the essay, called An Alien Mind, Pachocki points to alignment — making an AI act in accordance with human goals and values — as one of the main problems still open in the field.

One of the methods used to evaluate advanced systems is monitoring the so-called chain of thought, the reasoning process verbalized by the model. According to Pachocki, internal evaluations indicate that the ability to rely on that mechanism is decreasing as systems become more sophisticated.

Among the factors cited are more complex operating environments, greater interaction of agents with people, tools, and other AIs, and a growing capacity of the models themselves to control or alter the way they express their reasoning. Newer systems are also able to perform complex tasks without necessarily externalizing the entire process verbally, reducing the visibility available to researchers.

Pachocki relates this problem to the advance of recursive self-improvement, a scenario in which AI systems come to play an increasingly large role in the research and development of subsequent models. He said he expects the current pace of progress to continue in that direction and to be able to produce new leaps in capability in the coming years.

OpenAI intends to continue developing systems capable of acting as automated researchers and to use them also to improve its own safety techniques. For Pachocki, however, expanding capabilities without equivalent advances in alignment and monitoring should not be treated as a collective goal of the industry.

The proposal is to combine technical progress with coordinated reductions in the pace of development when confidence in safeguards is insufficient. Pachocki said he expects voluntary pauses of this kind to become part of the strategy of the leading labs until more robust and shared safety criteria exist.

More from Radar