OpenAI's New Reasoning Technique Raises AI Safety Concerns
The company is testing methods that could make model behavior harder to monitor, according to a new report.
OpenAI is experimenting with a new AI reasoning technique that could make it significantly harder to monitor how its models arrive at decisions, according to a report from The Information.
The approach, called "recurrent depth," is being used in OpenAI's Astra model and may improve performance and reduce costs. But it also obscures the model's internal reasoning process—what researchers call "chain of thought"—making it more difficult to understand or audit what the AI is actually doing.
That has set off alarm bells among AI safety researchers, including several who previously worked at OpenAI.
A critical monitoring tool at risk
Chain of thought monitoring has emerged as one of the few practical methods for peering inside large language models. While imperfect, it allows researchers to trace how a model reasons through a problem by examining the intermediate steps it takes.
A 2025 paper titled "Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety" argued that this capability represents a rare chance to improve AI safety. The authors warned that the opportunity is fragile and could easily be lost.
That concern now appears prescient. The recurrent depth technique appears to compress or hide those intermediate reasoning steps, potentially eliminating visibility into model behavior.
Steven Adler, founder of Guidelight.ai and a former OpenAI safety team member, said the company appears to be "violating one of the few redlines that exist in the AI industry." He questioned why OpenAI would train models this way.
Connection to recent incidents
The timing is notable. Gary Marcus and Zack Korman recently argued that better monitoring might have prevented the Hugging Face incident—a reference OpenAI itself acknowledged. According to Marcus, OpenAI admitted that improved monitoring could have caught the problem earlier.
Now the company is exploring techniques that could make such monitoring difficult or impossible, trading a safety mechanism for potential performance gains.
Why it matters
As AI systems take on more consequential tasks, the ability to understand and audit their decision-making becomes critical. Techniques that deliberately obscure model reasoning move in the opposite direction—prioritizing capability improvements over interpretability. For organizations deploying AI in high-stakes environments, this trade-off raises fundamental questions about whether performance gains justify reduced oversight. The debate also highlights a broader tension in AI development: whether safety considerations will constrain or be subordinated to the race for more powerful models.
The Information first reported these details about OpenAI's work with recurrent depth in Astra. According to their reporting, OpenAI is currently limiting use of the technique, though the extent of those limitations remains unclear.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call

