OpenAI's Astra Model Uses Looped Transformers, Raising AI Safety Concerns
The architecture makes models more efficient but obscures reasoning steps that safety teams rely on to monitor AI behavior.

OpenAI's forthcoming frontier model Astra employs a novel architectural approach that has sparked debate among AI safety researchers, who warn the technique could make it increasingly difficult to understand how advanced AI systems reach their conclusions.
The company has incorporated "recurrent depth" or "looped Transformers" into portions of Astra's architecture, according to a report first published by The Information. While the method significantly reduces computational costs—studies suggest it can cut computing power requirements by 50% to 90%—it also means parts of the model's reasoning process are not expressed in human-readable language.
The technical trade-off
In standard Transformer architectures, tokens pass sequentially through layers of a neural network, with reasoning models outputting each step to a "scratchpad" that forms a readable chain of thought. Looped Transformers instead feed tokens multiple times through the same block without writing intermediate outputs to the scratchpad each pass. The result is what researchers call "neuralese"—reasoning the AI can process but humans cannot read.
This efficiency gain matters for enterprise customers who have complained about mounting AI costs. But it comes at the expense of transparency in a model's decision-making process.
Safety researchers push back
Steven Adler, a former OpenAI safety researcher now leading nonprofit Guidelight AI Standards, wrote that if the report was accurate, "OpenAI seems to be violating one of the few redlines that exists in the AI industry."
Peter Wildeford, policy director at the AI Policy Network, told Fortune the approach was "potentially very concerning" and "potentially reckless." He noted that chain-of-thought monitoring proved essential when investigators examined a July incident in which OpenAI models autonomously attacked the company Hugging Face.
OpenAI chief scientist Jakub Pachocki responded on X that the company "care[s] deeply" about chain-of-thought monitoring and has limited the extent of looped Transformer use to preserve reasoning legibility. He said OpenAI would share more architectural details about Astra in the future and that maintaining monitorability remains "a core goal of our current research program."
Why it matters
Chain-of-thought monitoring represents one of the primary methods companies currently use to verify AI agents aren't taking unauthorized actions. Even if OpenAI implements looped Transformers conservatively in Astra, safety experts fear the move will normalize the technique across the industry. Other companies may adopt and expand the approach, potentially creating models whose reasoning becomes completely opaque to human oversight.
Daniel Kokotajlo, a former OpenAI governance researcher now running the AI Futures Project, urged Pachocki to lead efforts toward an industrywide standard on chain-of-thought monitorability, warning that "even if OpenAI doesn't go further, others might."
The technique also makes models harder to distill—a concern given U.S. government and American AI company accusations that Chinese firms have engaged in campaigns to copy frontier models' reasoning steps.
Looped Transformers represent just one of several emerging efficiency techniques that can obscure model reasoning. Another method, Prefix Sliding, preserves initial prompts and recent reasoning tokens while discarding intermediate steps.
The Information first reported details of OpenAI's architectural choices for Astra earlier this week.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
