Experimental AI Model Reasons Without Generating Lengthy Text
BDH-CQ solves logic puzzles by processing information internally rather than writing out each step, potentially reducing computing costs.

An experimental AI system demonstrates that artificial intelligence can tackle reasoning problems without generating lengthy chains of text for every intermediate step, according to research submitted to arXiv.org in August.
The model, called BDH-CQ (Dragon Hatchling), solved nearly 30 percent of puzzles on the ARC-AGI-1 benchmark when given two attempts—without writing out its thinking process in words. The approach contrasts sharply with current mainstream AI systems that rely on chain-of-thought reasoning, where models generate extensive intermediate text before reaching conclusions.
How the model works differently
Most AI systems keep training examples visible as they work through new problems. BDH-CQ instead uses each example to update a fixed-size internal memory that remains constant regardless of how many examples it processes. When a new puzzle arrives, the model works through it entirely internally.
"Nothing in between ever converts into language," says Zuzanna Stamirowska, CEO of AI company Pathway and one of the researchers on the project. The model processes information without turning each reasoning step into tokens—the basic units of text that AI systems generate.
Why it matters
Every token an AI model generates requires computing power, making verbose reasoning chains both slower and more expensive to produce. The researchers estimate each BDH-CQ puzzle query costs approximately $0.00070 to run—roughly one-eleventh the cost of GPT-5.6 Luna on the same benchmark, though the two figures were calculated using different methods. As organizations deploy AI systems at scale, these efficiency gains could translate into substantial cost savings and faster response times.
Performance and limitations
The model showed uneven performance across different puzzle types. It handled tasks involving rotation and movement of shapes more successfully than problems requiring color changes or complex rule combinations. On more difficult puzzles involving ordering and nesting, the system improved after seeing examples of similar difficulty.
The ARC-AGI-1 benchmark tests whether AI systems can infer visual rules from limited examples and apply them to new puzzles—tasks designed to resemble human cognitive assessments.
Expert perspectives
Yuntian Deng, a computer scientist at the University of Waterloo, called the efficiency result "interesting" but noted it doesn't prove BDH-CQ's architecture is superior to other approaches. He emphasized that more testing is needed to separate the effects of the model's design from its training methodology.
Deng also highlighted a trade-off: without written reasoning steps, the model becomes harder to inspect and understand. However, he noted that even explicit chains of thought may not accurately represent how models actually reach their answers.
Jonas Geiping, a machine learning researcher at the ELLIS Institute Tübingen and Max Planck Institute for Intelligent Systems, observed that BDH-CQ was built specifically for ARC-style problems, making direct comparisons with general-purpose AI systems difficult. Still, he praised the approach and noted the model can handle test problems without requiring retraining.
The research raises fundamental questions about whether language is necessary for all AI reasoning. "Human language is useful for communicating reasoning, but it need not be the most efficient representation for every intermediate computation," Deng said.
The study has not yet undergone peer review. Details were first reported by Science News, based on a paper by B. Engdahl and colleagues.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call