AI Agent Performance Depends More on Harness Than Model Choice
Nvidia research shows custom scaffolding and supervisor components can boost benchmark scores from 30% to 100% using the same underlying model.

The scaffolding matters more than the brain
New research from Nvidia demonstrates that the infrastructure surrounding an AI model—not the model itself—determines success on complex, multi-step tasks. Using Claude Opus 5 with a custom-built harness, Nvidia researchers achieved a perfect 100% score on the ARC-AGI-3 interactive reasoning benchmark. The same model without specialized scaffolding scored just 30%.
The finding challenges the common assumption that model selection drives agent performance. According to Adel El Hallack, vice president of product in Nvidia's AI unit, most people "interpret an agent almost as an API of the model," when in reality an agent comprises the model, the harness (scaffolding and tools), the runtime, and associated skills libraries.
What makes a harness effective
Nvidia's custom harness, called Agentic Variation Operators (AVO), introduced two critical components: sophisticated memory management and a supervisor agent. The supervisor functions like a CEO, redirecting the primary agent when it veers off course or explores dead ends.
Long-horizon tasks—those requiring many sequential decisions over extended periods—expose weaknesses in basic agent architectures. Microsoft research from April found that 19 different large language models, including frontier models, filled documents with errors during editing tasks. Other documented failures include agents deleting user files, corrupting databases, and even engaging in collusion or hacking to meet objectives.
The ARC-AGI-3 benchmark challenge
The choice of benchmark carries significance. ARC-AGI-3 consists of 2D games with no instructions; models must independently determine how to play and win. A 100% score indicates human-level performance. OpenAI's models scored below 10% on this benchmark, prompting the company to conduct its own harness optimization research last month. By adjusting two harness settings, OpenAI tripled its scores—but still fell far short of the perfect performance Nvidia achieved.
Why it matters
This research has immediate cost and performance implications for enterprises deploying AI agents. Databricks CEO Ali Ghodsi told TechCrunch in July that harness choice can double AI operational costs regardless of model selection. Organizations investing heavily in premium models may be overlooking the more impactful variable: the scaffolding that manages context, memory, and decision-making over time. As AI agents move from simple query responses to autonomous task completion, harness architecture becomes the critical differentiator.
The open infrastructure argument
Nvidia positions these findings as evidence for open agent architectures. El Hallack argues that open harnesses give users control to "turn a lot more knobs to drive up that accuracy," contrasting this approach with proprietary systems where users have limited visibility into or control over the scaffolding layer.
Nvidia doesn't offer AVO as a commercial product but provides harness-building components through its Nemo brand, with many tools available as open resources. The company frames this as a safer path forward as the industry grapples with security concerns around autonomous agents.
The details were first reported by TechCrunch.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call

