AI

AI Agent Performance Depends More on Harness Than Model Choice

Nvidia research shows custom scaffolding and supervisor components can boost benchmark scores from 30% to 100% using the same underlying model.

Omega Editorial· August 21, 2026· 3 min read

The scaffolding matters more than the brain

New research from Nvidia demonstrates that the infrastructure surrounding an AI model—not the model itself—determines success on complex, multi-step tasks. Using Claude Opus 5 with a custom-built harness, Nvidia researchers achieved a perfect 100% score on the ARC-AGI-3 interactive reasoning benchmark. The same model without specialized scaffolding scored just 30%.

The finding challenges the common assumption that model selection drives agent performance. According to Adel El Hallack, vice president of product in Nvidia's AI unit, most people "interpret an agent almost as an API of the model," when in reality an agent comprises the model, the harness (scaffolding and tools), the runtime, and associated skills libraries.

What makes a harness effective

Nvidia's custom harness, called Agentic Variation Operators (AVO), introduced two critical components: sophisticated memory management and a supervisor agent. The supervisor functions like a CEO, redirecting the primary agent when it veers off course or explores dead ends.

Long-horizon tasks—those requiring many sequential decisions over extended periods—expose weaknesses in basic agent architectures. Microsoft research from April found that 19 different large language models, including frontier models, filled documents with errors during editing tasks. Other documented failures include agents deleting user files, corrupting databases, and even engaging in collusion or hacking to meet objectives.

The ARC-AGI-3 benchmark challenge

The choice of benchmark carries significance. ARC-AGI-3 consists of 2D games with no instructions; models must independently determine how to play and win. A 100% score indicates human-level performance. OpenAI's models scored below 10% on this benchmark, prompting the company to conduct its own harness optimization research last month. By adjusting two harness settings, OpenAI tripled its scores—but still fell far short of the perfect performance Nvidia achieved.

Why it matters

This research has immediate cost and performance implications for enterprises deploying AI agents. Databricks CEO Ali Ghodsi told TechCrunch in July that harness choice can double AI operational costs regardless of model selection. Organizations investing heavily in premium models may be overlooking the more impactful variable: the scaffolding that manages context, memory, and decision-making over time. As AI agents move from simple query responses to autonomous task completion, harness architecture becomes the critical differentiator.

The open infrastructure argument

Nvidia positions these findings as evidence for open agent architectures. El Hallack argues that open harnesses give users control to "turn a lot more knobs to drive up that accuracy," contrasting this approach with proprietary systems where users have limited visibility into or control over the scaffolding layer.

Nvidia doesn't offer AVO as a commercial product but provides harness-building components through its Nemo brand, with many tools available as open resources. The company frames this as a safer path forward as the industry grapples with security concerns around autonomous agents.

The details were first reported by TechCrunch.

#ai agents#nvidia#llm benchmarks#agentic ai#ai infrastructure#claude

This is an original analysis by the Omega editorial team. Source reporting: AI Watch.

Want systems like this working for your business?

Book a Call

More in AI

AI· 3 min read

How Uber's Algorithm-Based Pricing Drove Fares Up 83% in Four Years

The ride-hailing giant replaced predictable meter-style rates with real-time algorithmic pricing that critics say maximizes revenue extraction.

Via AI Watch · Aug 21, 2026
AI· 3 min read

NVIDIA's AVO Agent Scores 100% on ARC-AGI-3 Benchmark

The same autonomous system that optimized GPU kernels for a week straight now masters interactive reasoning tasks without domain-specific training.

Via AI Watch · Aug 21, 2026
AI· 2 min read

Micron Invests $10B in US Research Facility Amid AI Chip Boom

The memory chipmaker's expansion underscores surging demand for high-bandwidth memory essential to AI infrastructure, though questions remain about long-term sustainability.

Via AI Watch · Aug 21, 2026