Pentagon AI Targeting Tempo Outpaces Evaluation Capacity
The Defense Department is deploying frontier AI models on 30-day timelines while the infrastructure to independently test them remains years behind.

The Speed-Safety Gap in Military AI
The Pentagon's rapid adoption of generative AI for military operations has created a dangerous mismatch: deployment timelines measured in weeks, evaluation capacity measured in years. Operation Epic Fury demonstrated this tension at scale, with Central Command using Anthropic's Claude through Palantir's Maven Smart System to generate roughly 1,000 targets in the first 24 hours—more than double the tempo of the 2003 Iraq invasion's opening phase. Project Maven now claims capacity for 5,000 targeting decisions in a single day.
Yet the most recent independent testing data dates to 2024, when Maven achieved 60 percent object identification accuracy compared to 84 percent for human analysts in 18th Airborne Corps evaluations. No one outside the program can verify whether that gap has closed, widened, or reversed during the 38-day campaign that reached 13,000 total strikes, according to Pentagon data.
Why it matters
The Pentagon is institutionalizing a pattern where operational deployment runs years ahead of independent verification. When machines generate targeting recommendations faster than humans can evaluate them, review becomes a formality rather than a safeguard. Without the workforce, compute infrastructure, and evaluation methodologies to independently assess what vendors deliver, the military risks fielding systems whose failure modes remain unmapped until something goes catastrophically wrong.
Policy Acceleration Without Capacity Building
The Pentagon's AI Acceleration Strategy, released in January, established a Barrier Removal Board with authority to waive non-statutory requirements across testing, contracting, and hiring. It mandated adoption of frontier models within 30 days of release. National Security Presidential Memorandum 11, signed June 5, reinforced this approach by calling for removal of "unnecessary barriers to rapid deployment."
The problem: waiving processes cannot solve a capacity problem. The department set July as the deadline for initial demonstrations across priority AI deployments but missed its own target without public explanation.
Five Critical Shortfalls
The Pentagon faces interconnected gaps that worsen as deployment accelerates:
Workforce: Security clearances alone take nearly a year before engineers can begin meaningful work. The department lacks machine learning engineers, acquisition officers who can write consumption-based contracts, and operators trained to catch fluent-sounding fabrications from large language models.
Evaluation infrastructure: The department relies on vendor-provided benchmarks rather than owning independent testing methodologies. The Responsible AI Office, which absorbed governance work in practice, lost staff to buyout programs and return-to-office mandates. The Director of Operational Test and Evaluation saw its staff cut roughly in half in May 2025 and dropped nearly 100 programs from oversight.
Compute and data: Pentagon-owned facilities run six-to-eighteen-month accreditation timelines. The department often cannot use data from its own operations to train AI because vendor contracts don't secure those rights.
Authority to operate: Single authorizations take twelve to eighteen months—a cadence built for systems that change rarely, not models that update every few weeks. A mid-2026 survey found no decision-makers reporting authorization timelines under six months.
Historical Precedent
In 2003, Patriot air defense batteries operating in largely automatic mode shot down a British Tornado and a Navy F/A-18 over Iraq, killing three aircrew. Operators had been conditioned to trust system outputs unconditionally. A Defense Science Board task force found that when the system's assumptions stopped holding, operators had no way to question what sensors told them. That was a mature, narrowly scoped system with decades of testing—unlike the large language models now being fielded on 30-day timelines.
What Needs to Change
Congress authorized three provisions in the FY26 National Defense Authorization Act: a testing sandbox, a cross-functional evaluation team, and a governance subcommittee for AI oversight. None has been established despite deadlines passing. The department needs dedicated leadership with budget authority, metrics that count actual outputs—engineers hired and cleared, compute delivered at each classification level, data rights clauses executed—rather than milestones.
Frontier model reliability varies enormously by task. On the HalluLens benchmark, these models hallucinate on roughly 27 to 85 percent of short factual queries and invent answers about nonexistent entities up to 94 percent of the time. When models generate articulate recommendations complete with precise coordinates and prioritized targets, they risk inspiring excessive operator deference—automation bias well-documented in human-machine teaming research.
The details in this analysis were first reported by War on the Rocks.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call

