AI Agents Complete 62% of Commerce Tasks in New Benchmark Test
Alibaba's CommerceAgentBench reveals where automation succeeds in real business workflows—and where it still fails.

AI Agents Complete 62% of Commerce Tasks in New Benchmark Test
The strongest frontier AI models successfully complete just 61.7% of real-world commerce tasks when measured on actual outcomes rather than conversational ability, according to new benchmark testing from Alibaba's Accio team.
The finding challenges the industry's focus on which model reasons best and redirects attention to a more practical question for businesses: which specific workflows can safely be automated, and which still require human oversight.
Why it matters
As thousands of small businesses deploy AI agents for sourcing, logistics, and customer service, individual errors risk becoming correlated failures across supply chains. A benchmark that measures task completion—not just plausible responses—gives companies a framework for precision delegation: knowing exactly where automation is ready to scale and where human judgment remains essential.
Testing outcomes, not eloquence
CommerceAgentBench, now open source on GitHub, contains 107 end-to-end tasks drawn from real e-commerce operations across procurement, logistics, product listings, fulfillment, and after-sales service. The test set was assembled from data covering 10 million active small-business users, 1.6 million conversations, and 200,000 execution traces.
Unlike traditional AI benchmarks that measure reasoning or coding ability, CommerceAgentBench grades whether the work actually got done correctly. Did the product listing go live with accurate attributes? Did freight move on a viable route? Did the customs form get filed properly?
Joshua Stancle, who runs Clean Saint from Los Angeles using AI for sourcing, marketing, and customer support, described the experience as having "ten of me." The challenge, as Alibaba's testing reveals, is knowing which of those ten versions can be trusted without supervision.
Where agents struggle
Failures clustered in predictable areas. Agents had difficulty spotting payment anomalies buried in long email threads, calculating landed costs with multiple variables, resolving after-sales disputes when documents contradicted each other, and routing multi-leg shipments.
One unexpected finding: no single model dominated across all categories. The model that ranked first on request-for-quote work and market research fell behind on claims settlement and listing compliance, where a different model led. A third excelled at publishing products and handling returns.
This category-specific performance means general reasoning scores offer little guidance for selecting the right system for a particular commercial task.
Precision delegation in practice
For businesses, the 62% completion rate is simultaneously encouraging and cautionary. Workflows like supplier comparison and routine listing work can now be safely automated. But unusual compliance questions, complicated negotiations, and the exceptions that comprise more of commerce operations than expected still require human involvement.
The gap between 62% and full reliability represents thousands of daily failures across real businesses—mistakes that matter when a customer receives the wrong product or freight sits idle because routing went wrong.
Alibaba President Kuo Zhang and the Accio team argue that every industry needs similar outcome-based benchmarks built by practitioners who understand what failures cost in their specific domains. These benchmarks should remain open source so buyers can verify vendor claims independently.
Authority will transfer to AI agents one workflow at a time, as each earns trust through measured performance. The details were first reported by Fortune.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call