Skip to main content

The Toughest Exam in E-Commerce: Why AI Agents Fail Real-World Tasks

A new benchmark, RealReplicaBench, tests AI agents in realistic e-commerce scenarios. Most models fail, revealing a gap between answering questions and completing actual work. The results challenge our expectations of AI.

A New Kind of Test for AI

Testing AI used to be simple. A few years ago, when AI was just a chatbot, you could grade it on writing, math, or code. The result was a score, a number that supposedly told you how smart the model was. But now AI has stepped out of the chat box. It can use tools, browse the web, handle files, and even run tasks across different systems. It's become an agent, and that changes everything.

You've probably seen AI on your phone, helping you shop on Taobao or adding things to your cart. But behind the scenes, on the merchant's side, AI is doing much more complex work. It's managing suppliers, processing orders, and coordinating logistics. And according to a new benchmark called RealReplicaBench, it's not doing a great job.

Why Most AI Models Fail the Real-World Test

RealReplicaBench was created by the team behind Accio Work, an AI agent designed specifically for e-commerce. It puts AI models through 107 real business tasks, and the results are sobering. Not a single model scored above 60—the passing mark. The highest score was 56.1, achieved by Claude Opus 5.

Thirteen models were tested, including GPT, Claude, GLM, Qwen, Deepseek, and Gemini (which, unsurprisingly, came in last). Even the best performers struggled. For example, Claude Opus 5 had a pass rate of 66 out of 107 tasks when using the Accio Work framework, but only 60 with OpenClaw and 65 with PI. The Accio framework generally seemed to help, but the overall picture was bleak.

From Answering Questions to Doing the Job

Traditional benchmarks are like a written driving test. You can get points for partially correct answers, even if you don't finish the problem. But in the real world, a half-done task is often useless. If a supplier isn't chosen correctly, the whole procurement process falls apart. If the price is wrong, the product can't be listed. If the logistics route doesn't meet the deadline, the booking is meaningless.

The team at Accio Work realized this and designed RealReplicaBench to be more like a road test. It's not enough to know when to signal a turn; you have to drive from point A to point B without crashing. In their philosophy, a task is only complete when the result can be directly used by the next step in the workflow. If 80% of the work is done but the critical 20% is left for a human to fix, it's a zero.

Building a Fake World to Test Real Work

To test real work, they had to create a realistic environment. RealReplicaBench doesn't just give the AI a text prompt; it replicates the entire business context. It includes front-end UIs, browser operations, command-line interfaces, APIs, file systems, and backend states. The AI has to interact with this environment as if it were a real business.

One task involves supplier procurement. The agent must sift through about 300 noisy emails to extract the real purchasing requirements, then select suppliers, tag emails, draft replies, and set up a Kickoff calendar. It's not about summarizing emails; it's about finding the truth in the noise and turning it into action.

Another task is even more complex: turning 5,383 customs records into a cross-system procurement control tower. The model has to aggregate data by category and supplier, filter the top three suppliers based on policy, and then create a dashboard in Google Workspace, a folder in Box, and tasks in Jira. All these pieces have to be linked correctly. If a dynamic ID is wrong, the entire chain breaks.

Verifying Completion: No More Self-Reported Success

One of the biggest challenges in testing agents is knowing when a task is truly done. A chatbot can just say, "I've finished," and you might believe it. But an agent's output is just part of the action. So RealReplicaBench uses a verifier that checks the final state of the environment, not the agent's report.

For example, in a logistics task, the agent must plan a route for a shipment from China to the US, considering sea freight, trucking, final delivery, insurance, customs, bonds, and platform fees. It has to exclude options that take more than 30 days or have invalid port connections. Then it must actually book the sea freight and verify the shipment. The final check isn't a text description; it's the actual Shipment ID that was generated.

This approach means that a task is only considered complete if it leaves a verifiable trace in the environment. That's why the grading is so strict. It's not about looking like an answer; it's about being an answer that works.

What This Benchmark Reveals About AI and Work

RealReplicaBench is not just a test; it's a window into how the Accio Work team thinks about AI. They believe that an agent's value lies not in its ability to answer questions but in its ability to complete work. This philosophy is reflected in the benchmark's design. The tasks are derived from real business data—about 1.6 million conversations, 200,000 execution traces, and 2,000 high-value workflows—which were then distilled into 107 reproducible tasks.

The team plans to keep adding new tasks and updating the environment. They also want to use the benchmark for model evaluation, training optimization, and even model routing, which means choosing the best model for each task based on effectiveness and cost.

This benchmark isn't just for e-commerce merchants, though it will help them choose better AI tools. It's also valuable for the broader tech community. As AI moves into business workflows, the ability to build a harness—the infrastructure that supports the agent—will be just as important as the model itself.

For users, what matters is not an agent that's "almost right" but a system that can take over a job and deliver results. The standard for AI is shifting from "what does it know?" to "can it get the job done?" RealReplicaBench is turning that simple question into a rigorous, repeatable test. It's a tough exam, but it's one that AI needs to pass if it's going to be trusted with real work.

Share this article:

Comments (0)

No comments yet. Be the first to comment!