research

AI agent failure statistics: what stalls an agent

this+that team
AI agent failure statistics: what stalls an agent

Agent adoption is close to universal. Results are rare. Several organizations with nothing in common commercially have now measured that gap.

Every study here measures whether agents fail. None isolates why down to a single cause, and several name different reasons. Where an analyst names a reason, we say which analyst.

Key takeaways

How often agent projects fail

Gartner’s June 2025 prediction is the most quoted number in the category, and it’s worth reading precisely. It says over 40% of agentic AI projects will be canceled by the end of 2027, and it attributes that to escalating costs, unclear business value and inadequate risk controls. It doesn’t say the agents can’t do the work.

The same release puts the number of vendors with substantial agentic capability at about 130, out of thousands marketing it, and gives the rest a name: agent washing, the rebranding of AI assistants, robotic process automation and chatbots. So some share of any failure statistic is a product that was never an agent.

Gartner’s own framing of the capability question is direct. Current models, it says, lack the maturity and agency to autonomously achieve complex business goals or follow nuanced instructions over time.

For adoption figures against those failure rates, we’ve written up two independent studies on agent adoption separately.

What an agent finishes on real work

The most useful number here isn’t a survey. Scale AI’s Remote Labor Index takes real remote work projects, the kind actually bought and paid for on freelance platforms, and hands the same briefs to frontier agent frameworks. Across the frameworks evaluated, the maximum automation rate is 2.5%.

Writers misquote that figure often, usually as 4%. To put it plainly, 2.5% is the ceiling rather than the average, and Scale AI measures it against paid professional output rather than against a benchmark the agents trained on.

A survey tells you what leaders believe about their agents. An index like this tells you what the agents did when someone checked.

The reasons on record

Three organizations have put a reason in writing, and they don’t agree with each other.

Gartner names cost, unclear value and risk controls, plus model maturity.

Deloitte names governance. In its 2026 State of AI in the Enterprise survey of 3,235 IT and business leaders across 24 countries, 21% say they have a mature governance model for agentic AI, leaving roughly 80% without clear boundaries on which decisions an agent can make alone, real-time monitoring, or audit trails. The same survey has 74% expecting at least moderate agent use by 2027, so the deployment curve is running ahead of the controls.

Camunda and Okta, covered in our adoption write-up, name identity and permissions.

Every one of those is a governance or economics answer. None of them is about whether the agent knows anything about the company it works for.

The variable nobody surveys

No study we can find measures agent failure against the quality of the context the agent was given. What exists instead is a long record of how badly organized company knowledge already is, gathered long before agents existed.

McKinsey put nearly 20% of the workweek into looking for internal information or tracking down colleagues, alongside 28% spent on email. McKinsey ran that research in 2012, more than a decade before anyone deployed an agent. Companies were already losing a fifth of the week to hunting for information.

Panopto’s YouGov survey of 1,001 US adults is the sharpest of these. 42% of institutional knowledge is unique to the individual who holds it. 60% found it difficult, very difficult or nearly impossible to get information vital to their job from a colleague. The average employee spends 5.3 hours a week waiting for that information. Scaled up, the study puts the annual loss at $47 million for a business averaging 17,700 employees.

Those two findings give the shape of the problem. Almost half of what a company knows sits with one person. Nobody ever wrote it into a system. An agent reads systems. So the agent starts every task missing the same information a new hire spends months collecting, and no amount of model capability closes that particular gap.

No study yet tests that link directly, so treat the last paragraph as our argument rather than a finding.

What follows for anyone deploying agents

Three things the numbers above support directly.

Check what the agent can read before you judge what it can do. A 2.5% ceiling on open-ended remote work is a fact about agents working without your context, not a fact about agents.

Expect the post-mortem to name the wrong cause. A review records cost and governance, because those are the things it can see.

The knowledge problem is older than the agent problem. McKinsey measured it in 2012 and Panopto in 2018. Anything you deploy now inherits it.

The missing 95%

The most quoted number in this category isn’t above, so here’s why.

In August 2025 the NANDA project at the MIT Media Lab published The GenAI Divide: State of AI in Business 2025, and the press turned one line of it into a headline: 95% of enterprise generative AI pilots produce no return. Forbes, Fortune and Inc. all ran it. Markets moved on it.

Two things keep it out of this post.

You can’t read it. The report’s public link now redirects to the NANDA group page, where the reports link asks you to request access. That gating is old news: writing on August 26, 2025, days after publication, Futuriom’s R. Scott Raynovich already found the report behind a request form. Every public copy today is a third party’s re-host, so there’s no primary to cite.

The number didn’t survive review. Kevin Werbach, Wharton professor and chair of legal studies and business ethics, read the document repeatedly and couldn’t find support for the 95% anywhere in it. The report does contain a 5% figure, for custom enterprise tools reaching what it calls a marked and sustained productivity or P&L impact. That’s a much narrower claim, and falling short of it explicitly doesn’t mean zero return. Werbach’s conclusion was that NANDA should release the underlying data or retract the report. Raynovich catalogued the rest: unlabeled axes, sample sizes the report never gives, and a jump from a subset to a headline about everyone.

Worth knowing about the source, too. NANDA builds decentralized agent infrastructure on top of Anthropic’s Model Context Protocol and Google’s Agent-to-Agent protocol, which gives it an interest in a finding that current enterprise AI approaches are failing.

None of that makes the underlying worry wrong. Plenty of pilots do stall, which is what Gartner’s cancellation figure above measures. It makes 95% a number to stop repeating.

How we sourced this

Every figure links to the organization that ran the research. We don’t cite aggregators or statistics roundups, and where a number is commonly attributed to the wrong source, we say so.

One more correction. The 4% figure often attached to the Remote Labor Index doesn’t appear in it. The number is 2.5%.

Frequently asked questions

What percentage of AI agent projects fail? Gartner predicts over 40% of agentic AI projects will be canceled by the end of 2027, attributing that to escalating costs, unclear business value and inadequate risk controls.

How much work can an AI agent actually complete? On Scale AI’s Remote Labor Index, which uses real paid remote work projects, the best frontier agent framework completed 2.5% of them.

Why do AI agents fail? The reasons on record differ by who you ask. Gartner names cost, unclear value, risk controls and model maturity. Deloitte names governance, with only 21% of organizations reporting a mature agentic AI governance model. Camunda and Okta name identity and permissions.

Are most AI agent products actually agents? Gartner puts the count with substantial agentic capability at about 130 out of thousands, and calls the rest agent washing.