Picture the customer service ticket from last month that never got resolved. The AI agent answered instantly. It was courteous. It followed every output rule its compliance team spent a quarter writing. It cited nothing false, invented nothing, defamed nobody. It also never fixed your billing problem, never escalated to a human, and closed the conversation with a cheerful summary of the nothing it had accomplished. According to research released this week, that interaction, not the fabricated court citation or the invented biography, is now the typical enterprise AI failure.

Behind the numbers is ChatSee.ai, a startup that has spent the past year collecting and classifying more than 10,000 grounded examples of enterprise AI failures into a taxonomy of 157 distinct categories. The headline finding, published July 29 through PR Newswire, lands like a correction to two years of industry anxiety: hallucination-related failures accounted for less than 10 percent of the observed events. The largest single bucket, at 31.1 percent, was resolution and escalation breakdowns. And the fastest-growing problem was execution and action failures, up 62 percent relative to the company's Q2 2024 baseline.

The Failure Moved From The Answer To The Action

That shift tracks exactly with what enterprises have been doing to their AI deployments. In 2024, most corporate AI was a chatbot: a system that answers questions. If the answer was wrong, that was the failure, and hallucination testing caught a meaningful share of it. In 2026, the same companies are deploying agents: systems that perform work, call tools, update records, issue refunds, classify transactions. A system that performs work has entirely new ways to fail that have nothing to do with whether its sentences are true.

Its taxonomy reads like a field guide to those new ways: tool-call failures, where the agent invokes the wrong system or invokes the right one incorrectly, and breakdowns across what the researchers describe as the scoping, reasoning and execution phases of a task. An agent can scope a task wrong and confidently do half the job. It can reason correctly and then execute against a stale record. It can complete nine steps of a ten-step workflow and leave the last one hanging without telling anyone. None of those events would trip a hallucination detector, because at no point did the agent say anything false.

Most organizations still evaluate AI risk through hallucination checks, output guardrails and prompt testing. If those account for less than 10 percent of real failures, then the standard enterprise AI safety review is auditing a rounding error while the actual failure surface goes unwatched.

Who Is Telling You This, And What They Are Selling

An honest caveat belongs in the middle of this story, not a footnote. ChatSee is not a university lab. It is a venture-backed startup that raised $6.5 million in June, in a seed round led by True Ventures with First Rays Venture Partners and Seven Hills Ventures participating, to sell exactly the thing this research argues enterprises are missing: a failure intelligence layer for AI agents. Co-founder and CEO Sekhar Sarukkai describes the product as a failure knowledge base that agents reference at the platform level, so that when one agent gets corrected by a human, the correction propagates to every other agent in the system. A company that monetizes agent failures publishing research showing agent failures are rampant and undercounted is marketing, and should be read as marketing.

Two things can be true at once, though. The incentive is obvious, and the direction of the finding matches what independent incidents have shown all year. This site has documented a Gap-owned brand's chatbot recommending sex toys and discussing Nazis without hallucinating a single product fact, a McKinsey enterprise agent compromised in about two hours, and an OpenAI model that escaped a security test and breached a rival company's infrastructure. Not one of those failures was a wrong answer. All of them were wrong actions. The vendor has a motive, but the trendline it is describing has been visible in the incident record for months.

The Number Nobody Budgeted For

Of the two headline numbers, the 31.1 percent figure deserves a closer look than the hallucination number, because it describes the failure mode that costs enterprises money invisibly. A resolution breakdown does not generate a screenshot that goes viral. It generates a customer who asked for help, received a fluent conversation instead, and left. An escalation breakdown is worse: the system was supposed to recognize that a human needed to take over, and did not. Every company that replaced a support tier with an agent this year has some percentage of these happening right now, and by definition the agent is not flagging them, because failing to flag is the failure.

Meanwhile the 62 percent rise in execution and action failures is the forward-looking warning. That number is growing because agent deployment is growing, and it measures the gap between what enterprises think they bought, a digital worker, and what they actually installed, a probabilistic system with write access. E-commerce catalog validation, pricing decisions, transaction labeling, merchant code classification: these are the workloads ChatSee says agents are running in production today. Each one is a place where a silent execution failure becomes a wrong price, a mislabeled transaction, or a compliance problem that surfaces in an audit months later.

What This Means For The Hallucination Industry

An entire cottage industry of AI safety tooling was built on the premise that the model saying false things is the main event. Guardrail vendors, output filters, prompt test harnesses, hallucination leaderboards. If the ChatSee numbers are even directionally right, that industry is defending a shrinking slice of the failure surface while the dominant risks, the unresolved ticket, the missed escalation, the wrong tool call, the abandoned workflow, sail through every check. The chatbot era's safety playbook is being carried into the agent era mostly unmodified, and the failure data says it does not transfer.

None of this makes hallucinations harmless. Courts are still sanctioning lawyers over fabricated citations, and this site's archive of those cases grows monthly. The point is proportion. The failures that get screenshots are under 10 percent of the failures that happen. The other 90 percent look like nothing at all, which is exactly why they compound.

The Verdict

Research built on 10,000-plus logged enterprise AI failures says hallucinations are less than 10 percent of the problem, resolution and escalation breakdowns are 31.1 percent, and execution failures are up 62 percent since mid-2024. The source is a vendor selling the cure, so apply salt. But the shift it describes, from wrong answers to wrong actions, matches every major AI incident this site has documented since spring.