What survives the pilot
Almost everyone is running an AI agent somewhere. Almost nobody is running one in production. This issue is about the difference — which first tasks make it through, which die in the pilot, and what separates them.
An intent-to-execution ratio of roughly eight to one, and a further five-fold drop between "in production" and "integrated across business functions". The interesting question is not why the technology fails. It is which tasks make it all the way down this chart, because a substantial number do.
Six numbers
Each traced to a dated sourceThe AI employee market talks almost exclusively about capability, which is odd, because capability is not what most buyers are deciding. They are deciding what to hand over first. Get that choice right and a deployment compounds. Get it wrong and it joins the eighty-something percent that never leave the pilot, having spent real money to find out.
Abandonment is getting worse, not better
The intuitive expectation is that the pilot-to-production gap closes as the technology matures and organisations learn. The data says the opposite. The share of companies abandoning most of their AI initiatives went from 17% to 42% in a single year, and the average enterprise scrapped 46% of its pilots before they reached production.
Two readings are available and both are probably true. Organisations are attempting harder things than they were a year ago, so more attempts fail. And organisations have become willing to kill projects they would previously have kept alive — which, if you are the one paying, is progress rather than decline.
A failed agent project costs about $340,000 before anyone calls it. The expensive decision is made in the first week, when the task is chosen.
Sources: S&P Global abandonment data and the 46% pilot-scrap figure via Olakai, Feb 2026 · $340k per failed project via Digital Applied, Mar 2026 · Gartner and WEF figures via Hendricks, Mar 2026.
Narrow survives. Broad does not.
A March 2026 survey of 650 enterprise technology leaders found 78% running at least one pilot and only 14% with an agent scaled to organisation-wide operation. Of the stalled deployments, 72% had been stuck for more than six months with no resolution path. Comparing the two groups gives the clearest scoping rule in the current evidence.
Deployments that scaled started with an agent scoped to a single, well-defined task with a measurable output — a document classifier, a data-enrichment step, a routing decision. Deployments that stalled attempted multi-function agents covering broad task categories from the beginning, producing a combinatorial explosion of edge cases that no evaluation harness could cover and no dashboard could attribute.
The successful group expanded scope only after the narrow version held stable for 90 days or more. That is the whole rule, and it is unglamorous enough that most teams skip it.
Sources: 650-leader survey, scoping comparison and the 90-day rule, AI Agent Scaling Gap, Mar 2026 · median time-to-value and ship/kill discipline, ClarityArc, Jul 2026.
Where the demand actually concentrated
Function-level adoption tells a more useful story than organisation-level adoption, because agents are deployed against a specific job to be done rather than against a company. The distribution is a barbell: customer service and software engineering are saturated relative to legal and HR.
Customer service is the clear leader at 62% of enterprises, and it is also the only function with published outcome data granular enough to be worth reading. That data is more equivocal than the adoption number suggests.
Sources: function-level adoption, deflection, CSAT delta and intervention rate, AI Agent Adoption 2026, Apr 2026. Small-business figures from the same source: 14% in production, 38% piloting, 0.7 agents on average.
Nobody hands over a whole job
The most useful measurement of delegation in the wild comes from software engineering, where usage is instrumented more thoroughly than anywhere else. Developers now use AI in roughly 60% of their work. They can fully delegate 0–20% of tasks. And 86% treat AI-generated output as a starting point rather than a final answer.
Those three numbers describe partial delegation, and partial delegation is the normal case rather than a transitional stage. A product that assumes the whole job transfers is mispricing what it is selling and mis-designing where the human sits.
Sources: Anthropic, 2026 Agentic Coding Trends Report, as compiled in AI Agent Statistics 2026, May 2026 · consistent with the delegation-gap findings in Issue 03, section 06.
Building it yourself fails three times as often
Across an analysis of more than a hundred agent deployments, purchasing from a specialised vendor or building through a partnership succeeded roughly 67% of the time. Internal builds succeeded about 22% of the time. We are a vendor, so treat our reading of this with the scepticism it deserves — but the mechanism is not mysterious and it is not about engineering talent.
An internal build starts at zero on everything that turns out to matter: evaluation harness, audit logging, permission scoping, rollback, and an institutional memory of which failure modes are real. Those are the things the surveys keep naming as the reason pilots stall — governance bolted on late, no agreed definition of "good enough", reliability never measured during the pilot. None of them is the model.
The buy-versus-build question is really a question about whether you want to discover the failure modes yourself, at $340,000 a discovery.
The honest caveat: buyers who purchase are often the ones with a clearer use case to begin with, which would inflate the gap. The published analysis does not control for that, and we are not going to pretend it does.
Sources: buy-versus-build success rates across 100+ deployments, Olakai, Feb 2026 · stall causes, ClarityArc, Jul 2026 and Vinsys, Jul 2026.
Choosing the first task
Six tests, drawn from what separated the deployments that scaled from the ones that stalled. Run a candidate task through them before anyone builds anything. A task that fails two or more is a task that will cost you a pilot to reject.
The rubric is ours; each test traces to a finding in the sections above. It is not specific to our platform and works the same if you build it yourself.
Graduated autonomy, and where deployments actually settle
The deployments that reach production do not switch autonomy on. They graduate it, and the reported pattern is consistent enough to be treated as a default. What is worth noticing is where they stop: most production agents live permanently in the middle two stages. Full autonomy is rare and reserved for well-bounded tasks with strong guardrails.
That is the same conclusion Figure 04 reaches from usage data and Issue 03 reaches from failure data. Three different measurements, one answer: the human does not leave.
Sources: graduated-autonomy stages, AnAr Solutions, Apr 2026 · co-pilot against autonomous failure-rate argument, Fiddler AI, Jul 2026.
Methods, and what this issue cannot tell you
Every figure comes from published work, dated and linked at the foot of its section and listed again in the references. Where a finding is qualitative rather than measured — the graduated-autonomy resting point in Figure 06, the barbell shape of function-level adoption — the note says so instead of drawing a bar.
The limits are real and worth stating plainly. Reported production rates disagree: 11%, 12%, 14% and 2% at full scale all appear in this issue from different surveys with different definitions of "production", and we have not reconciled them because they cannot be reconciled from the outside. The abandonment trend rests on two data points. Several sources are vendor-published analyses rather than peer-reviewed work, and vendors have an interest in the size of the problem they solve — including us. Where a figure appears in only one such source, treat it as indicative.
We are a vendor writing about why buying beats building. Read section 05 with that in front of you.
What no published source has, and we do: the revealed preference. Which task a business actually hands over first, what it hands over next, and which handovers get taken back. A later issue publishes that from our own platform, aggregated, with cohort definitions and the n behind every cut — including the abandonments.
References
Nine sources · every figure in this issue traces to one of them-
01AI Pilot to Production: Why 95% of AI Projects Stall, Olakai, Feb 2026 — S&P Global abandonment data (42% in 2025 against 17% in 2024), 46% of pilots scrapped, 33 prototypes to 4 in production, buy 67% against build 22% across 100+ deployments. olakai.aiUsed in: Figure 01, sections 01, 05
-
02AI Agent Scaling Gap March 2026, Digital Applied, Mar 2026 — survey of 650 enterprise technology leaders: 78% piloting, 14% scaled, 72% of stalled deployments stuck past six months, narrow-against-broad scoping comparison, the 90-day stability rule. digitalapplied.comUsed in: Figure 02, sections 02, 06
-
03AI Agent Adoption 2026: 120+ Enterprise Data Points, Digital Applied, Apr 2026 — function-level adoption barbell, customer service at 62%, deflection 39% tier-1 and 17% tier-2, cost-per-task reduction 40–70%, CSAT +2 / −4, HITL intervention 32%, small-business figures. digitalapplied.comUsed in: Figure 03, section 03
-
04Why 89% of AI Agent Projects Never Reach Production, Hendricks, Mar 2026 — Kore.ai and Deloitte production rate of 11%, 86% planning, 2% at full operational scale, Gartner 40% abandonment by 2027, WEF January 2026 finding that 60% of CEOs slowed deployment. hendricks.aiUsed in: opening funnel, section 01
-
05Why 88% of AI Agents Fail Production, Digital Applied, Mar 2026 — $340,000 average direct spend before abandonment, 12% of pilots reaching sustained production. digitalapplied.comUsed in: sections 01, 05
-
06Anthropic, 2026 Agentic Coding Trends Report — AI involved in roughly 60% of developer work, 0–20% of tasks fully delegated, 86% treating output as a starting point. Compiled in AI Agent Statistics 2026, May 2026Used in: Figure 04, section 04
-
07Why Most AI Agent Pilots Never Reach Production, ClarityArc, Jul 2026 — median time-to-value around five months, the four-to-eight-week prove-or-kill window, graduation criteria written before kickoff, named stall causes. clarityarc.comUsed in: Figure 05, sections 02, 05, 06
-
08Why 88% of Agentic AI Pilots Never Reach Production, AnAr Solutions, Apr 2026 — the four-stage graduated-autonomy pattern and the observation that most production agents remain in stages two and three. anarsolutions.comUsed in: Figure 06, section 07
-
09AI Agent Failure Rate, Fiddler AI, Jul 2026 — co-pilot against autonomous failure-rate comparison, confidence-threshold routing, retry-loop cost exposure. Cross-checked against Vinsys, Jul 2026, which places the reported failure range at 77–95%. fiddler.aiUsed in: sections 05, 07, 08
Several of these are vendor-published analyses rather than peer-reviewed research. Where a figure rests on a single such source it is described as reported rather than established, and section 08 names the reconciliation problems we could not solve from the outside.