Adoption & scoping Published 25 August 2026 · Evidence current to August 2026 # What survives the pilot Almost everyone is running an AI agent somewhere. Almost nobody is running one in production. This issue is about the difference — which first tasks make it through, which die in the pilot, and what separates them. From intent to operating scale Planning to deploy 86% Running a pilot 78% In production 11% At full scale 2% An intent-to-execution ratio of roughly eight to one, and a further five-fold drop between "in production" and "integrated across business functions". The interesting question is not why the technology fails. It is which tasks make it all the way down this chart, because a substantial number do. ## Six numbers Each traced to a dated source 42% of companies abandoned most of their AI initiatives in 2025 — up from 17% the year before. Abandonment is accelerating, not settling. S&P Global · reported Feb 2026 $340k average direct spend on a failed agent project before it is abandoned. The pilot is not the cheap option. Reported Mar 2026 62% of enterprises run a customer-service agent in production — the highest share of any function by a wide margin. Enterprise adoption survey · Apr 2026 90 days of stability before scope expansion, in the deployments that scaled. The ones that stalled widened scope from the start. 650 tech leaders · Mar 2026 67 / 22 success rate, in percent, for buying from a vendor against building internally. The build path fails three times as often. 100+ deployments analysed · Feb 2026 0–20% of tasks developers can fully delegate, while using AI in roughly 60% of their work. Partial delegation is the normal case. Anthropic Agentic Coding Trends · 2026 The AI employee market talks almost exclusively about capability, which is odd, because capability is not what most buyers are deciding. They are deciding what to hand over first. Get that choice right and a deployment compounds. Get it wrong and it joins the eighty-something percent that never leave the pilot, having spent real money to find out. 01 ## Abandonment is getting worse, not better The intuitive expectation is that the pilot-to-production gap closes as the technology matures and organisations learn. The data says the opposite. The share of companies abandoning most of their AI initiatives went from 17% to 42% in a single year, and the average enterprise scrapped 46% of its pilots before they reached production. Two readings are available and both are probably true. Organisations are attempting harder things than they were a year ago, so more attempts fail. And organisations have become willing to kill projects they would previously have kept alive — which, if you are the one paying, is progress rather than decline. A failed agent project costs about $340,000 before anyone calls it. The expensive decision is made in the first week, when the task is chosen. Fig 01 Companies abandoning most AI initiatives Year on year · two points, one direction 17% 42% 2024 2025 Alongside this, Gartner forecasts that more than 40% of agentic AI projects will be abandoned or substantially restructured by 2027, and the World Economic Forum reported in January 2026 that 60% of CEOs had actively slowed their agent timelines over error rates and accountability. The pattern is not confined to one survey. Sources: S&P Global abandonment data and the 46% pilot-scrap figure via [Olakai , Feb 2026 · $340k per failed project via Digital Applied , Mar 2026 · Gartner and WEF figures via Hendricks , Mar 2026. 02 ## Narrow survives. Broad does not. A March 2026 survey of 650 enterprise technology leaders found 78% running at least one pilot and only 14% with an agent scaled to organisation-wide operation. Of the stalled deployments, 72% had been stuck for more than six months with no resolution path. Comparing the two groups gives the clearest scoping rule in the current evidence. Deployments that scaled started with an agent scoped to a single, well-defined task with a measurable output — a document classifier, a data-enrichment step, a routing decision. Deployments that stalled attempted multi-function agents covering broad task categories from the beginning, producing a combinatorial explosion of edge cases that no evaluation harness could cover and no dashboard could attribute. The successful group expanded scope only after the narrow version held stable for 90 days or more. That is the whole rule, and it is unglamorous enough that most teams skip it. Fig 02 Two ways to start The scoping decision that separated scaled deployments from stalled ones Scaled One task, measurable output A classifier, an enrichment step, a routing decision. Narrow enough that "did it work" has an answer, and small enough that the edge cases are enumerable. Scope widened only after 90+ days stable Stalled Broad category, open-ended A multi-function agent aimed at a whole department from day one. Quality variance compounds, edge cases multiply, and no dashboard can attribute a regression to a cause. 72% stuck past six months 78% running at least one pilot 14% scaled to organisation-wide use ~5 months median time-to-value when it works Sources: 650-leader survey, scoping comparison and the 90-day rule, AI Agent Scaling Gap , Mar 2026 · median time-to-value and ship/kill discipline, ClarityArc , Jul 2026. 03 ## Where the demand actually concentrated Function-level adoption tells a more useful story than organisation-level adoption, because agents are deployed against a specific job to be done rather than against a company. The distribution is a barbell: customer service and software engineering are saturated relative to legal and HR. Customer service is the clear leader at 62% of enterprises, and it is also the only function with published outcome data granular enough to be worth reading. That data is more equivocal than the adoption number suggests. Fig 03 The most-deployed function, measured Customer-service agents in production · averages across surveyed enterprises 39% / 17% ticket deflection on tier-1 against tier-2 inquiries. Difficulty more than halves the benefit. 40–70% cost-per-task reduction in deflection use cases, with the top decile reporting 78%. +2 / −4 median CSAT delta against a human-only baseline: better on quick resolutions, worse on complex multi-touch cases. 32% of agent-handled conversations are escalated or supervised — roughly one in three. Read the four together and the shape of a viable first task appears: high-volume, low-complexity, reversible, with an obvious success signal. The same deployment that wins on tier-1 volume loses satisfaction points on the complex tail — which is an argument for scoping the agent to the tier, not to the queue. Note also the 32% intervention rate against the 20–30% escalation band in Issue 03](https://olakai.ai/blog/ai-pilot-to-production/): the most-deployed function in the market sits just above the band. Sources: function-level adoption, deflection, CSAT delta and intervention rate, [AI Agent Adoption 2026 , Apr 2026. Small-business figures from the same source: 14% in production, 38% piloting, 0.7 agents on average. 04 ## Nobody hands over a whole job The most useful measurement of delegation in the wild comes from software engineering, where usage is instrumented more thoroughly than anywhere else. Developers now use AI in roughly 60% of their work. They can fully delegate 0–20% of tasks. And 86% treat AI-generated output as a starting point rather than a final answer. Those three numbers describe partial delegation, and partial delegation is the normal case rather than a transitional stage. A product that assumes the whole job transfers is mispricing what it is selling and mis-designing where the human sits. Fig 04 Touched versus handed over Share of a developer's work, on a single 0–100% axis AI involved at all 60% Fully delegated 0–20% 0% 25% 50% 75% 100% The gap between the two bars is the product surface almost nobody designs for: work the human stays inside. The upper bar is the reported 60% involvement; the lower bar is drawn at the top of the reported 0–20% full-delegation range. Sources: Anthropic, 2026 Agentic Coding Trends Report , as compiled in AI Agent Statistics 2026 , May 2026 · consistent with the delegation-gap findings in Issue 03](https://www.digitalapplied.com/blog/ai-agent-adoption-2026-enterprise-data-points), section 06. 05 ## Building it yourself fails three times as often Across an analysis of more than a hundred agent deployments, purchasing from a specialised vendor or building through a partnership succeeded roughly 67% of the time. Internal builds succeeded about 22% of the time. We are a vendor, so treat our reading of this with the scepticism it deserves — but the mechanism is not mysterious and it is not about engineering talent. An internal build starts at zero on everything that turns out to matter: evaluation harness, audit logging, permission scoping, rollback, and an institutional memory of which failure modes are real. Those are the things the surveys keep naming as the reason pilots stall — governance bolted on late, no agreed definition of "good enough", reliability never measured during the pilot. None of them is the model. The buy-versus-build question is really a question about whether you want to discover the failure modes yourself, at $340,000 a discovery. The honest caveat: buyers who purchase are often the ones with a clearer use case to begin with, which would inflate the gap. The published analysis does not control for that, and we are not going to pretend it does. Sources: buy-versus-build success rates across 100+ deployments, [Olakai , Feb 2026 · stall causes, ClarityArc , Jul 2026 and Vinsys , Jul 2026. 06 ## Choosing the first task Six tests, drawn from what separated the deployments that scaled from the ones that stalled. Run a candidate task through them before anyone builds anything. A task that fails two or more is a task that will cost you a pilot to reject. The rubric is ours; each test traces to a finding in the sections above. It is not specific to our platform and works the same if you build it yourself. Fig 05 Six tests for a first task Score a candidate before you scope it Test Pass condition Why Volume It happens often enough that a month of running it produces a readable signal, not three anecdotes. Narrow high-volume tasks are what scaled. Bounded You can list the edge cases on one page. If you cannot, the evaluation harness will not cover them either. Combinatorial edge cases stalled the broad deployments. Reversible A wrong output can be undone without contacting a customer, a regulator, or a bank. Irreversible steps make an approval gate mandatory and slow. Measurable "Did it work" has an answer that does not require a meeting. Write the success rate, exception rate and cost per completed task before kickoff. No agreed definition of good enough is a named stall cause. Owned A named person signs the risk and can kill it. Set a ship-or-kill date at the start; four to eight weeks is usually enough. A pilot past a quarter with no decision is a zombie. Partial by design You know which part the human keeps. Full handover is the exception, not the target state. 0–20% of tasks are fully delegated even by heavy users. Deliberately absent: anything about the model. None of the six tests is answered by a benchmark score, and none of them changes if you switch provider. 07 ## Graduated autonomy, and where deployments actually settle The deployments that reach production do not switch autonomy on. They graduate it, and the reported pattern is consistent enough to be treated as a default. What is worth noticing is where they stop: most production agents live permanently in the middle two stages. Full autonomy is rare and reserved for well-bounded tasks with strong guardrails. That is the same conclusion Figure 04 reaches from usage data and Issue 03](https://olakai.ai/blog/ai-pilot-to-production/) reaches from failure data. Three different measurements, one answer: the human does not leave. Fig 06 Four stages, and where they stop The graduated-autonomy pattern reported across deployments that reached production Stage 1 Suggest The agent analyses and proposes. A human decides. Passed through Stage 2 Act, reviewed The agent acts, a human reviews every action before it lands. Most production agents live here Stage 3 Routine autonomy Routine decisions run independently. Edge cases route to a human. …or here Stage 4 Full autonomy No routine human involvement. Only for well-bounded tasks with strong guardrails. Rare The stages are as reported. "Most production agents live permanently in Phase 2 or Phase 3" is a qualitative finding, not a measured distribution — no source we found publishes the split, and we have not invented one. Sources: graduated-autonomy stages, AnAr Solutions , Apr 2026 · co-pilot against autonomous failure-rate argument, Fiddler AI , Jul 2026. 08 ## Methods, and what this issue cannot tell you Every figure comes from published work, dated and linked at the foot of its section and listed again in the references. Where a finding is qualitative rather than measured — the graduated-autonomy resting point in Figure 06, the barbell shape of function-level adoption — the note says so instead of drawing a bar. The limits are real and worth stating plainly. Reported production rates disagree: 11%, 12%, 14% and 2% at full scale all appear in this issue from different surveys with different definitions of "production", and we have not reconciled them because they cannot be reconciled from the outside. The abandonment trend rests on two data points. Several sources are vendor-published analyses rather than peer-reviewed work, and vendors have an interest in the size of the problem they solve — including us. Where a figure appears in only one such source, treat it as indicative. We are a vendor writing about why buying beats building. Read section 05 with that in front of you. What no published source has, and we do: the revealed preference. Which task a business actually hands over first, what it hands over next, and which handovers get taken back. A later issue publishes that from our own platform, aggregated, with cohort definitions and the n behind every cut — including the abandonments. ## References Nine sources · every figure in this issue traces to one of them - 01 AI Pilot to Production: Why 95% of AI Projects Stall , Olakai, Feb 2026 — S&P Global abandonment data (42% in 2025 against 17% in 2024), 46% of pilots scrapped, 33 prototypes to 4 in production, buy 67% against build 22% across 100+ deployments. olakai.ai Used in: Figure 01, sections 01, 05 - 02 AI Agent Scaling Gap March 2026 , Digital Applied, Mar 2026 — survey of 650 enterprise technology leaders: 78% piloting, 14% scaled, 72% of stalled deployments stuck past six months, narrow-against-broad scoping comparison, the 90-day stability rule. digitalapplied.com Used in: Figure 02, sections 02, 06 - 03 AI Agent Adoption 2026: 120+ Enterprise Data Points , Digital Applied, Apr 2026 — function-level adoption barbell, customer service at 62%, deflection 39% tier-1 and 17% tier-2, cost-per-task reduction 40–70%, CSAT +2 / −4, HITL intervention 32%, small-business figures. digitalapplied.com Used in: Figure 03, section 03 - 04 Why 89% of AI Agent Projects Never Reach Production , Hendricks, Mar 2026 — Kore.ai and Deloitte production rate of 11%, 86% planning, 2% at full operational scale, Gartner 40% abandonment by 2027, WEF January 2026 finding that 60% of CEOs slowed deployment. hendricks.ai Used in: opening funnel, section 01 - 05 Why 88% of AI Agents Fail Production , Digital Applied, Mar 2026 — $340,000 average direct spend before abandonment, 12% of pilots reaching sustained production. digitalapplied.com Used in: sections 01, 05 - 06 Anthropic, 2026 Agentic Coding Trends Report — AI involved in roughly 60% of developer work, 0–20% of tasks fully delegated, 86% treating output as a starting point. Compiled in AI Agent Statistics 2026 , May 2026 Used in: Figure 04, section 04 - 07 Why Most AI Agent Pilots Never Reach Production , ClarityArc, Jul 2026 — median time-to-value around five months, the four-to-eight-week prove-or-kill window, graduation criteria written before kickoff, named stall causes. clarityarc.com Used in: Figure 05, sections 02, 05, 06 - 08 Why 88% of Agentic AI Pilots Never Reach Production , AnAr Solutions, Apr 2026 — the four-stage graduated-autonomy pattern and the observation that most production agents remain in stages two and three. anarsolutions.com Used in: Figure 06, section 07 - 09 AI Agent Failure Rate , Fiddler AI, Jul 2026 — co-pilot against autonomous failure-rate comparison, confidence-threshold routing, retry-loop cost exposure. Cross-checked against Vinsys , Jul 2026, which places the reported failure range at 77–95%. fiddler.ai Used in: sections 05, 07, 08 Several of these are vendor-published analyses rather than peer-reviewed research. Where a figure rests on a single such source it is described as reported rather than established, and section 08 names the reconciliation problems we could not solve from the outside.