Measuring your AI Employee
Three years of capability gains, and the median company still cannot get an agent past one function. This report separates what the evidence establishes from what the market asserts: how far autonomy has actually come, where deployments break, what it is doing to entry-level hiring, and which obligations became enforceable this month.
Five findings
What counts as an AI employee
The term covers three quite different things, and most market statistics silently merge them. A copilot sits inside a tool and responds to a prompt; the human holds the plan. An agent executes a multi-step workflow — retrieving, calling tools, checking its own output — inside a scope someone else defined. An AI employee holds a standing role: a persistent scope of responsibility, its own tool credentials, work that arrives as a queue rather than a prompt, and a supervisor who reviews output rather than issues instructions.
This report is about the second and third. Where a source measures something broader, we say so. The distinction matters because it explains most of the apparent contradiction in the adoption data: near-universal AI use and near-absent agent deployment are both true, of different objects.
McKinsey's working definition — systems "capable of acting in the real world, planning and executing multiple steps in a workflow" — is the one most of the survey evidence below is built on (McKinsey, Nov 2025).
The funnel narrows fast
McKinsey's 2025 Global Survey on AI (n=1,993, 105 countries, fielded June–July 2025) is the most-cited enterprise baseline. Its headline is adoption; its substance is attrition. Each step down the funnel loses more than half the previous one.
The bottom of the funnel is the number to hold onto: in any individual business function, no more than 10% of respondents say their organization is scaling agents. Most of the 23% who are scaling anything are doing it in one or two functions only.
Cross-check: Gartner's 2026 CIO survey puts deployed agent adoption at 17%, and a Deloitte survey of 3,235 leaders found only 25% had moved 40% or more of their agent pilots into production. Both figures reached us through secondary compilations rather than the original releases, so we treat them as corroborating direction, not precision.
The failure rate nobody can pin down
"95% of AI pilots fail" is the most repeated statistic in this market. It comes from MIT NANDA's State of AI in Business 2025, it refers to generative AI pilots delivering expected ROI — not agents specifically — and in our checks it was reachable only through secondary summaries. Treat it as a slogan with a paper behind it, not a measurement you can act on.
The honest version is that published estimates range from roughly 77% to 95%, and they disagree because they count different things: pilots that never ship, pilots that ship without ROI, and projects forecast to be cancelled. The spread is the finding.
Capability is not the constraint
Two measurement programmes matter here, and they measure different axes. GDPval asks whether the output is good enough. METR's time horizon asks how long a job the system can hold together on its own.
GDPval, released by OpenAI in October 2025, is 1,320 tasks drawn from the real work of professionals averaging 14 years' experience, across 44 occupations in the nine largest sectors of US GDP. Expert graders compared model deliverables against human ones blind. At release, Claude Opus 4.1 was rated as good as or better than the human expert on 47.6% of tasks, and OpenAI reported performance more than tripling from GPT-4o to GPT-5 in a single year.
METR measures the other axis. It times human experts on each task in its suite, then finds the task length at which a model succeeds half the time. In January 2026 it published a revised suite — 228 tasks, 31 of them eight hours or longer — and re-estimated fourteen models on it.
The trajectory is the finding. Across 2019–2025 the doubling time was about 196 days. Restricted to models released since 2023 it is 131 days. Since 2024, 89 days.
Where they break
A 50% success rate is a strange thing to hire. The headline horizon figures describe the length at which a model succeeds half the time — which means the operationally relevant number, the 80% horizon, is considerably shorter, and the failure modes are not random.
Three limits recur across the literature, and none of them are fixed by a better model alone.
Limits 01 and 02 are from a 2026 synthesis of agent failure modes citing Kwa et al. 2025 and arXiv:2512.04307; limit 03 is METR's own statement. The practical reading: agentic work needs to be decomposed into short, checkpointed units with fresh context — which is a workflow design problem, not a model selection problem.
Why deployments stall
McKinsey's 2026 AI Trust Maturity Survey put the question directly to about 500 organizations' governance and risk owners between December 2025 and January 2026. The answer was unambiguous, and it was not about the models.
As the authors frame it: in the agentic era, organizations can no longer worry only about systems saying the wrong thing — they must contend with systems doing the wrong thing. Average responsible-AI maturity rose to 2.3 from 2.0 the year before, but only about a third of organizations reach level three or higher in strategy, governance, or agentic controls.
The money is real but local
Value shows up at the use-case level and mostly disappears on the way to the income statement. 39% of respondents attribute any EBIT impact to AI, and most of those put it below 5%. Cost benefits concentrate in software engineering, manufacturing and IT; revenue benefits in marketing and sales, strategy and corporate finance, and product development.
The qualitative picture is stronger than the financial one: a majority say AI has improved innovation, and nearly half report gains in customer satisfaction and competitive differentiation. Whether that eventually converts is the open question of the next two survey cycles.
The labor signal is narrow and widening
Brynjolfsson, Chandar and Chen's Canaries in the Coal Mine is the closest thing to a live instrument on this question: ADP payroll records covering roughly one in six US workers, updated monthly. The August 2026 revision runs through June 2026.
Their first finding is the one usually skipped: there is no widespread, economy-wide displacement. The aggregate barely moves. The composition underneath it moves a lot.
The new material in this revision is a mechanism. Employment fell among young workers in occupations built on codified knowledge — formal, documented, teachable from a textbook. It rose among experienced workers in occupations built on tacit knowledge acquired through practice and judgment. That is a sharper prediction than "AI takes jobs", and a testable one.
The gaps shrink once education is controlled for. Some divergent trends predate generative AI. The estimated gaps are larger in the ADP sample than in national survey benchmarks, concentrated in education, health care and public administration. And improvements to their own data pipeline left the raw patterns intact but made firm-adjusted estimates more sensitive to specification choices.
Their conclusion, verbatim in spirit: no single study — including theirs — settles this. Anyone citing the 19% as proof that AI is eliminating entry-level jobs is citing it against the authors' own reading.
What became enforceable this month
The EU AI Act's timeline moved in 2026, and the part that moved is not the part most compliance budgets were aimed at. The Digital Omnibus on AI — provisionally agreed 7 May 2026, the first amendment package since the Act's 2024 adoption — deferred the high-risk regime by sixteen months. It did not defer Article 50.
For anyone deploying agents that talk to customers, that is the operative fact: transparency and disclosure obligations applied from 2 August 2026, while conformity assessment, risk management systems and technical documentation for Annex III high-risk systems do not bite until December 2027.
What the 6% do differently
McKinsey ran a relative-weights analysis across 31 organizational variables to find what actually separates high performers. Two factors dominated, and neither is a technology choice.
Every factor that predicts success in this dataset is a management artefact, not a model artefact: a redefined workflow, a written approval gate, a named owner, a supervisor who reads the output. These are the things organizations already know how to build for human employees, and have not yet built for AI ones.
That is the thesis Sista is built on, and it is why we publish this: an AI employee that arrives without an onboarding, a scope, a review cadence and a place its work is visible will fail for the same reasons a human hire would. The evidence in this report is not primarily about models getting better. It is about the supervision layer being missing.
What we could not establish
There is no credible market size for AI employees
Every projection we found traced back to paid report vendors without published methodology. One statistics compiler explicitly removed its own market-size figures in July 2026 for exactly this reason. We report no market size rather than launder one.
No public benchmark measures a standing role
GDPval measures deliverables. METR measures task length. Neither measures what an AI employee actually is: durable context across weeks, a queue of related work, handoffs, and the judgment to escalate. The thing the market is selling has no benchmark.
Causality in the labor data is unresolved
The strongest available study declines to claim it. Education controls attenuate the effect; the ADP sample overstates it relative to national benchmarks. Anyone who tells you AI has definitively eliminated entry-level jobs is ahead of the data.
The capability trend may not hold
The 89-day doubling is fitted to two years of data on a suite METR itself says is near saturation, measuring mostly software tasks — the domain labs optimise hardest. A separate no-chain-of-thought analysis finds a 373-day doubling on a broader task set. Both trends deserve tracking; neither should be extrapolated alone.