Sistava
Field report · Agentic coworkers · Published 3 August 2026

Measuring your AI Employee

Three years of capability gains, and the median company still cannot get an agent past one function. This report separates what the evidence establishes from what the market asserts: how far autonomy has actually come, where deployments break, what it is doing to entry-level hiring, and which obligations became enforceable this month.

Scope
Agentic coworkers — systems that plan and execute multi-step work, not copilots
Evidence base
11 primary sources, cross-checked; all figures dated
Data through
June 2026 payroll · January 2026 benchmarks

Five findings

10%

is the ceiling. In no single business function do more than 10% of organizations report scaling AI agents, even though 62% are experimenting with them.

89 days

is how long it now takes for the length of task a model can finish autonomously to double, measured on models released since 2024 — down from roughly seven months across 2019–2025.

47.6%

of real professional deliverables were graded as good as or better than an industry expert's, on the strongest model at GDPval's release. Capability is not the binding constraint.

2 in 3

governance and risk leaders name security and risk — not model quality, not regulation — as the top barrier to scaling agentic AI.

19%

is the shortfall in employment for 22–25 year-olds in highly AI-exposed occupations, versus their less-exposed peers. It was 15% a year ago. The mechanism is stalled hiring, not layoffs.

Read together: the models cleared the bar for a large share of knowledge work before most organizations built anything capable of supervising them. The gap between the two is the whole story of 2026.

01

What counts as an AI employee

The term covers three quite different things, and most market statistics silently merge them. A copilot sits inside a tool and responds to a prompt; the human holds the plan. An agent executes a multi-step workflow — retrieving, calling tools, checking its own output — inside a scope someone else defined. An AI employee holds a standing role: a persistent scope of responsibility, its own tool credentials, work that arrives as a queue rather than a prompt, and a supervisor who reviews output rather than issues instructions.

This report is about the second and third. Where a source measures something broader, we say so. The distinction matters because it explains most of the apparent contradiction in the adoption data: near-universal AI use and near-absent agent deployment are both true, of different objects.

McKinsey's working definition — systems "capable of acting in the real world, planning and executing multiple steps in a workflow" — is the one most of the survey evidence below is built on (McKinsey, Nov 2025).

02

The funnel narrows fast

Evidence: strong — single large survey, consistent with two independent sources

McKinsey's 2025 Global Survey on AI (n=1,993, 105 countries, fielded June–July 2025) is the most-cited enterprise baseline. Its headline is adoption; its substance is attrition. Each step down the funnel loses more than half the previous one.

The bottom of the funnel is the number to hold onto: in any individual business function, no more than 10% of respondents say their organization is scaling agents. Most of the 23% who are scaling anything are doing it in one or two functions only.

Figure 1 · The enterprise agent funnel
Using AI in ≥1 business function
88%
At least experimenting with agents
62%
Reporting any EBIT impact from AI
39%
Scaling an agentic system anywhere
23%
Scaling agents in any one function
≤10%
"AI high performers" (>5% EBIT from AI)
6%
Source: McKinsey Global Survey on the State of AI, published 5 Nov 2025, fielded 25 Jun – 29 Jul 2025, n=1,993. Bars are shares of all respondents and are not strictly nested subsets.

Cross-check: Gartner's 2026 CIO survey puts deployed agent adoption at 17%, and a Deloitte survey of 3,235 leaders found only 25% had moved 40% or more of their agent pilots into production. Both figures reached us through secondary compilations rather than the original releases, so we treat them as corroborating direction, not precision.

03

The failure rate nobody can pin down

Evidence: thin — estimates differ by definition, and the loudest figure is the least traceable

"95% of AI pilots fail" is the most repeated statistic in this market. It comes from MIT NANDA's State of AI in Business 2025, it refers to generative AI pilots delivering expected ROI — not agents specifically — and in our checks it was reachable only through secondary summaries. Treat it as a slogan with a paper behind it, not a measurement you can act on.

The honest version is that published estimates range from roughly 77% to 95%, and they disagree because they count different things: pilots that never ship, pilots that ship without ROI, and projects forecast to be cancelled. The spread is the finding.

Figure 2 · Four claims, four denominators
Figure
Source
What it actually counts
Traceability
95%
MIT NANDA, 2025
GenAI pilots not delivering expected ROI. Not agent-specific.
Secondary only
>40%
Gartner forecast
Agentic projects predicted to be cancelled by end-2027. A forecast, not an observation.
Secondary only
75%
Deloitte, 3,235 leaders
Share who have not moved 40%+ of agent pilots into production.
Secondary only
90%
McKinsey, n=1,993
Not scaling agents in a given function. Measured, published methodology.
Primary
The only figure here we could verify against a published methodology is McKinsey's, and it is a statement about scaling, not failure. Anyone quoting a single failure percentage for agentic AI is quoting a number that does not exist yet.
04

Capability is not the constraint

Evidence: strong — two independent benchmark programmes, open methodology, wide confidence intervals disclosed

Two measurement programmes matter here, and they measure different axes. GDPval asks whether the output is good enough. METR's time horizon asks how long a job the system can hold together on its own.

GDPval, released by OpenAI in October 2025, is 1,320 tasks drawn from the real work of professionals averaging 14 years' experience, across 44 occupations in the nine largest sectors of US GDP. Expert graders compared model deliverables against human ones blind. At release, Claude Opus 4.1 was rated as good as or better than the human expert on 47.6% of tasks, and OpenAI reported performance more than tripling from GPT-4o to GPT-5 in a single year.

Figure 3 · GDPval — share of deliverables rated as good as or better than an industry expert
Human expert baseline
50%
Claude Opus 4.1 · Sep 2025
47.6%
GPT-5.2 Thinking · Dec 2025
70.9%
Sources: OpenAI, GDPval (Oct 2025) and Patwardhan et al., arXiv:2510.04374. The 70.9% figure is vendor-reported at model launch and has not been independently replicated; it is shown in amber for that reason. The 50% line is where a model ties the human, by construction.

METR measures the other axis. It times human experts on each task in its suite, then finds the task length at which a model succeeds half the time. In January 2026 it published a revised suite — 228 tasks, 31 of them eight hours or longer — and re-estimated fourteen models on it.

The trajectory is the finding. Across 2019–2025 the doubling time was about 196 days. Restricted to models released since 2023 it is 131 days. Since 2024, 89 days.

Figure 4 · 50%-time-horizon by model — the length of task a model finishes half the time
Horizontal axis is logarithmic. Human-expert minutes.
GPT-4 (0314) · 2023
3.5 min
GPT-4 (1106) · 2023
3.6 min
Claude Sonnet 3.7 · 2025
60 min
Claude Opus 4 · 2025
101 min
o3 · 2025
121 min
GPT-5 · 2025
214 min
Claude Opus 4.5 · 2025
320 min
196 days
doubling time, 2019–2026
131 days
models released since 2023
89 days
models released since 2024
Source: METR, Time Horizon 1.1, 29 Jan 2026. Confidence intervals are wide — Opus 4.5's runs 170–729 minutes. METR notes it measured human baseline times for only 5 of its 31 long tasks, and that its suite is close to saturation at the frontier.
05

Where they break

A 50% success rate is a strange thing to hire. The headline horizon figures describe the length at which a model succeeds half the time — which means the operationally relevant number, the 80% horizon, is considerably shorter, and the failure modes are not random.

Three limits recur across the literature, and none of them are fixed by a better model alone.

Limit 01

The cliff at four hours

Frontier models succeed near-100% of the time on tasks a skilled human finishes in under four minutes, and under 10% of the time on tasks taking a human more than about four hours. The curve does not decay gently; it falls off.

Limit 02

Context length is not context

Web agents scoring 40–50% on a short version of a task can drop below 10% on the identical task embedded in a longer interaction history — even when the needed information is still technically inside the context window.

Limit 03

The ruler has run out

METR states plainly that its suite has relatively few tasks the latest models cannot already do, and that it is prioritising raising the ceiling. Benchmark saturation makes the next year's capability claims harder, not easier, to verify.

Limits 01 and 02 are from a 2026 synthesis of agent failure modes citing Kwa et al. 2025 and arXiv:2512.04307; limit 03 is METR's own statement. The practical reading: agentic work needs to be decomposed into short, checkpointed units with fresh context — which is a workflow design problem, not a model selection problem.

06

Why deployments stall

Evidence: moderate — single survey, ~500 respondents, self-reported maturity

McKinsey's 2026 AI Trust Maturity Survey put the question directly to about 500 organizations' governance and risk owners between December 2025 and January 2026. The answer was unambiguous, and it was not about the models.

As the authors frame it: in the agentic era, organizations can no longer worry only about systems saying the wrong thing — they must contend with systems doing the wrong thing. Average responsible-AI maturity rose to 2.3 from 2.0 the year before, but only about a third of organizations reach level three or higher in strategy, governance, or agentic controls.

Figure 5 · The governance gap
Cite security & risk as top barrier to scaling agents
~65%
Identify inaccuracy as a highly relevant risk
74%
Identify cybersecurity as a highly relevant risk
72%
Cite knowledge & training gaps as top blocker
~60%
Reach maturity level 3+ in agentic governance
~30%
Report an AI-related incident
~8%
Incident rate held steady at roughly 8% year over year — but nearly 60% of those who had an incident rated their organization's response as merely satisfactory or worse. Preparedness is falling behind system complexity, not incident volume.
Organizations with a clearly accountable owner for responsible AI score 2.6 on maturity. Those without score 1.8. Naming an owner is the single cheapest intervention in the dataset.
Source: McKinsey, State of AI trust in 2026, published 25 Mar 2026. Survey fielded Dec 2025 – Jan 2026, n≈500. Values marked ~ are read from the published narrative rather than an exhibit table.
07

The money is real but local

Value shows up at the use-case level and mostly disappears on the way to the income statement. 39% of respondents attribute any EBIT impact to AI, and most of those put it below 5%. Cost benefits concentrate in software engineering, manufacturing and IT; revenue benefits in marketing and sales, strategy and corporate finance, and product development.

The qualitative picture is stronger than the financial one: a majority say AI has improved innovation, and nearly half report gains in customer satisfaction and competitive differentiation. Whether that eventually converts is the open question of the next two survey cycles.

64%
say AI is enabling innovation
39%
report any enterprise EBIT impact
6%
clear 5% EBIT — the "high performers"
more likely to have redesigned workflows
08

The labor signal is narrow and widening

Evidence: moderate — large administrative dataset, but the authors state these are descriptive, not causal

Brynjolfsson, Chandar and Chen's Canaries in the Coal Mine is the closest thing to a live instrument on this question: ADP payroll records covering roughly one in six US workers, updated monthly. The August 2026 revision runs through June 2026.

Their first finding is the one usually skipped: there is no widespread, economy-wide displacement. The aggregate barely moves. The composition underneath it moves a lot.

The new material in this revision is a mechanism. Employment fell among young workers in occupations built on codified knowledge — formal, documented, teachable from a textbook. It rose among experienced workers in occupations built on tacit knowledge acquired through practice and judgment. That is a sharper prediction than "AI takes jobs", and a testable one.

Figure 6 · Employment change, Nov 2022 – Jun 2026, workers aged 22–25
Two most AI-exposed quintiles
−11%
Three least AI-exposed quintiles
+10%
The gap, widening
15% 19%
July 2025 vintage → June 2026 vintage
Mechanism
Reduced hiring of young workers, not increased separations. Adjustment shows up in headcount, not base pay.
Experienced workers
No comparable gap between more- and less-exposed occupations. The divergence is specific to the entry rung.
Source: Brynjolfsson, Chandar & Chen, August 2026 revision, Stanford Digital Economy Lab. Exposure measures from Eloundou et al. (2024).
What the authors say against their own result

The gaps shrink once education is controlled for. Some divergent trends predate generative AI. The estimated gaps are larger in the ADP sample than in national survey benchmarks, concentrated in education, health care and public administration. And improvements to their own data pipeline left the raw patterns intact but made firm-adjusted estimates more sensitive to specification choices.

Their conclusion, verbatim in spirit: no single study — including theirs — settles this. Anyone citing the 19% as proof that AI is eliminating entry-level jobs is citing it against the authors' own reading.

09

What became enforceable this month

Evidence: strong — statutory text and adopted amendments

The EU AI Act's timeline moved in 2026, and the part that moved is not the part most compliance budgets were aimed at. The Digital Omnibus on AI — provisionally agreed 7 May 2026, the first amendment package since the Act's 2024 adoption — deferred the high-risk regime by sixteen months. It did not defer Article 50.

For anyone deploying agents that talk to customers, that is the operative fact: transparency and disclosure obligations applied from 2 August 2026, while conformity assessment, risk management systems and technical documentation for Annex III high-risk systems do not bite until December 2027.

Figure 7 · EU AI Act — the dates that survived the Omnibus
2 Aug 2025
GPAI model obligations apply to providers of foundation models
In force
2 Aug 2026
General application. Article 50 transparency — every chatbot and agent must disclose it is AI. Article 101 enforcement powers activate: up to 3% of worldwide turnover or €15M for GPAI providers.
Live now
2 Dec 2026
Deferred Article 50(2) machine-readable marking for pre-existing synthetic-content systems. New prohibitions on non-consensual intimate imagery and CSAM take effect.
Pending
2 Aug 2027
Member States must have at least one national regulatory sandbox — deferred by a year
Pending
2 Dec 2027
Annex III high-risk obligations — conformity assessment, registration, risk management, logging, human oversight. Deferred 16 months from Aug 2026.
Pending
2 Aug 2028
Annex I product-embedded high-risk systems — deferred one year
Pending
Sources: Covington & Burling, 18 May 2026, on the provisional Digital Omnibus agreement of 7 May 2026; maximum penalty of €35M or 7% of global turnover is set by the Act itself. Also newly in the 3%/€15M fining band: breaches of Article 25(2) and (4) — the duty of an upstream provider to hand a downstream provider the technical documentation, known failure modes, and test access it needs. That one lands squarely on anyone shipping agents built on someone else's model.
10

What the 6% do differently

McKinsey ran a relative-weights analysis across 31 organizational variables to find what actually separates high performers. Two factors dominated, and neither is a technology choice.

~3×

Workflow redesign

High performers are nearly three times as likely to have fundamentally redesigned individual workflows. Of every factor tested, this had one of the strongest contributions to business impact. Bolting an agent onto an unchanged process is the modal failure.

65% vs 23%

Defined human validation

High performers are far more likely to have explicit processes determining how and when model output needs human sign-off. Not a review culture — a defined gate, per workflow, written down.

2.6 vs 1.8

A named owner

Organizations that assign clear accountability for responsible AI score a full maturity level higher than those that leave it distributed. Ownership precedes control.

Leadership that uses it

High performers are three times more likely to strongly agree that senior leaders demonstrate ownership of AI initiatives — including role-modelling the use of the tools themselves.

Our reading

Every factor that predicts success in this dataset is a management artefact, not a model artefact: a redefined workflow, a written approval gate, a named owner, a supervisor who reads the output. These are the things organizations already know how to build for human employees, and have not yet built for AI ones.

That is the thesis Sista is built on, and it is why we publish this: an AI employee that arrives without an onboarding, a scope, a review cadence and a place its work is visible will fail for the same reasons a human hire would. The evidence in this report is not primarily about models getting better. It is about the supervision layer being missing.

11

What we could not establish

A

There is no credible market size for AI employees

Every projection we found traced back to paid report vendors without published methodology. One statistics compiler explicitly removed its own market-size figures in July 2026 for exactly this reason. We report no market size rather than launder one.

B

No public benchmark measures a standing role

GDPval measures deliverables. METR measures task length. Neither measures what an AI employee actually is: durable context across weeks, a queue of related work, handoffs, and the judgment to escalate. The thing the market is selling has no benchmark.

C

Causality in the labor data is unresolved

The strongest available study declines to claim it. Education controls attenuate the effect; the ADP sample overstates it relative to national benchmarks. Anyone who tells you AI has definitively eliminated entry-level jobs is ahead of the data.

D

The capability trend may not hold

The 89-day doubling is fitted to two years of data on a suite METR itself says is near saturation, measuring mostly software tasks — the domain labs optimise hardest. A separate no-chain-of-thought analysis finds a 373-day doubling on a broader task set. Both trends deserve tracking; neither should be extrapolated alone.

Sources

The state of AI in 2025: Agents, innovation, and transformation — McKinsey & Company, 5 Nov 2025. Global survey, n=1,993, fielded 25 Jun – 29 Jul 2025.
State of AI trust in 2026: Shifting to the agentic era — McKinsey & Company, 25 Mar 2026. AI Trust Maturity Survey, n≈500, Dec 2025 – Jan 2026.
Time Horizon 1.1 — METR, 29 Jan 2026. 228-task suite, 14 models re-estimated.
Measuring AI Ability to Complete Long Software Tasks — METR, 19 Mar 2025. Original time-horizon methodology.
Measuring the performance of our models on real-world tasks — OpenAI, Oct 2025. Announcement and headline GDPval results.
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks — Patwardhan et al., OpenAI, arXiv:2510.04374. 1,320 tasks, 44 occupations, 9 sectors.
Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence — Brynjolfsson, Chandar & Chen, Stanford Digital Economy Lab, August 2026 revision. ADP payroll data through June 2026.
Canaries Dashboard — Stanford Digital Economy Lab & ADP Research. Monthly-updating series behind the paper.
EU AI Act Update: Timeline Relief, Targeted Simplification, and New Prohibitions — Covington & Burling, 18 May 2026. Analysis of the Digital Omnibus on AI.
Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents — 2026. Long-horizon degradation and context-management failure clusters.
Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models — 2026. The 373-day doubling counterpoint on a broader task distribution.
Figures attributed to Gartner, Deloitte and MIT NANDA in sections 02 and 03 reached us only through secondary compilations. They are labelled as such throughout and none carry an argument on their own.