When Not to Hire an AI Employee, and Who Should Wait
Guide — — by Mahmoud Zalt
Most enterprise AI pilots show no measurable P&L impact. Here is what the research says fails, who should not buy yet, and what to fix first.
The number that should slow you down
MIT's NANDA initiative studied 52 executive interviews, 153 leader surveys, and 300 public deployments, and reported that 95% of generative AI pilots produced no measurable profit and loss impact. The funnel behind that is just as telling: around 60% of organisations evaluated something, about 20% ran a pilot, and roughly 5% reached live production.
Be careful with what that measures. It is not a finding that 95% of the tools failed to work. It is a finding that 95% of pilots could not be traced to a number on a financial statement, which is a different and in some ways more uncomfortable result. Plenty of those pilots produced output people liked and nothing anyone could count.
At a Glance
- 95%
- Of pilots showed no measurable P&L impact
- 60%
- Of organisations evaluated something
- 20%
- Ran an actual pilot
- 5%
- Reached live production
That drop from evaluated to live is where most of the money and most of the disappointment lives. Very few of those projects were killed by the technology. They were killed by a vague brief, an absent owner, or an outcome nobody had defined well enough to measure.
We sell an AI employee product, so this is an awkward statistic for us to lead with. It is also the single most useful thing you can know before you spend anything, because the same study says exactly which side of the line the successes were on. The product is Sistava, and we would rather you ran the readiness checklist further down this page than a demo, because a buyer who skips it usually ends up in the 95%.
What actually worked, and what did not
The wins in the study were back office. Document automation and review produced $2 million to $10 million in annual savings. Risk monitoring produced around $1 million annually. Procurement showed real results. Content and marketing operations cut agency spend by about 30%.
The disappointments were the sales and marketing focus, and this is the part worth sitting with, because that is where the budgets were going. Organisations were spending on the front office while the returns were appearing in operations. The money and the results were pointed in different directions.
| Area | What the research found |
|---|---|
| Document automation and review | Succeeded, $2M to $10M annual savings reported |
| Risk monitoring | Succeeded, around $1M annual savings reported |
| Procurement | Succeeded |
| Content and marketing operations | Succeeded, roughly 30% reduction in agency spend |
| Sales and marketing as the primary focus | Disappointed, despite receiving the larger share of budget |
| Internally built tools | Succeeded at half the rate of tools built by external vendors |
One more finding, and it cuts against a very common instinct: tools built by external vendors succeeded twice as often as internal builds. Building it yourself feels like control and usually buys you a longer road to the same place, with your own team maintaining it forever.
Seven cases where the answer is no, or not yet
Here is where this gets specific. If you recognise yourself in any of these, the honest recommendation is to fix the underlying thing first. Buying now will produce output you cannot measure and a bill you can.
- Your first target is sales or marketing outcomes. This is the exact profile the research found disappointing despite the budget. If you want revenue lift as the first win, start somewhere back-office where the wins were, and come back to the front office once you have a working pattern.
- You cannot write the job down. If you cannot describe the task in five bullet points a new starter could follow, nobody can specify it for software either. A vague brief is the most common cause of a pilot that produces nothing countable.
- Nobody will own the review. Every employee needs a named person who reads its output and owns whether it was right. If that person does not exist or has no time, the work will drift and you will not notice until something goes wrong publicly.
- Your bottleneck is decisions, not throughput. If your week is slow because you are waiting on choices, approvals, or information from other people, more output does not help. You will just generate work that queues behind the same decision.
- You need a P&L number this quarter. The study's central finding is that most pilots never produced one. If your case for buying depends on a traceable financial result in ninety days, define that metric before you spend anything, and be prepared for the honest answer that it is not measurable yet.
- The work is physical, or needs an accountable signature. Anything requiring a physical presence and anything requiring a legally accountable person is out. Not slower, not harder, out. No configuration changes that.
- You are planning to build it yourself. The research found vendor-built tools succeeded twice as often as internal builds. If you are choosing between buying and building, the published odds are not close, and the maintenance burden of the internal path never ends.
If none of those seven applies to you, that is genuinely a good sign, and it says more about your readiness than any demo could. The profile that works is unglamorous: a repetitive back-office job with a countable output, a named person who reads the results, and a definition of done you could write on an index card. That is the shape the successful five percent had, and it is available to a small team just as much as to an enterprise. A hire at Sistava is built around that same shape, one job written in plain English, an activity feed the owner reads, and approval gates on anything you do not want happening unattended.
The work that should stay with a person
Separate from readiness, there is a category of work that is simply not a candidate. The published task list for an assistant is heavy with it: prioritising an inbox by what actually matters this week, following up with a real person until they respond, re-coordinating travel when a flight is cancelled, preparing for a meeting where reading the room is the point, and writing an SOP for how this particular company genuinely works.
There is a measured reason those resist handover. The OSWorld 2.0 benchmark of authentic professional tasks names three failure patterns, and the third is that agents struggle to maintain hidden state across steps, repeating work or executing from a plan that went stale several steps ago. Holding the thread of a week-long piece of work, and noticing when the premise quietly changed, is exactly what that judgment work is made of.
That is a technical finding about how software behaves over long horizons, not a claim about anyone's colleagues. It is also the most reliable filter you have. If a job's difficulty is mostly about holding context nobody wrote down, it stays where it is.
What to fix before you buy anything
The readiness checklist
- Pick a back-office job, not a revenue job — Document handling, review, reconciliation, reporting, research. That is where the measured wins were, and it is where a wrong result costs you a correction rather than a customer.
- Write the job down in five bullets — Inputs, steps, definition of done, what must never happen, who reviews it. If you cannot fill all five, that is the work to do first, and it is worth doing regardless.
- Name the number you expect to move — Hours returned, backlog cleared, errors caught, response time. Write down today's value before you start, because you cannot reconstruct a baseline afterwards.
- Name the human owner — One person who reads the activity feed and owns whether the output was right. Not a committee, and not the person who is already the most overloaded.
- Decide what needs holding back — Reads and drafts can run unattended. Anything a customer sees or anything that changes a system of record belongs behind an approval gate until a person releases it.
- Set a stop date — Four weeks, then a real decision against the number you wrote down. Pilots that never end are how the 95% happens.
That checklist is deliberately boring, and it is the difference between the 5% and everyone else. Most of it is work you would benefit from even if you never bought anything, which is a reasonable test of whether advice is honest.
There is a broader caution worth knowing too, though it comes to us through secondary reporting rather than a source we could read directly, so treat it as directional. Gartner has been reported as forecasting that more than 40% of agentic AI projects will be cancelled by the end of 2027 on escalating costs, unclear business value, and inadequate risk controls, and as estimating that of the thousands of vendors claiming agentic capability, only around 130 offer genuinely agentic features.
And if you are not ready, not buying is a perfectly good outcome for a Tuesday. The readiness checklist above costs nothing, improves your operation on its own, and will still be waiting whenever the timing is better. Nothing in the research suggests that moving early is what separates the winners. If you want a feel for the output quality without spending anything, the free tools at Sistava run a single job in the browser without an account, which is enough to judge whether the drafts would be worth reviewing at all.
FAQ
Why do most AI pilots fail?
MIT's NANDA study found 95% of generative AI pilots produced no measurable P&L impact, with about 60% of organisations evaluating, 20% piloting, and only 5% reaching production. The failures were rarely technical. They came from vague briefs, no named owner for the output, and no metric defined before the work started, which makes a result impossible to trace to a financial statement afterwards.
Should I use AI for sales and marketing first?
The research suggests not. MIT's study found sales and marketing focus disappointed even though budgets favoured it over operations, while back-office work delivered the measured wins: document automation and review, risk monitoring, procurement, and marketing operations. Start where the results appeared, build a working pattern with a countable output, then apply what you learned to the front office.
Is it better to build our own AI tools or buy them?
The same study found tools built by external vendors succeeded roughly twice as often as internal builds. Building feels like control, but it puts your team on a longer road to the same destination and leaves them maintaining it indefinitely. If you are genuinely choosing between the two, the published odds are not close, and the maintenance burden is the part usually underestimated.
What kind of work should not be given to an AI employee?
Judgment and relationship work: prioritising by what actually matters, following up with a real person until they respond, re-coordinating when plans collapse, preparing for a meeting where reading the room is the point, and documenting how your company really works. Benchmarks show agents struggle to maintain hidden state across steps, which is precisely what that work requires. Anything physical or needing an accountable signature is also out.
How do I know if my business is ready for an AI employee?
Six checks. You can name a back-office job with a countable output, you can write it down in five bullets, you can state the number you expect to move and its value today, you can name one person who will read the output, you know which actions need holding behind approval, and you have set a stop date for a real decision. Missing any of those means fix that first.
How long should an AI pilot run before I decide?
Set four weeks and a hard stop with a real decision at the end, measured against a baseline you wrote down before you started. Pilots without an end date are how most projects drift into the 95% that never showed measurable impact. A month is long enough to surface the edge cases a single week hides and short enough that nothing has been built around it yet.
The most valuable thing in the MIT findings is not the failure rate. It is that the successes had a recognisable shape: back office, well-specified, owned, and measured. That shape is available to a two-person company, and it does not depend on buying anything in particular.
So do the unglamorous half first. Write the job down, name the owner, record the baseline, decide what needs holding back, and set the date you will decide. If you finish that and the case still holds, you are in a much better position than nineteen out of twenty pilots ever reach.