Sistava

What AI Still Cannot Do: An Honest List of the Limits

Question — — by Mahmoud Zalt

A straight account of what AI still cannot do at work, why each limit exists, and how to design around them instead of finding out the expensive way.

We sell AI Employees, so an article about limits is an odd thing for us to write. We write it because the alternative is worse. People who buy on an inflated promise discover the ceiling in month two, and by then the damage is a bad experience they attribute to the whole category.

The limits below are not the ones that get patched next quarter. They are the ones that come from what this technology fundamentally is, and knowing them is what lets you use it confidently rather than nervously.

It helps to separate two kinds of limit. Some things AI does badly today and will do better next year, like handling messy interfaces or very long documents. Those are engineering problems, and betting against them is usually a mistake.

Other limits will not move, because they are about accountability, trust, and knowledge that does not exist in writing anywhere. This article is about the second kind. Everything an AI Employee does on Sistava is built around them, which is why approval gates on irreversible actions are a permanent design choice rather than a beginner setting.

At a Glance

6
Limits that are structural rather than temporary
0
Amount of responsibility AI can hold for an outcome
Always
How often an irreversible action needs a human on the trigger
Undo test
The one question that sorts safe work from unsafe work

What can AI still not do?

Six things, and they hold regardless of which product you buy. It cannot be accountable for a result. It cannot reliably signal its own uncertainty. It has no access to knowledge that was never written down. It cannot carry a relationship. It cannot read a situation. And it cannot tell you that your instructions were wrong.

Read them slowly, because each one implies a different design decision on your side. They are not reasons to avoid AI. They are the shape of the fence you build around it.

What does hitting a limit actually feel like?

It feels like something working beautifully for six weeks and then producing one output that makes you go cold. Aisha runs a boutique recruitment firm placing finance staff, mostly by relationship, mostly repeat clients. She handed over candidate screening and it went well for a month.

The AI Employee read every application, checked qualifications against the role, wrote a two paragraph summary, and ranked the shortlist. It saved her around five hours a week and the summaries were genuinely good. Then it ranked a candidate third who should have been first.

The candidate had a two year gap. On paper that reads as a risk and the AI scored it accordingly. Aisha knew, from a conversation eighteen months earlier, that the gap was caring for a parent and that this person was the best analyst she had ever placed. That fact existed in exactly one place, which was her memory.

Two things are true about that story. The AI did not fail, it did precisely what the brief said. And the brief could never have been complete, because you cannot write down every piece of context you carry about people you have known for years. She kept the screening and stopped using the ranking. That is the correct response to a limit.

Which limits are actually temporary?

Several of the things people list as permanent limits are just current engineering. Handling awkward websites, working through very long documents without losing the thread, staying reliable across a long chain of steps, and cost per task are all improving fast, and each has improved measurably in the last year.

It matters that you sort them correctly, because building your whole process around a limit that disappears is as expensive as ignoring one that does not. The table below is our honest split, and we would rather be wrong in public than vague.

LimitPermanent or temporaryWhat to do about it
Cannot be accountable for an outcomePermanentName a human owner for every process, always
Cannot reliably flag its own uncertaintyPermanent in kind, improving in degreeRequire a clickable source for anything factual
Does not know undocumented contextPermanent, but you can shrink itWrite the rules down. This helps your people too
Cannot hold a relationship or read a roomPermanentKeep the human conversations human
Struggles with messy interfaces and odd systemsTemporaryWork around it now, expect it to fade
Loses the thread over very long tasksTemporaryBreak work into smaller steps for now
Cost per task at high volumeTemporary, falling steadilyDo not redesign your business around today's price

If you want the other side of this ledger, the capability list rather than the limit list, we wrote it plainly in what AI can actually do for you. The two articles are meant to be read together.

How do I design around these limits?

You apply one test to every task: if this goes wrong, how expensive is undoing it. Cheap and reversible means the AI can run it unattended. Expensive or permanent means a human approves before it happens. That single question resolves most of the hard cases without any philosophy.

The rest is ordinary management. Write the rules down, keep a source requirement on facts, review most closely when things seem to be going well, and never let one process run for months without anybody actually reading the output.

Five habits that keep the limits from hurting you

  1. Apply the undo test to every task — Ask what it costs to reverse a mistake. Cheap and reversible runs freely. Expensive or permanent keeps a human on the trigger, permanently, no matter how good the track record is.
  2. Demand a source for every fact — If a number or a claim cannot be traced to a document, an email, or a page it just read, treat it as unverified. This one rule removes most of the risk from the uncertainty limit.
  3. Write down the rules that live in your head — Every exception, every special case, every customer who gets different treatment. This is the only way to shrink the undocumented context problem, and it makes onboarding humans easier too.
  4. Review hardest when it is going well — Month one gets attention automatically. Month four is when a plausible wrong answer slips through because nobody is really reading any more. Put a recurring spot check in the calendar.
  5. Keep the relationships human — Difficult conversations, negotiations, apologies, and anything with a long standing customer stay with a person. The AI can prepare the background. It should not be the voice.

None of these habits are exotic. They are what a decent manager does with a capable but new team member: clear instructions, checkable work, and a firm line around the things that cannot be undone.

The businesses that get burned are not the ones who trusted AI too much on day one. They are the ones who stopped paying attention in month four.

Does knowing the limits make AI less useful?

The opposite. Vague expectations produce cautious half use, where you never quite trust anything and never really save time. A clear fence lets you hand over everything inside it without a second thought, which is where the actual hours come from.

Most of what a small business loses time on sits comfortably inside the fence. Chasing, sorting, drafting, summarizing, and reporting are all reversible, checkable, and rule based. The limits bite on maybe a tenth of the work, and that tenth was always going to be yours.

If your worry is more personal than operational, the question of what this means for your own job is a fair one and it deserves a straight answer rather than reassurance. We wrote that one separately.

FAQ

What can AI still not do in a business setting?

Six things, and they are structural rather than temporary. It cannot take responsibility for an outcome, cannot reliably tell you when it is guessing, cannot know context that was never written down, cannot hold a relationship, cannot read an unspoken situation, and cannot tell you that your instructions were wrong. Everything else on the usual limits list is a version of one of these.

Why does AI sound confident when it is wrong?

Because confidence and accuracy are produced by the same process. The output is generated the same way whether the underlying information is solid or half remembered, so tone is not a signal you can use. The practical fix is a rule rather than a judgment call: anything factual must trace back to a document, an email, or a page it actually read, and anything it cannot source is treated as unverified.

Will these limits go away as models improve?

Some will and some will not. Handling messy interfaces, staying reliable across very long tasks, and cost per task are all engineering problems that have improved fast and will keep improving. Accountability, relationships, reading a room, and knowledge that exists only in someone's head are not engineering problems, so plan for those to stay. Sorting a limit into the wrong column is expensive in both directions.

What is the safest way to decide what to delegate?

Apply the undo test. Ask what it costs to reverse a mistake on that task. If reversing is cheap, such as a bad draft or a mislabeled email, let it run unattended. If reversing is expensive or impossible, such as money leaving, a contract signed, or a record deleted, keep a human approving before the action happens, permanently, regardless of how good the track record has been.

When is a limit most likely to bite?

In month four, not month one. Early on everyone reviews carefully because they do not trust it yet. Once the output has been good for weeks, attention drops, and that is exactly when a plausible wrong answer slips through. The countermeasure is a recurring spot check in the calendar rather than relying on anyone to keep paying attention voluntarily.

Should these limits stop me using AI at all?

No, and knowing them clearly usually increases how much you use it. Vague fear produces half hearted use where you never trust anything and never save real time. A clear fence lets you hand over everything inside it without hesitating. For most small businesses the limits only bite on around a tenth of the work, and that tenth was always going to stay with a person.

AI cannot be accountable, cannot flag its own guesses, cannot know what you never wrote down, and cannot carry a relationship. Build the fence there, hand over everything inside it, and check the work hardest when it seems to be going well. That is the whole honest position.