Sistava

What to Expect From an AI Workforce

Guide — by Mahmoud Zalt

The standard a serious AI workforce has to meet: hiring, deliverables, bounded memory, approval gates, reach, and proof of what actually happened.

Set the bar before you go shopping

Most people evaluate this category backwards. They watch three demos, get impressed by whichever one demos best, and only afterwards work out what they actually needed. By then the standard has been written by the salesperson rather than the buyer.

The fix is unglamorous. Write down what a good outcome looks like first, in your own words, then hold every product against it. A demo can make almost anything look capable for four minutes. A written standard survives the fourth minute.

This article is that standard. Twelve expectations, each one a capability you should insist on, and for each one the reason it matters more than it sounds. Some of them are obvious. The ones buyers usually skip are the ones that decide whether this is still running in month three.

One warning about vocabulary. Nearly every tool in this space now calls itself an agent, an assistant or a workforce, and the words have stopped carrying information. Judge by what happens after the model produces text, because that is where the products separate.

Expect it to finish the job, not the sentence

This is the first expectation and the one everything else rests on. Most AI tools stop at the answer, and everything expensive happens after the answer. If a product hands text back to you and calls that done, it has moved the work rather than removed it.

Try it with a request you actually have. Ask a chat window for twenty qualified prospects and you get a list. You still open the CRM, create the records, write the emails, move the files, schedule the follow-ups, assign the work, review the result, and check nothing sensitive went out. Every one of those steps was yours before you asked, and every one of them is still yours.

That gap is the whole product category. An AI workforce is built for the after: it carries a request through research, planning, execution, communication, review, storage and follow-up, and reports back when the chain is finished rather than when the paragraph is.

Comparison

DimensionTraditionalWith Sista
What you get backText in a window, ready to be copiedRecords created, files produced, messages sent, work assigned
Who does the middleYou, in six other tabsThe employee, across the connected tools it has access to
What happens tomorrowYou remember to ask againIt runs on a schedule it owns, with no reminder from you
Where the output livesWherever you paste itIn the workspace, findable, attached to the task it came from
What is proven afterwardsThe chat scrollback, if you kept itAn activity record, a work journal and a cost line

Use that table as a test question in any sales call. Ask what the product does with its own output once it has produced it. If the honest answer is nothing, you are buying a faster way to write, which is a real thing to buy, but it is not a workforce and should not be priced like one.

Expect to hire a role, not configure a tool

The setup experience tells you how the product thinks. If step one is a blank canvas, a node graph and a list of triggers, you have bought an automation platform with a friendly coat on. The standard is that you pick a role, the way you would when staffing a job.

So expect a marketplace of roles you can hire an individual from, expect to be able to hire a complete team in one move, and expect a builder for a role that does not exist in the catalog yet. The catalog is a starting point, not a ceiling.

The detail buyers rarely think to ask for: expect to interview a candidate before you hire it. You talk to it first, exactly like a real interview, give it a scenario from your own week, watch how it reasons and what it asks for, then decide. Nobody hires a person off a job description alone, and there is no reason to do it here either.

Benefits

A real identity

A name, an avatar, a job title, a persona and a communication style. The same character shows up in email, in chat and in meetings, so your team learns who to ask.

How deeply it thinks

Reasoning depth as a setting on each employee. Routine work stays cheap and quick, and the work that deserves more thinking gets it, without you paying for depth everywhere.

An interview before the hire

Talk to a candidate, give it a live scenario, and decide afterwards. Judging a role by conversation is a better test than judging it by a feature list.

A status lifecycle

Active, paused, moved to the bench when the workload changes, or kept as a former team member with its history intact. Winding something down should not mean erasing the record.

The reason this expectation is high on the list is momentum. A product you hire from is usable on day one by someone who has never drawn a workflow diagram. A product you configure gets set up once, by one person, and quietly stops evolving the moment that person gets busy.

Hiring one employee is the easy part of the standard. The harder question is what happens when you have six, because six independent workers reporting into one human inbox is not a workforce, it is six more things to manage. How the product handles that is the next expectation, and it is where a lot of otherwise decent tools quietly fall over.

Expect a team to have an actual shape

A workforce should be organised, not stacked. Expect to hire a whole team, and expect that team to have an AI leader that delegates work to the members and reviews what comes back. You brief one place, and the coordination happens below you rather than through you.

That single capability separates a multi-agent system from a folder of chatbots, and most competitors cannot demonstrate it. Ask for it directly: who assigns the work, and who checks it before it reaches me?

Then expect the rest of the scaffolding a real team has. Team OKRs and KPIs, goals and written guidelines, an org chart of your workforce, and sprints with a goal, a weekly rhythm, a review and a written account of what got done. There is also a 3D office view on desktop, generated from your real teams, where each employee sits at a desk you can click and message in real time.

How you already organise peopleWhat the standard maps it to
A department with a managerA team with an AI leader that delegates and reviews
Quarterly objectives and monthly numbersTeam OKRs and KPIs, plus goals and guidelines
The way things are done around hereWritten team guidelines the employees work inside
A two week cycle and a reviewSprints with a goal, a rhythm and a written account
The org chart on the wallAn org chart of your workforce, and a desktop 3D office
Someone on leave or between projectsAn employee paused or moved to the bench, history intact

Two honest caveats on the office view, because vendors oversell this one. It runs on desktop, and its layout is generated from your real teams rather than decorated by you. It is a live view of your workforce, not a game you furnish.

Expect the work to cover your whole company

A narrow product solves one department and leaves you shopping again in a quarter. The standard is coverage across the functions a small company actually has, which in practice means over 150 live capabilities spanning sixteen areas of work.

Coverage is not the same as depth, so probe it. Pick the two functions that consume most of your week and ask for a walkthrough of each, end to end, using your own vocabulary. Breadth is easy to list on a website and hard to fake in a live account.

Expect deliverables you can open, and revisions in place

Output should arrive as real files, not as text you reformat. Expect Word documents, PowerPoint presentations, spreadsheets, PDFs, CSVs and images, plus a document editor and a presentation builder inside the workspace, and text extraction so an employee can read what you upload.

Then expect the harder version of the same thing: it should edit an existing file rather than rebuilding it. Anyone who has received a revision that was correct in substance and wrong in every visual detail knows why this is a capability and not a footnote.

The review loop is where this expectation earns its place. You should be able to highlight the exact sentence in a document, or pin a comment to a specific part of an image, and have the employee receive that comment with its context. No rewriting your own feedback as instructions, no explaining which paragraph you meant.

What a revision should look like

  1. The work lands as a file — A document, a deck, a sheet or an image, in a format the person receiving it can open without you converting anything.
  2. You mark up the exact spot — Highlight the sentence. Pin the comment to that part of the image. The context travels with the comment.
  3. It revises in place — The file is edited rather than regenerated, so everything you did not comment on comes back exactly as it was.
  4. The next round follows the same path — Feedback attaches to the same deliverable, and the version history stays with the work instead of living in your inbox.

If a vendor cannot show you this, watch what your week actually becomes. You get faster first drafts and an unchanged amount of formatting, reformatting and version chasing, which is often where the hours were in the first place.

Expect memory that stops where the role stops

This is the expectation buyers most often fail to set, and the one that hurts most when it is missing. A workforce that learns is valuable. A workforce that learns without a boundary is a data incident waiting for a busy week.

Expect it to learn from documents, websites, uploaded files, internal notes, past conversations, connected apps, databases, Notion and Google Drive. Then expect the limit: memory and knowledge should follow the employee's role. If an employee may not access a resource, it must not be able to reach that information through memory, through knowledge search, or by being asked in a different conversation. What it writes back to memory is scoped the same way, and changing the role changes the access with it.

The part to insist on is where that limit is enforced. It has to sit in the backend, below the model. Security cannot depend on asking a model to keep a secret, because an instruction is a request and a long conversation can drift away from it. An access rule that never puts the information in reach is a boundary.

Here is the question that settles it in one sentence. Ask a vendor how the system stops an employee from surfacing information it is not allowed to see. If the answer describes an instruction, a system prompt or careful wording, you have found the ceiling of that product. If the answer describes an access layer that refuses, keep talking.

Setting a standard this detailed raises a fair objection: who is going to configure all of it? A platform with this much surface would be a burden to set up by hand, which is exactly why the next expectation is that you should not have to.

That assistant matters more than it looks, so it gets its own section later. For now, treat it as part of the standard rather than a nicety: any platform broad enough to be worth buying is broad enough to be abandoned during setup, and the thing that prevents that is a control layer you can talk to in plain language.

Expect one workspace instead of scattered outputs

Work that lands in six different places is work you will spend your Fridays reassembling. Expect a company Drive, a content workspace, a built-in CRM, mailboxes, a company calendar with meeting workflows, tasks, routines, projects, sprints and dashboards, in one place rather than as separate purchases.

Expect global search across pages, employees, teams, tasks, projects, files, sprints and CRM contacts, so a half-remembered deliverable takes one search instead of four tools. And expect your employee to be able to take you there, opening the exact page or resource rather than describing where it lives.

Expect the work to repeat without being asked

Anything that only runs when you remember to ask is not capacity, it is a tool with extra steps. The standard is that ownership is durable: duties define what an employee owns, skills define what it can use, routines put recurring work on a schedule, and tasks and projects hold the one-off pieces.

Expect playbooks, which are the exact procedures and standards your company follows: how a report is structured, what happens before anything goes to a customer, the sequence you always run at the start of a project. A playbook should be strict when the process is the point, and it should also allow you to hand over the outcome and let the employee find the path.

Expect rules for behaviour in specific situations, working hours so the cadence matches your business rather than firing at random, and sprints with objectives. And expect a work journal, where each employee writes what it did, what it decided and what it hit. That journal is usually the fastest honest answer to what happened last week.

Expect it to have its own presence

An employee nobody outside your company can reach is a private tool. Expect every employee to get its own email address and to send and receive independently, so a customer can write to whoever handles their account and get a reply from that address.

Expect the other shape too: connect Gmail or Outlook and it sends and replies from your real inbox under your name, when the relationship is yours to hold. Beyond email, expect web chat, Slack, Telegram and a personal mailbox, plus meetings where you add a Zoom, Google Meet or Teams call to the calendar, pick which employee attends, and it joins, takes notes, listens, can speak, and turns the conversation into follow-up work.

Expect control, because control is what makes autonomy usable

This is the section to read twice. Every serious objection to AI doing real work is a control objection: what if it sends the wrong thing, spends too much, sees something it should not, or loops forever. A product that answers those with enthusiasm rather than mechanism is not ready for your business.

Notice that the objections are not really about intelligence. They are about consequence, and consequence is governed by gates, limits, checks and records. Judge this section harder than any other, because it is the one that decides whether you can leave the thing running while you sleep.

Benefits

Human approval before selected actions

Sending external email, publishing, spending money, deleting data, triggering a workflow, sharing confidential information, or anything over a threshold you set. Rules written in plain language, so approval depends on the situation rather than one global switch.

Guardrails

Protect sensitive information, enforce policy, filter unsafe behaviour, reduce prompt-injection risk, block information-boundary crossings, and prevent repeated or runaway actions. Where an answer would expose restricted information, it can refuse or ask approval for that one piece rather than being handed broad access.

Output evaluations

Define what a good result must contain and must avoid. Output is checked before it reaches you, and a failed check goes back for revision automatically. Your quality bar runs every time instead of when someone remembers.

Budget and model control

Daily or monthly spending limits, credit monitoring, and routing simple work to efficient models while reserving stronger ones for work that needs the thinking.

Observability you can actually read

Dashboards, an activity timeline, work journals, cost tracking and an action inspector, plus a view of what your employee was actually told, showing the real context it received before it acted.

An AI Mentor over the workforce

It watches the workforce, chases stalled work, resolves common blockers, retries where sensible, and escalates only when a human decision is genuinely needed.

Two of those deserve extra weight because almost nobody asks for them. Output evaluations turn your quality standard into something the system enforces on itself, which is the difference between reviewing everything and reviewing exceptions. And being able to see exactly what your employee was told is the only way to explain a strange result without guessing.

The approval design matters as much as its existence. A single on or off switch pushes you into one of two bad states: approving everything until you stop reading, or approving nothing and never getting the capacity you paid for. Rules written in plain language let the gate track the situation, which is how a real delegation works.

A practical way to test control before you commit: pick one process that touches the outside world, turn the approval gate on for every outbound action, and run it for two weeks. You will see exactly how often it wanted to do the right thing, and you will see it before a customer does.

That two week run also tells you something a demo cannot. It shows you which parts of your own process were never written down anywhere, because those are the parts the employee will ask about. Most companies discover their playbook by watching something try to follow it.

Autonomy is not a setting you turn up. It is what is left over once the gates, the limits and the record are good enough that you stop watching.

Mahmoud Zalt, founder of Sistava

Expect it to reach the software you already run

A capable worker with no access to your systems is a very expensive writer. Reach is the difference between advice and action, so treat integration count as a floor rather than a feature, and then ask what happens for the tools that are not on the list.

At a Glance

874
connected apps and services
150+
live capabilities you can hand to an employee
16
areas of work covered across the catalog

Beyond connectors, expect a REST API, MCP, A2A, and inbound and outbound webhooks, so the workforce can be reached by your own systems as well as reach into theirs. That two way street is what stops it becoming another island.

Then expect the abilities that do not depend on an integration existing at all: web search, website scraping, browser automation, computer control, screen vision, terminal commands, file organisation, data-entry automation, and meeting attendance with transcription.

The one to ask about specifically is using apps through your existing authenticated session. With permission, an employee can operate the same browser and desktop applications you already use, which means a tool does not need a good API, or any API, to be reachable. Every business has at least one system that will never appear on an integrations page.

One precision worth holding vendors to, in this space generally and here specifically: Shopify is the only guided store workflow, and every action that changes a connected Shopify store is approval-gated at runtime, enforced in the platform rather than left as a toggle you might forget. Reads run unattended. Nobody should imply that every store platform behaves the same way, because they do not.

Expect the platform to run itself for you

A standard this long has an obvious failure mode: the buyer agrees with all of it, signs up, opens an empty workspace, and closes the tab. Breadth is only an advantage if something else carries the setup, which is why the personal assistant belongs in the standard rather than in the marketing.

Expect it to set up the workspace, understand your company, hire and configure employees, create projects and tasks, connect apps, establish routines, find a file or a contact, navigate you anywhere in the product, review activity, and check progress and spending. Expect it to explain the platform itself when you are not sure what something does, which is the difference between a product you finish adopting and one you keep meaning to.

Where the standard does not apply

A standard is more credible when it names its own edges. Judgement calls that depend on knowing your market, your history and the people in the room are not what this changes. Neither is the decision about what your company should be doing next quarter.

What changes is how much of a week goes to the middle: the research, the drafting, the formatting, the record keeping, the chasing, the reporting. Hold the standard high on that work, keep the strategic calls yours, and be suspicious of any vendor who tells you the second category is also solved.

FAQ

What should I expect an AI workforce to actually do, beyond writing?

Finish work. Expect it to create records in a CRM, produce real files, send and receive email, publish to connected channels, attend meetings and turn them into follow-up work, run recurring routines on a schedule, and store the output where the rest of the work lives. If the product's output is text you then have to carry somewhere else, it is a writing tool. The useful test is asking what the product does with its own output after producing it.

How is this different from a chatbot with plugins?

Ownership and continuity. A chatbot responds to whoever is typing and forgets between sessions. A workforce is hired into a role, owns duties, follows playbooks, runs routines on a schedule, works inside a team with a leader that delegates and reviews, and keeps a work journal of what it did. The conversation is one way to reach it, not the only place it exists.

How do I stop an AI employee from sending something before I have seen it?

Approval gates, and you should insist on them being granular. Expect human approval before selected actions including external email, publishing, spending money, deleting data, triggering a workflow and sharing confidential information, with rules written in plain language so approval depends on the situation rather than one global switch. Output evaluations add a second layer by checking work against what it must contain and must avoid before it reaches you.

Can an AI employee see company information it should not have access to?

Not if memory is bounded by role, which is the standard to hold every vendor to. Memory and knowledge should follow the employee's role, so information it may not access cannot be reached through memory, through knowledge search, or by asking in a different conversation, and what it writes back is scoped the same way. The critical part is that the backend enforces this rather than a model being instructed to keep a secret.

How much should I expect to supervise it?

More at the start, less once your standards are written down. Early on you approve outbound actions and review output while playbooks and evaluations take shape. After that, supervision shifts to exceptions: the AI Mentor chases stalled work, resolves common blockers and retries where sensible, escalating only when a human decision is genuinely needed. Dashboards, activity timelines and work journals cover the rest.

What if the tool my company depends on has no integration?

That is exactly the case to test before buying. Beyond 874 connected apps and services there is a REST API, MCP, A2A, and inbound and outbound webhooks, plus browser automation, computer control, screen vision and terminal commands. With permission an employee can operate the same browser and desktop applications you already use, so an awkward internal system does not have to have a clean API to be reachable.

Should I expect real documents, or text I have to reformat?

Real documents. Word files, PowerPoint presentations, spreadsheets, PDFs, CSVs and images, with a document editor and a presentation builder in the workspace and text extraction so it can read what you upload. Expect it to edit an existing file rather than rebuilding it, and expect to review in place by highlighting an exact sentence or pinning a comment to part of an image.

How do I control what this costs me each month?

Through limits and routing rather than hope. Expect daily or monthly spending caps, credit monitoring, and reasoning depth as a per-employee setting so routine work stays cheap while harder work gets more thinking. Simple work can be routed to efficient models with stronger ones reserved for work that needs them, and cost tracking shows where the spend actually went. Plan details are on the pricing page, and a free tier exists to test the standard before committing.

How will I know what it did while I was not watching?

Expect a record, not a recollection. Dashboards, an activity timeline, per-employee work journals, cost tracking and an action inspector, plus a view of what the employee was actually told, showing the real context it received before acting. Sprints add a written account of each cycle, which is usually the fastest way to answer what happened last month.

You chat, they work

That is the whole standard compressed into four words. A chatbot produces an answer and hands the remaining nine steps back to you. A workforce moves the request through research, planning, execution, communication, review, storage, measurement and follow-up, and comes back when the chain is done.

So the question to carry into every demo is not whether it writes well. It is what happens in the hour after it writes, who is allowed to see what, what it does tomorrow without being asked, and what it can prove afterwards. Twelve expectations, written down before anyone starts selling to you, and the honest test is whether the product still looks good on the fourth one you check rather than the first.