Sistava

Review the Quality of Every Deliverable

After an employee does real work, a quality reviewer scores each deliverable against a standard and shows the verdict in the activity inspector. Everyday chat is left alone, so you only grade real output.

Every real deliverable your employees produce, a written file, an email sent, a document drafted, a schedule created, a post published, gets scored against a quality standard by a reviewer, and the verdict shows up in the activity inspector right next to everything else that ran. Documents are read in full before being scored, not judged by filename, and emails your employees send are reviewed too. The reviewer covers twelve deliverable types in total: email, file, schedule, post, asset, message, task, sprint, report, research, plan, and summary, each against its own standard, plus a generic fallback for anything that does not map cleanly to one of those.

Read every verdict with the reason behind it, filter your activity to good, needs work, or bad, and see which employee produced each piece and when. Everyday back-and-forth chat is never reviewed, so the score only ever reflects real work: a fixed list of deliverable-producing activity types is what makes it into the queue at all, so a greeting or a quick clarification never shows up with a grade attached.

Turn it on from the AI Evaluation page in Technical settings. It's off by default for every workspace, plan rank aside, because the reviewer is itself a paid model call layered on top of the work it's checking, so a workspace opts in deliberately rather than being billed for it silently. Where this differs from a spell-checker or a linter: the review also runs before some deliverables ship, not only after. Before an employee sends an email or message, or writes or updates a file, the same reviewer checks the draft first, and a failing draft is bounced back to the employee to revise, up to two revisions, before it ever reaches the send step. What you see afterward in the inspector is the after-the-fact record of every deliverable, sent or written, whether it passed clean or needed a fix along the way.

This is what closes the trust gap most AI tools leave open: you're not just told an AI employee finished a task, you're shown whether the work was actually good, with a reason, the same way you'd review a human hire's output before it goes out the door.

A Standard You Can See, Not a Black Box

The AI Evaluation page shows the quality standards a deliverable is graded against, in plain English, not a hidden scoring model. There's a card per deliverable type, covering email, file, schedule, post, asset, message, task, sprint, report, research, plan, and summary. Each starts as the platform default, marked as such, and an Edit turns it into your own version: add your own standards alongside the defaults if your business has criteria the platform wouldn't know to check on its own, brand voice, a required disclaimer, a house format.

Every score comes with a reason, not just a number. A 0 to 1 score maps to one of three verdicts: pass at 0.8 or above, needs work between 0.5 and 0.8, and bad below 0.5. Filter your activity to good, needs work, or bad, and see which employee produced each piece and when, so a pattern across an employee's work is visible, not buried one deliverable at a time. A reviews log on the same settings page tracks pass rate, average score, and recent failures per workspace, so a drift in quality shows up as a trend, not just a one-off surprise.

Only Real Work Gets Graded

Everyday back-and-forth chat, greetings, quick clarifications, is never reviewed. The reviewer only scores actual deliverables, drawn from a fixed list of activity types: an email or message sent, a file written or updated, a task or sprint created, a post published. A plain reply in the chat thread never produces one of those activity records, so it never enters the review queue at all. That keeps the signal meaningful instead of diluted by conversational noise.

The reviewer is a smaller, cheaper model than the one doing the actual work, run through the same billed gateway every employee call goes through, so a review never bypasses billing or hides its own cost. If the review call itself fails or times out for any reason, the deliverable ships unscored instead of the run failing, so turning this on never puts a customer-facing task at risk waiting on a grade.

Catch a Bad Employee Before It Costs You a Customer

Without a quality signal, the first time you find out an AI employee's work was off is when a customer or a colleague notices. Quality Evaluation surfaces that gap the moment the deliverable is produced, so you're reviewing a score and a reason instead of discovering a problem after it already went out.

For the specific case of an outbound email, message, or file, the check runs twice in different roles: before it ships, a failing draft sends the employee back to revise it, up to two rounds, so a wrong recipient or a leftover placeholder often never reaches the send step at all. After it ships, the same review is what you see in the activity inspector, the permanent record of what actually went out and how it graded, whether it needed a fix along the way or not.

How It Works

A second, cheaper model reads the finished work and grades it against a written standard

When an employee finishes a work turn, the platform looks at what it actually produced: a file, a sent email, a schedule, a post, a task, a plan, a report, and so on. Each of those deliverable types has its own plain-English standard, for example a good email addresses the right recipient, has a clear subject, invents no facts, and leaves no leftover placeholders like [NAME] or TODO.

A judge model, run through the same billed gateway every employee call goes through, reads the original request and the finished output side by side against that standard, reasons briefly, and returns a score with a one-sentence reason. The score maps to pass, needs work, or fail, and both the verdict and the reason land in the Safety section of the Activity Inspector for that run.

The judge only ever grades, it never edits, blocks, or retries the employee's work, and a plain conversational reply is never sent through it at all, only real deliverables are. If the grading call itself errors for any reason, the deliverable ships unscored instead of the run failing, so turning this on never puts a customer-facing task at risk.

Use Cases

Spot-check a busy employee's output without reading everything

Instead of opening every activity an employee ran this week, filter to needs work or bad and go straight to the handful of deliverables actually worth a second look, with the reviewer's reason already attached.

Catch a bad email before a client sees it happened

An employee sends dozens of emails a week on your behalf. Before send, a failing draft is bounced back to the employee to revise, up to two rounds, and the outcome, revised or not, is graded against a standard for tone, correctness, and completeness so it shows up as a scored entry in your activity history instead of surfacing only when a client mentions it.

Build trust in autonomous, unattended work

For employees running on a schedule with nobody watching in real time, the verdict is what tells you the unattended run actually held up to standard, not just that it finished without an error.

Tighten the bar for one deliverable type without touching the rest

If reports keep coming back too casual but everything else is fine, edit only the report standard on the AI Evaluation page. The other eleven deliverable types keep the platform default untouched.

FAQ

Is Quality Evaluation on by default?

No, it's off by default for every workspace. Turn it on from the AI Evaluation page in Technical settings. It requires the Builder plan or above, since each review is itself a paid model call.

Does it review casual chat, or only real deliverables?

Only real deliverables: emails, messages, files, schedules, posts, tasks, sprints, reports, research, plans, and summaries. A fixed list of deliverable-producing activity types is what enters the review queue, so everyday conversational chat never gets a score.

Can I see why a deliverable got a particular score?

Yes, every verdict comes with a written reason, visible in the activity inspector next to the score. Scores map to pass (0.8 and above), needs work (0.5 to 0.8), or bad (below 0.5).

Can I set my own quality standards, or only use the defaults?

Both. Default standards work out of the box for every deliverable type, and you can edit or add your own alongside them from the AI Evaluation page. An edited standard becomes your workspace's version and overrides the default for that type.

Does the review only happen after the work is done, or can it stop a bad email before it sends?

Both. Outbound emails, messages, and file writes are also checked before they ship: a failing draft is sent back to the employee to revise, up to two rounds, before the send actually happens. Everything you see afterward in the activity inspector is the full record either way, whether a revision was needed or not.

Does a failed review ever block or fail the employee's task?

No. The reviewer only judges, it never blocks a task outright. If a review call itself errors, the deliverable ships unscored rather than the run failing, so a reviewer outage never stops real work from getting done.

Where Review the Quality of Every Deliverable fits

Review the Quality of Every Deliverable is part of How you see what they did.

A real-time activity feed shows who is working on what right now. The step-by-step inspector traces every tool call, decision, and reasoning chain. Cost tracking breaks down spend per message, per employee, and per team. Walk through your live 3D office to see your workforce at a glance.

Read the guide

More in Monitoring

Explore