Sistava

Let Your Employee See Your Screen

Your employee looks at your actual screen and works from what it sees, not just an API. It captures your desktop through the Sista desktop app, reads the pixels the same way you would, and can click, type, and scroll based on what it finds. That covers any app you have open, including native software, dashboards, and design tools that never had an integration.

Screen Vision is one of four capabilities that make up Computer Controller, alongside app control, file access, and terminal execution. When a task needs it, the employee takes a screenshot of your desktop, reasons over the pixels the same way you would look at your own screen, then decides what to click, type, or scroll next. It only fires when a task genuinely calls for it, not on a timer or a continuous feed, so an ordinary browser or API task never triggers a capture.

Each capture stitches every connected display into one desktop image, so a two-monitor setup is read as a single wide screenshot rather than two separate ones. The full-resolution PNG is uploaded to private, tenant-scoped storage and shows up in the employee's Drive tab so you can review exactly what it saw. What goes to the model is a separate, lighter copy: resized to a maximum width of 1280 pixels and re-encoded as JPEG at quality 85, the resolution both OpenAI and Anthropic recommend for reliable vision accuracy. Sending an unscaled Retina screenshot would cost more Credits for no gain in what the model can actually read from it.

This is what separates Screen Vision from an integration built against an API: it works on anything visible on screen, including native apps like Photoshop or Final Cut that were never built with an API for an AI to call. It differs from Browser Controller the same way. Browser Controller reads a web page's structure and clicks elements by name; Screen Vision reads pixels and clicks by coordinate, so it is the fallback for whatever is not a browser tab. It also has a quiet safety net most tools do not: if your employee's main model cannot read images at all, one middleware routes only that screenshot's model call to a separate, inexpensive vision model, so the employee does not go blind mid-task just because its main model is text-only.

None of this happens without the desktop app running on your own machine. Screen Vision works through the Sista desktop companion you install and sign in to, so nothing can be captured while that app is closed or you are not connected. It runs on macOS, Windows, and Linux, and the capability stays off for every employee until you switch it on for that specific employee from its Tools tab.

On a Mac there is one extra step nobody can take on your behalf. Screen recording permission has to be granted manually in System Settings, because macOS gives an app no way to ask for it in a dialog, and until it is granted a capture comes back showing only your wallpaper instead of your actual screen. After you grant it, the app has to restart before the first real capture works. The setup checklist inside the app tracks this and updates itself, so you can see when the permission is actually in place.

Every capture is stored privately and scoped to your workspace, served only to somebody signed in to it, and never exposed to the open internet. What goes to the model is a separate, lighter copy, which is why an oversized screenshot never costs you more without telling the model anything it could not already read. If you want it to stop entirely, turning the tool off for that employee or quitting the desktop app both end it on the spot.

Setup and Plan

Screen Vision requires the Sista desktop app, the same app that also powers app control, file access, and terminal execution. It is available from the Founder plan and above, and like every desktop capability it is off by default until you turn it on per employee from the Tools tab.

What It Does Not Do

Screen Vision is not a live video feed and it does not watch your screen continuously. Every capture is a single point-in-time screenshot the employee takes on demand, right before it needs to decide what to do next, then again afterward to confirm the result. If an interface is ambiguous, icon-only buttons or an unfamiliar layout, the employee may pause and ask you to clarify rather than guess from the pixels.

It is also not a way around a locked screen or somebody else's machine. Captures come only from the computer running the companion app you installed and signed in to, so an employee sees what you would see sitting at that desk, nothing on a colleague's laptop and nothing on a server somewhere. Close the app and the eyes close with it.

How It Works

The employee captures your screen, reads the pixels, and acts on what it sees.

When a task needs eyes, the employee takes a capture through the companion app and interprets it: text, layout, buttons, fields, charts. It reasons about what it sees the way a person reading the screen would.

Vision pairs with app control for the full loop: see the state, act, look again. Captures happen only through the companion you installed, and the capability can be disabled or removed at any time.

Nothing runs on a schedule. The employee decides a capture is needed as part of doing the task, takes it, then takes another afterwards to check the action actually landed, which is what makes the loop reliable rather than hopeful. If the screen is genuinely ambiguous, icon only buttons or a layout it has not seen before, it will often stop and ask you rather than guess and click something.

Use Cases

Native apps without an API

Nudge layers in Photoshop, adjust a timeline in Final Cut, or format a sheet in Excel, since none of these tools expose an API an integration could call instead.

Internal dashboards

Read a metrics dashboard or admin panel built only for a browser, with no export and no API, by looking at what is rendered on screen and acting on it directly.

Design review

Look at a Figma file or design mockup, then click and adjust based on what is visible, when the change needs a human eye on the actual layout rather than raw data.

Confirming a result

Take a second screenshot after acting, so the employee, and you, from the saved image in Drive, can verify the click or edit actually landed before moving to the next step.

Comparison

BeforeAfter
If there is no API, there is no automation.If it renders on screen, it can be read and used.
Pixel-matching scripts break on any change.Meaning is read from the screen, not coordinates.
You describe what you see to get help.The employee looks for itself and answers.
Visual checks require a human.An extra pair of eyes on demand.

FAQ

What plan do I need for Screen Vision?

Founder plan or above. It is one of the desktop capabilities gated to that tier, alongside app control, file access, and terminal execution.

Does it work with two monitors?

Yes. All connected displays are captured together as one stitched desktop image, not a separate screenshot per monitor.

Where do the screenshots go?

Every full-resolution screenshot uploads to private, tenant-scoped storage and appears in the employee's Drive tab, so you have a record of exactly what it saw.

Does it cost more than other tools?

It draws from your normal Credit balance like any tool call. The copy sent to the model is downscaled to 1280 pixels wide as JPEG before analysis, which keeps the vision cost down without losing what the model needs to read the screen.

Can my AI employee see what is on my screen?

Yes, once you install the Sista desktop app and turn Screen Vision on for that employee. It takes a screenshot on demand when a task needs one, reads what is on the screen the way you would, then clicks, types, or scrolls based on what it found. It is available from the Founder plan and above and stays switched off until you enable it.

Does my employee watch my screen all the time?

No. There is no live video feed and no continuous monitoring. Each capture is a single point in time screenshot taken on demand, right before the employee decides what to do next and again afterwards to confirm the result, so an ordinary browser or API task never triggers one at all.

Can my employee work in an app that has no API, like Photoshop?

Yes, that is exactly what Screen Vision is for. It reads pixels rather than calling an interface, so a photo editor, a video editor, or an internal dashboard with no export is workable even though none of them were built for an AI to call. Browser Controller handles web pages by their structure; Screen Vision covers everything that is not a browser tab.

How do I stop my employee from seeing my screen?

Turn the tool off for that employee from its Tools tab, or quit the desktop companion app. Either one ends it immediately, and with the app closed nothing on your machine can be captured at all. On a Mac you can also revoke screen recording permission in System Settings.

Where Let Your Employee See Your Screen fits

Let Your Employee See Your Screen is part of Where their work lives.

Your AI agents manage their own workspace. Scheduled tasks run daily, weekly, or on custom cron cadences. A built-in kanban board tracks what is in progress. Every document lands in a personal Drive. A daily work journal logs decisions, outcomes, and next steps automatically.

Read the guide

More in Workspace

Explore