Efficiency degrades sharply
Getting to a 50% partial score demands disproportionately more tokens than getting halfway there. Effort does not scale linearly with progress, so the last stretch of a long task costs far more than the first.
Question — — by Mahmoud Zalt
Benchmarks show real partial progress on long professional work, and about one task in five finished unattended. Here is what that means for your week.
In OSWorld 2.0, a benchmark of 108 authentic tasks across 31 self-hosted web environments and seven professional domains, the best system scored 20.6% binary completion at a 500-step budget, with a 54.8% partial score. Other systems landed between 8% and 18% completion and 34% to 50% partial.
Those two columns say very different things, and the gap between them is the most useful fact in this article. Partial scores around half mean the work genuinely moves forward. Completion around a fifth means you cannot count on it crossing the finish line without someone looking.
It also tells you the shape of the tasks being measured. These are not one-line requests. Nearly seven in ten take a person more than an hour, and they average over 250 steps, which is the kind of work a week is actually made of.
So the honest headline is this: frontier agents make real measurable partial progress on long professional workflows and finish only about one in five of them unattended. That is neither the disappointment nor the miracle you may have been sold, and it has a very specific design consequence.
The benchmark names three failure patterns, and each one maps onto something you will recognise from your own week. None of them is mysterious once you see it written down.
Getting to a 50% partial score demands disproportionately more tokens than getting halfway there. Effort does not scale linearly with progress, so the last stretch of a long task costs far more than the first.
End-to-end completion collapses on the longest workflows. A task that is fine at fifty steps is a different proposition at three hundred, which is why chopping work into shorter units changes outcomes so much.
Agents struggle to maintain hidden state across steps, repeating work already done or executing from a plan that went stale several steps ago. This is the pattern that matters most for a week of work.
That third one deserves a moment. Hidden state is everything you are holding in your head that was never written down: the client changed their mind on Tuesday, the invoice is on hold until the contract lands, this supplier always replies late so chase early. Software loses that thread across a long run. A good assistant holds it without effort and without being asked.
Design the week around the completion number rather than the partial number. That means short jobs with a countable output, run often, each one small enough that a person can glance at the result and know immediately whether it is right.
In practice that looks less like a project and more like a rhythm. A daily inbox sweep that drafts replies and flags what needs you. A twice-weekly research pass that returns a written summary. A weekly report assembled from sources you already trust. Each of those is finishable, checkable, and useless to fake.
Add those up across five days and the volume is not small. A daily job running once each morning is twenty runs a month, each one landing something you can accept or reject in under a minute. The value shows up as accumulation rather than as any single impressive result, which is a quieter kind of progress than the demos suggest and a far more reliable one. It is also the shape a Sistava personal assistant is built around, a standing morning job that lands in your lap finished rather than an open-ended request you keep checking on.
The other advantage of the rhythm is that it fails safely. If Tuesday's run goes sideways you lose Tuesday, not the week, and you find out on Tuesday rather than on Friday when you finally open the folder. Short cycles convert a rare, expensive failure into a frequent, cheap one, which is almost always the better trade.
The instinct to hand over one big thing is understandable, because the big thing is what is actually costing you time. But the benchmark is clear that the big thing is exactly where completion collapses, and the fix is not a better prompt. It is cutting the big thing into segments that each end in something reviewable. You still get the whole job done. You just get it in pieces you can verify, which is the difference between output you use and output you re-do.
Structured that way, a week produces a genuinely useful pile: a cleared inbox with drafts waiting, a research summary you did not have to assemble, a report you only had to read. None of it is a heroic unattended marathon, and all of it is work you did not do yourself. The parts holding it together, Tool Rules, approval gates, memory that survives between runs, and a recorded activity feed, are ordinary Sistava features rather than scaffolding you have to build first.
Half of the published assistant task list stays where it is. Prioritising a messy inbox by what actually matters this week. Following up with a real person until they answer. Re-coordinating travel when a flight is cancelled. Preparing for a meeting where the point is reading the room. Writing an SOP for how this specific company works.
That is the hidden-state failure mode again, seen from the other side. Every one of those tasks is mostly about holding a thread across days and noticing when the premise changed, which is precisely the thing the benchmark measures software as being worst at. It is a technical finding about software, not a comment about anyone's colleagues.
Anything physical stays out too, along with anything that needs a legally accountable person to sign it. Those are not gaps a better setup closes. Planning around them is what makes the rest of the week realistic.
One week is also the right length for your first honest assessment. It is long enough to hit the edge cases that a single day hides, and short enough that you have not built a process around something that turns out not to work. Judge it on what you did not have to redo. If you want to run that week for yourself, Sistava is where you would start it, with one job rather than five.
On the OSWorld 2.0 benchmark of 108 authentic professional tasks, the best system finished about 20.6% of them end to end at a 500-step budget, while scoring 54.8% on partial progress. Others ranged from 8% to 18% completion. Read that as substantial real progress on long work, with unreliable unattended completion, which is why short segments with review points outperform one long run.
The benchmark names three patterns. Efficiency degrades sharply, so the last part of a long task costs disproportionately more effort. Task horizon creates hard limits, and end-to-end completion collapses on the longest workflows. And agents struggle to maintain hidden state across steps, repeating work already done or acting on a plan that went stale several steps earlier.
A job with a countable output that ends at a natural checkpoint: drafted replies, a written research summary, a list of flagged discrepancies, an assembled weekly report. Countable output means you can verify a run in seconds, which is what makes daily review sustainable. Avoid starting with the biggest, longest job even though that is the one costing you the most time.
Write the context down instead of expecting it to be held in flight. Constraints you keep repeating belong in Tool Rules, attached to a specific tool in plain English and binding on every run. Keep runs short so less state has to survive between steps, and rely on the employee's memory across runs to carry context rather than one long unbroken session.
Partly. Reads, research, and drafting can run unattended because a wrong result costs you a glance. Anything a customer will see, or anything that changes a system of record, belongs behind an approval gate that holds the action until you release it. Every action is recorded with a screenshot, so the review afterwards shows how a result was reached, not just what it was.
The judgment half of assistant work: prioritising by what actually matters, following up with a real person until they respond, re-coordinating when plans collapse, preparing for a meeting where reading the room is the point, and documenting how your company genuinely works. Anything physical and anything needing a legally accountable signature are also out, and no setup changes that.
Expectations are what make or break the first month. If you arrive expecting an unattended week of finished work, the benchmark says you will be disappointed. If you arrive expecting real progress on long work with checkpoints you review, you will get roughly what the research predicts, which is a lot.
Cut the work small, write the context down, gate the consequential actions, and read the feed. Do that and a week returns something concrete every time, which is a far better foundation than a single ambitious run that either works or quietly does not.