
How to read this finding
These are separate benchmarks with different tasks and environments. Their percentages should not be averaged or interpreted as a universal probability of an agent completing office work. The report identifies continuing long-horizon failures.