All the major AI labs are developing AI software agents that can operate a computer just like a person, visually parsing pixels, moving the mouse, and pressing keys. These are Computing User Agents (CUAs), and I wrote in-depth about them last week.
Today’s best CUAs are able to complete about 45% of the tasks in the popular OSWorld benchmark, up from just 6% when the benchmark was created sixteen months ago. What happens when they reach 100%?
In this post, I’ll explore what these benchmarks really measure, what they leave out, and how to prepare for the moment that AI UI execution becomes a solved problem.



