Key Takeaways
- A UC Berkeley study found that popular AI models scored below 25 percent on real-world professional tasks, highlighting their limited capabilities in executing complex workflows across various industries.
- The AI models, including OpenAI's ChatGPT-5.5, struggled with sustained reasoning and execution-heavy workflows, averaging a mere 2.6 percent success rate on the most challenging tasks.
- While currently ineffective at handling complex tasks, researchers believe that AI will increasingly automate repetitive, routine jobs, although decision-intensive roles may remain secure for longer.
Popular artificial intelligence language models scored below 25 percent when researchers from the University of California Berkeley tested their capabilities in more than 50 industries.
The Berkeley Center for Responsible, Decentralized Intelligence conducted the “Agent’s Last Exam,” the school’s real-world professional workflow test designed to determine whether AI is “job-ready,” according to the study.
“Today’s agents can solve a meaningful fraction of professional tasks. However, when we look at the hardest tasks that require sustained reasoning, deep domain expertise, and reliable execution over long horizons, they are still far from human-level performance,” it states.
“On ALE’s hardest tier, every frontier agent we tested, including Fable 5, achieved a 0% success rate,” the study states.
The test covered “more than 1,500 expert-sourced tasks spanning 55 occupations,” including finance, law, and manufacturing.
Fable 5, GPT-5.5, and Composer 2.5, among others, failed to complete more than one-fourth of the test correctly. Out of all of the models that were tested, OpenAI’s ChatGPT-5.5 model had the highest scores with a 24% passage rate, according to the study.
Yiyou Sun, a postdoctoral researcher who led the study with Berkeley Professor Dawn Song, told The College Fix that one of the most surprising elements of this study was that the AI models “still lack practical experience when it comes to execution-heavy workflows.”
When asked what skills AI might still struggle with, Sun said that its ability to choose how to complete tasks, as well as to change methods when failure occurs, remains difficult.
“It is still too early for us to draw conclusions about the rate of progress,” Sun added. “As researchers and model developers focus on these shortcomings, I expect many of the current limitations to improve substantially over time.”
Asked which industries would be most affected, Sun said, “The key factor is not the industry itself but the nature of the work.”
The repeatability of tasks, rather than the trade it’s conducted in, is a larger factor in the opportunity for automation through AI, Sun said.
“Highly repetitive workflows generate large amounts of data that are easier to collect, verify, and use for training AI systems. As a result, these tasks tend to be learned and automated more quickly,” Sun said.
Despite the low passage rates of current AI models, Sun believes repetitive human tasks will be replaced by AI.
“Even if current pass rates remain relatively low, occupations dominated by routine and well-defined procedures are likely to experience disruption first, while decision-intensive roles will remain more resilient for longer,” Sun said.
Similarly, Peyton Hornberger, a spokesperson for the Alliance for Secure AI, told The College Fix that “while AI may not be able to completely replace humans yet, the technology is trending towards a fully automated workforce.”
“AI is better at coding than virtually all humans, and that is just the beginning of its practical capabilities,” Hornberger said.
She called for “independent third-party evaluations by reliable entities, safety testing before deployment, and, most importantly, transparency to the public about the models’ capabilities.”
“There is a lot of fear around AI taking people’s jobs … Leaders in DC have an opportunity to set the standard for AI safety testing,” Hornberger said.
However, policymakers might not be getting the full picture on AI capabilities, as the lack of a standard benchmark makes it hard to assess industry-reported results, she said.
Ultimately, Hornberger recommended a set of three standards for future evaluations of AI systems, asking for “independent testing, transparency about limitations and capabilities, and transparency to the public about dangers and risks.”
The College Fix reached out to OpenAI and Anthropic for their response to the study. Neither responded.