Benchmark

OSWorld Verified

OSWorld is a benchmark designed to evaluate AI agents in real computer environments. These agents are AI systems capable of processing text, images, and other data, with benchmarks including open-ended tasks such as file management or software usage. OSWorld Verified is an upgraded version that provides more accurate testing results by addressing bugs and refining processes. It supports multiple operating systems, including Ubuntu, Windows, and macOS, and enables AI to learn through interactive tasks.

Limitations of Existing Benchmarks

Several current AI benchmarks rely on simulated environments rather than actual computers, which means their results don't accurately reflect real-world usage. Key issues include:

  • Simulated setups cannot handle diverse software or OS files.
  • Testing often requires manual checks, hindering automation and reproducibility.
  • Tasks are narrowly focused, overlooking complex multi-software workflows.

These drawbacks make it challenging to assess AI performance in everyday computer tasks reliably.

Origins and Goals

OSWorld was developed through a collaboration between the University of Hong Kong, Salesforce Research, Carnegie Mellon University, and the University of Waterloo. Its first version was released in 2024, with the research paper presented at NeurIPS 2024. OSWorld Verified launched on July 28, 2025, incorporating AWS cloud support for faster testing and resolving over 300 community-reported issues. The benchmark aims to overcome evaluation gaps by assessing AI agents in authentic computer settings. Specific goals include:

  • Creating a scalable environment that supports diverse operating systems.
  • Using automated scripts for result verification to minimize manual effort.
  • Designing open-ended tasks to evaluate AI's interface recognition, operational skills, and planning abilities.

Testing Methods and Workflow

OSWorld Verified leverages virtual machines to create realistic computer environments. AI agents analyze scenarios through screenshots and UI structure trees, then execute actions like mouse clicks or keyboard inputs. Outcomes are automatically evaluated via scripts that check file contents or software states. The benchmark comprises 369 tasks (361 when excluding eight that require network access for Google Drive). These tasks span multiple categories:

Task TypeCountExample
Chrome Browser46Browsing websites, adjusting settings
GIMP Image Editing26Modifying images
LibreOffice Calc47Calculating data
LibreOffice Impress47Creating presentations
LibreOffice Writer23Editing documents
Multi-Software Collaboration93Combining multiple programs to complete work
OS Operations23Managing files, configuring systems
Thunderbird Email15Sending and receiving emails
VLC Media Playback17Playing videos
VS Code23Writing code

The testing process involves:

  1. Setting up task initial states with VM snapshots and scripts.
  2. Allowing AI agents to take up to 100 actions.
  3. Automatically computing success rates via scripts.

It supports both local execution and cloud-based parallel testing, typically completing within an hour. To ensure consistency, 134 verification scripts are included, along with tools for manual checks.

Performance of Leading AI Models

Based on data from February 2026, here's how selected models performed on OSWorld Verified (success rates from single runs with up to 100 steps):

RankModel NameRelease DateSuccess Rate (%)Type
1Claude Sonnet 4.6 (Anthropic)2026-02-1772.5General-purpose
2Claude Opus 4.6 (Anthropic)2026-02-1772.7General-purpose
3Kimi K2.5 (Moonshot AI)2026-01-3063.3General-purpose
4GPT-5.3 Codex (OpenAI)2025-12 (approx.)64.7Specialized

Humans achieved a 72.36% success rate on the same tasks. AI performance has surged from an initial 12.24% in earlier versions.

Key observations include:

  • AI agents still struggle with screen element recognition and software operation, causing some failures.
  • Higher-resolution screenshots boost success by 5-10%.
  • Textual action histories outperform screenshot-only approaches.
  • While AI is sensitive to UI layout changes, it maintains consistency across different operating systems.

Additionally, models like ByteDance's Seed-1.8 reached 61.9%, highlighting the multi-task strengths of general-purpose models.

Value and Future Directions

OSWorld Verified advances AI agent development by leveraging authentic environments and automated evaluation. It uncovers flaws in task planning and execution, offering valuable data for refinement. Looking ahead, the benchmark could expand to additional operating systems and task types, while also facilitating research in learning methodologies and AI safety.

Comments (0)

Share:XHatena

Post a Comment

Loading...