Computer use

OSWorld-Verified open model results

Real-computer benchmark for multimodal agents completing open-ended tasks across desktop and web applications.

Selected open-weight/open-source systems from the official verified-results workbook. All rows use no accessibility tree, no coding action, one rollout, and the listed maximum step budget. Results from the older OSWorld release are intentionally excluded.
SystemModelScoreEvaluation scopeEvidence
OpenCUA-32B100 max steps; screenshot-only; single rollout; mean of 33.8%, 35.7%, and 34.8% OpenCUA-32BOpen-source system 34.8% OSWorld-Verified (three official runs) Organizer verified
UI-TARS-72B-DPO100 max steps; screenshot-only; single rollout UI-TARS-72B-DPOOpen-source system 27.1% OSWorld-Verified (361-task run) Organizer verified
OpenCUA-Qwen2-7B100 max steps; screenshot-only; single rollout Qwen2-VL-7B based OpenCUAOpen-source system 23.1% OSWorld-Verified (360-task run) Organizer verified
Qwen2.5-VL-72B-Instruct baseline100 max steps; screenshot-only; single rollout Qwen2.5-VL-72B-InstructOpen weights 5% OSWorld-Verified (361-task run) Organizer verified

Primary sources

Official benchmark Original paper Evaluation repository Data or result file