Computer use
OSWorld-Verified open model results
Real-computer benchmark for multimodal agents completing open-ended tasks across desktop and web applications.
Selected open-weight/open-source systems from the official verified-results workbook. All rows use no accessibility tree, no coding action, one rollout, and the listed maximum step budget. Results from the older OSWorld release are intentionally excluded.
| System | Model | Score | Evaluation scope | Evidence |
|---|---|---|---|---|
| OpenCUA-32B100 max steps; screenshot-only; single rollout; mean of 33.8%, 35.7%, and 34.8% | OpenCUA-32BOpen-source system | 34.8% | OSWorld-Verified (three official runs) | Organizer verified |
| UI-TARS-72B-DPO100 max steps; screenshot-only; single rollout | UI-TARS-72B-DPOOpen-source system | 27.1% | OSWorld-Verified (361-task run) | Organizer verified |
| OpenCUA-Qwen2-7B100 max steps; screenshot-only; single rollout | Qwen2-VL-7B based OpenCUAOpen-source system | 23.1% | OSWorld-Verified (360-task run) | Organizer verified |
| Qwen2.5-VL-72B-Instruct baseline100 max steps; screenshot-only; single rollout | Qwen2.5-VL-72B-InstructOpen weights | 5% | OSWorld-Verified (361-task run) | Organizer verified |