OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
2024-04-12 · Verified
Paper info
Start research from this paper
Choose a research mode to carry this paper into the workspace context.
Reproducibility status
No structured reproducibility record has been added yet.
Model-role topology
This view only shows recorded model-role relations; it does not invent workflow edges.
This paper record has no workflow edges; the topology above only shows recorded roles and does not treat Actor, Environment, Critic, or Optimizer relations as facts.
Model roles
GPT-4
Policy · weights not updated · GPT-4 evaluated as an OSWorld computer-use agent.
GPT-4V
Policy · weights not updated · GPT-4V evaluated as a multimodal OSWorld agent.
Claude 3 Opus
Baseline · weights not updated · Claude 3 Opus benchmarked on OSWorld.
Gemini 1.5 Pro
Baseline · weights not updated · Gemini 1.5 Pro benchmarked on OSWorld.