Search the atlas

Esc to close · Cmd/Ctrl + K to open

Why these models were selected

Roles and weight updates come from the paper case record. Selection reasons are written only when the paper or code states them; unknown is never inferred as fact.

Paper model roles and selection-evidence status
ModelRoleWeights updatedSelection basis
GPT-4policyNoNot recorded: verify in the paper/code
GPT-4VpolicyNoNot recorded: verify in the paper/code
Claude 3 OpusbaselineNoNot recorded: verify in the paper/code
Gemini 1.5 ProbaselineNoNot recorded: verify in the paper/code

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

2024-04-12 · Verified

Paper info

Evolution targets
Tool use, Multimodal policy, Actor policy
Benchmark
OSWorld
Category
Self-evolving agent, Tool use, Web agent
Model relations
4

Start research from this paper

Choose a research mode to carry this paper into the workspace context.

Reproducibility status

CodeNot verified
CheckpointNot verified
ConfigNot verified
EnvironmentNot verified

No structured reproducibility record has been added yet.

Model-role topology

This view only shows recorded model-role relations; it does not invent workflow edges.

Policy
GPT-4GPT-4V
Baseline
Claude 3 OpusGemini 1.5 Pro

This paper record has no workflow edges; the topology above only shows recorded roles and does not treat Actor, Environment, Critic, or Optimizer relations as facts.

Model roles

GPT-4

Policy · weights not updated · GPT-4 evaluated as an OSWorld computer-use agent.

GPT-4V

Policy · weights not updated · GPT-4V evaluated as a multimodal OSWorld agent.

Claude 3 Opus

Baseline · weights not updated · Claude 3 Opus benchmarked on OSWorld.

Gemini 1.5 Pro

Baseline · weights not updated · Gemini 1.5 Pro benchmarked on OSWorld.