SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
2023-10-10 · Verified
Paper info
Start research from this paper
Choose a research mode to carry this paper into the workspace context.
Reproducibility status
No structured reproducibility record has been added yet.
Model-role topology
This view only shows recorded model-role relations; it does not invent workflow edges.
This paper record has no workflow edges; the topology above only shows recorded roles and does not treat Actor, Environment, Critic, or Optimizer relations as facts.
Model roles
Claude 2
Baseline · weights not updated · Claude 2 benchmarked on SWE-bench.
GPT-4
Baseline · weights not updated · GPT-4 benchmarked on SWE-bench.
GPT-3.5 Turbo
Baseline · weights not updated · GPT-3.5 Turbo benchmarked on SWE-bench.