Search the atlas

Esc to close · Cmd/Ctrl + K to open

Why these models were selected

Roles and weight updates come from the paper case record. Selection reasons are written only when the paper or code states them; unknown is never inferred as fact.

Paper model roles and selection-evidence status
ModelRoleWeights updatedSelection basis
Claude 2baselineNoNot recorded: verify in the paper/code
GPT-4baselineNoNot recorded: verify in the paper/code
GPT-3.5 TurbobaselineNoNot recorded: verify in the paper/code

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

2023-10-10 · Verified

Paper info

Evolution targets
Agent evaluation, Tool use
Benchmark
SWE-bench, SWE-bench-Lite
Category
Tool use, Reasoning, Self-evolving agent
Model relations
3

Start research from this paper

Choose a research mode to carry this paper into the workspace context.

Reproducibility status

CodeNot verified
CheckpointNot verified
ConfigNot verified
EnvironmentNot verified

No structured reproducibility record has been added yet.

Model-role topology

This view only shows recorded model-role relations; it does not invent workflow edges.

Baseline
Claude 2GPT-4GPT-3.5 Turbo

This paper record has no workflow edges; the topology above only shows recorded roles and does not treat Actor, Environment, Critic, or Optimizer relations as facts.

Model roles

Claude 2

Baseline · weights not updated · Claude 2 benchmarked on SWE-bench.

GPT-4

Baseline · weights not updated · GPT-4 benchmarked on SWE-bench.

GPT-3.5 Turbo

Baseline · weights not updated · GPT-3.5 Turbo benchmarked on SWE-bench.