AgentBench: Evaluating LLMs as Agents
2023-08-07 · Verified
Paper info
Start research from this paper
Choose a research mode to carry this paper into the workspace context.
Reproducibility status
No structured reproducibility record has been added yet.
Model-role topology
This view only shows recorded model-role relations; it does not invent workflow edges.
This paper record has no workflow edges; the topology above only shows recorded roles and does not treat Actor, Environment, Critic, or Optimizer relations as facts.
Model roles
GPT-4
Baseline · weights not updated · GPT-4 benchmarked as an AgentBench agent.
GPT-3.5 Turbo
Baseline · weights not updated · GPT-3.5 Turbo benchmarked as an AgentBench agent.
Llama 2 70B Chat
Baseline · weights not updated · Llama-2-70B-Chat benchmarked as an open AgentBench agent.
Vicuna 13B
Baseline · weights not updated · Vicuna-13B benchmarked as an open AgentBench agent.