Search the atlas

Esc to close · Cmd/Ctrl + K to open

Why these models were selected

Roles and weight updates come from the paper case record. Selection reasons are written only when the paper or code states them; unknown is never inferred as fact.

Paper model roles and selection-evidence status
ModelRoleWeights updatedSelection basis
GPT-3.5 TurbopolicyNoNot recorded: verify in the paper/code
text-davinci-003criticNoNot recorded: verify in the paper/code
Llama 2 70B ChatbaselineNoNot recorded: verify in the paper/code

CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

2023-05-19 · Verified

Paper info

Evolution targets
Self-correction, Tool use
Benchmark
HotpotQA, GSM8K, HumanEval
Category
Self-evolving agent, Reflection, Tool use
Model relations
3

Start research from this paper

Choose a research mode to carry this paper into the workspace context.

Reproducibility status

CodeNot verified
CheckpointNot verified
ConfigNot verified
EnvironmentNot verified

No structured reproducibility record has been added yet.

Model-role topology

This view only shows recorded model-role relations; it does not invent workflow edges.

Policy
GPT-3.5 Turbo
Critic
text-davinci-003
Baseline
Llama 2 70B Chat

This paper record has no workflow edges; the topology above only shows recorded roles and does not treat Actor, Environment, Critic, or Optimizer relations as facts.

Model roles

GPT-3.5 Turbo

Policy · weights not updated · GPT-3.5 Turbo as initial generator corrected via tool-interactive critique.

text-davinci-003

Critic · weights not updated · text-davinci-003 evaluated as a CRITIC backbone.

GPT-3.5

Llama 2 70B Chat

Baseline · weights not updated · Llama-2-70B-Chat evaluated as open-source CRITIC backbone.