跳到主要内容
↗
模型图谱
智能体基础模型选择地图
开始研究
模型
论文
学习指南
对比
英文
◐
开始研究
模型
论文
学习指南
对比
英文
◐
搜索
⌘ K
全站搜索
×
按 Esc 关闭 · Cmd/Ctrl + K 打开
自进化智能体论文
论文记录单独维护,通过模型 ID 与供应层建立关系。角色比单一“使用了某模型”更重要。
论文—模型矩阵
论文
Claude 2
Claude 3.5 Sonnet
Claude 3 Opus
Gemini 1.5 Pro
GLM-5.2
GPT-3.5 Turbo
GPT-4
GPT-4V
GPT-4o
Llama 2 13B Chat
Llama 2 70B Chat
Qwen 7B Chat
Qwen2.5-3B-Instruct
Qwen2.5-7B-Instruct
Qwen3-1.7B
text-davinci-003
Vicuna 13B
AdaPlanner: Adaptive Planning from Feedback with Language Models
—
—
—
—
—
—
—
—
—
—
—
—
—
—
—
策略
—
AgentBench: Evaluating LLMs as Agents
—
—
—
—
—
基线
基线
—
—
—
基线
—
—
—
—
—
基线
AgentBoard: An Analytical Evaluation Board of LLM Agents
—
—
—
—
—
基线
基线
—
—
—
基线
基线
—
—
—
—
—
AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors
—
—
—
—
—
执行者
批评者
—
—
—
—
—
—
—
—
—
—
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Agents
—
基线
—
—
—
—
基线
—
策略
—
—
—
—
—
—
—
—
CAMEL: Communicative Agents for Mind Exploration of Large Scale Language Model Society
—
—
—
—
—
执行者
—
—
—
—
—
—
—
—
—
—
—
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
—
—
—
—
—
执行者
策略
—
—
—
—
—
—
—
—
—
—
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
—
—
—
—
—
策略
—
—
—
—
基线
—
—
—
—
批评者
—
DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
—
—
—
—
—
策略
—
—
—
优化器
—
—
—
—
—
—
—
ExpeL: LLM Agents Are Experiential Learners
—
—
—
—
—
反思器
策略
—
—
—
—
—
—
—
—
—
—
Generative Agents: Interactive Simulacra of Human Behavior
—
—
—
—
—
策略
—
—
—
—
—
—
—
—
—
—
—
MetaGPT: Meta Programming for Multi-Agent Collaborative Framework
—
—
—
—
—
—
策略
—
—
—
—
—
—
—
—
—
—
OpenHands: An Open Platform for AI Software Developers as Generalist Agents
—
执行者
—
—
—
—
—
—
策略
—
—
—
—
—
—
—
—
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
—
—
基线
基线
—
—
策略
策略
—
—
—
—
—
—
—
—
—
Reflexion: Language Agents with Verbal Reinforcement Learning
—
—
—
—
—
—
策略
—
—
—
—
—
—
—
—
—
—
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
—
—
—
—
分析器
—
—
—
—
—
—
—
执行者
执行者
执行者
—
—
Self-Refine: Iterative Refinement with Self-Feedback
—
—
—
—
—
策略
批评者
—
—
—
—
—
—
—
—
—
—
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
—
—
—
—
—
—
策略
—
—
—
—
—
—
—
—
—
—
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
基线
—
—
—
—
基线
基线
—
—
—
—
—
—
—
—
—
—
Voyager: An Open-Ended Embodied Agent with Large Language Models
—
—
—
—
—
—
策略
—
—
—
—
—
—
—
—
—
—
WebArena: A Realistic Web Environment for Building Autonomous Agents
—
—
—
—
—
基线
基线
—
—
—
基线
—
—
—
—
基线
—