全站搜索

按 Esc 关闭 · Cmd/Ctrl + K 打开

论文—模型矩阵

论文Claude 2Claude 3.5 SonnetClaude 3 OpusGemini 1.5 ProGLM-5.2GPT-3.5 TurboGPT-4GPT-4VGPT-4oLlama 2 13B ChatLlama 2 70B ChatQwen 7B ChatQwen2.5-3B-InstructQwen2.5-7B-InstructQwen3-1.7Btext-davinci-003Vicuna 13B
AdaPlanner: Adaptive Planning from Feedback with Language Models策略
AgentBench: Evaluating LLMs as Agents基线基线基线基线
AgentBoard: An Analytical Evaluation Board of LLM Agents基线基线基线基线
AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors执行者批评者
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Agents基线基线策略
CAMEL: Communicative Agents for Mind Exploration of Large Scale Language Model Society执行者
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models执行者策略
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing策略基线批评者
DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines策略优化器
ExpeL: LLM Agents Are Experiential Learners反思器策略
Generative Agents: Interactive Simulacra of Human Behavior策略
MetaGPT: Meta Programming for Multi-Agent Collaborative Framework策略
OpenHands: An Open Platform for AI Software Developers as Generalist Agents执行者策略
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments基线基线策略策略
Reflexion: Language Agents with Verbal Reinforcement Learning策略
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning分析器执行者执行者执行者
Self-Refine: Iterative Refinement with Self-Feedback策略批评者
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering策略
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?基线基线基线
Voyager: An Open-Ended Embodied Agent with Large Language Models策略
WebArena: A Realistic Web Environment for Building Autonomous Agents基线基线基线基线