We present PCA-Bench, a multimodal decision-making benchmark for evaluating the integrated capabilities of Multimodal Large Language Models (MLLMs). Departing from previous benchmarks focusing on simplistic tasks and individual model capability, PCA-Bench introduces three complex scenarios: autonomous driving, domestic robotics, and open-world games. Given task instructions and diverse contexts, the model is required to seamlessly integrate multiple capabilities of Perception, Cognition, and Action in a reasoning chain to make accurate decisions. Moreover, PCA-Bench features error localization capabilities, scrutinizing model inaccuracies in areas such as perception, knowledge, or reasoning. This enhances the reliability of deploying MLLMs. To balance accuracy and efficiency in evaluation, we propose PCA-Eval, an automatic evaluation protocol, and assess 10 prevalent MLLMs. The results reveal significant performance disparities between open-source models and powerful proprietary models like GPT-4 Vision. To address this, we introduce Embodied-Instruction-Evolution (EIE), an automatic framework for synthesizing instruction tuning examples in multimodal embodied environments. EIE generates 7,510 training examples in PCA-Bench and enhances the performance of open-source MLLMs, occasionally surpassing GPT-4 Vision (+3\% in decision accuracy), thereby validating the effectiveness of EIE. Our findings suggest that robust MLLMs like GPT4-Vision show promise for decision-making in embodied agents, opening new avenues for MLLM research.

PCA-Bench是一个用于评估多模态大型语言模型（MLLMs）综合能力的多模态决策基准，引入了三个复杂场景：自动驾驶、家庭机器人和开放世界游戏，并提出了误差定位能力和自动评估协议PCA-Eval对10种著名MLLM进行评估结果显示开源模型和GPT-4 Vision等强大专有模型之间存在显著性能差异，通过引入基于体验环境的自动框架Embodied-Instruction-Evolution（EIE），在PCA-Bench中生成了7,510个训练示例，并提高了开源MLLM的性能，偶尔超越GPT-4 Vision（+3％决策准确性），验证了EIE的有效性，发现GPT4-Vision之类的鲁棒MLLM对体验型代理的决策具有潜力，为MLLM研究开辟了新的道路。

PCA-Bench: 评估感知-认知-行动链中的多模态大型语言模型