The VNHSGE (VietNamese High School Graduation Examination) dataset, developed exclusively for evaluating large language models (LLMs), is introduced in this article. The dataset, which covers nine subjects, was generated from the Vietnamese National High School Graduation Examination and comparable tests. 300 literary essays have been included, and there are over 19,000 multiple-choice questions on a range of topics. The dataset assesses LLMs in multitasking situations such as question answering, text generation, reading comprehension, visual question answering, and more by including both textual data and accompanying images. Using ChatGPT and BingChat, we evaluated LLMs on the VNHSGE dataset and contrasted their performance with that of Vietnamese students to see how well they performed. The results show that ChatGPT and BingChat both perform at a human level in a number of areas, including literature, English, history, geography, and civics education. They still have space to grow, though, especially in the areas of mathematics, physics, chemistry, and biology. The VNHSGE dataset seeks to provide an adequate benchmark for assessing the abilities of LLMs with its wide-ranging coverage and variety of activities. We intend to promote future developments in the creation of LLMs by making this dataset available to the scientific community, especially in resolving LLMs' limits in disciplines involving mathematics and the natural sciences.

介绍了一个新的Vietnamese National High School Graduation Examination数据集，用于评估大型语言模型(LLMs)在多任务情况下的表现，其中包含文本和相关图像，并使用ChatGPT和BingChat对其进行评估，结果表明大型语言模型在文学、英语、历史、地理和公民教育方面能达到人类水平，但在数学、物理、化学和生物等领域还有提升的空间。该数据集旨在为评估LLMs的能力提供足够的基准，并督促未来更多的LLMs在数学和自然科学领域的发展。

VNHSGE: 用于大型语言模型的越南高中毕业考试数据集