Effective evaluation of multi-hop tool use is critical for analyzing the understanding, reasoning, and function-calling capabilities of large language models (LLMs). However, progress has been hindered by a lack of reliable evaluation datasets. To address this, we present ToolHop, a dataset comprising 995 user queries and 3,912 associated tools, specifically designed for rigorous evaluation of multi-hop tool use. ToolHop ensures diverse queries, meaningful interdependencies, locally executable tools, detailed feedback, and verifiable answers through a novel query-driven data construction approach that includes tool creation, document refinement, and code generation. We evaluate 14 LLMs across five model families (i.e., LLaMA3.1, Qwen2.5, Gemini1.5, Claude3.5, and GPT), uncovering significant challenges in handling multi-hop tool-use scenarios. The leading model, GPT-4o, achieves an accuracy of 49.04%, underscoring substantial room for improvement. Further analysis reveals variations in tool-use strategies for various families, offering actionable insights to guide the development of more effective approaches. Code and data can be found in https://huggingface.co/datasets/bytedance-research/ToolHop.

本研究解决了大型语言模型（LLMs）在多跳工具使用评估中的可靠数据集缺乏问题，提出了ToolHop数据集，包含995个用户查询和3,912个相关工具。通过一种新颖的查询驱动数据构建方法，ToolHop确保了查询的多样性和工具的本地可执行性，为进一步提升LLMs的多跳工具使用能力提供了重要数据支持。

ToolHop: 一个用于评估大型语言模型多跳工具使用的查询驱动基准