清华等开源「工具学习基准」ToolBench，微调模型ToolLLaMA性能超越ChatGPT

WBOY 2023-06-06 11:12 1422浏览转载

人类具有创造和利用工具的能力，使得我们可以突破身体的限制，探索更广阔的世界。

人工智能基础模型也类似，如果仅靠训练阶段得到的权重，使用场景就会非常受限，而最近提出的工具学习（tool learning），将特定领域的专用工具与大规模基础模型相结合，可以实现更高的效率、性能。

不过目前工具学习的相关研究还不够深入，也缺乏相关的开源数据和代码。

最近，清华大学自然语言处理实验室等支持的开源社区OpenBMB （Open Lab for Big Model Base）发布了ToolBench项目，可以帮助开发者构建开源、大规模、高质量的指令调优数据，促进构建具有通用工具使用能力的大型语言模型。

清华等开源「工具学习基准」ToolBench，微调模型ToolLLaMA性能超越ChatGPT

仓库链接：https://github.com/OpenBMB/ToolBench

ToolBench仓库中提供了相关数据集、训练和评估脚本，以及在ToolBench上微调的功能模型ToolLLaMA，具体特点为：

1. 支持单工具和多工具方案

其中单工具设置遵循LangChain提示风格，多工具设置遵循AutoGPT的提示风格。

2. 模型回复不仅包括最终答案，还包含模型的思维链过程、工具执行和工具执行结果

3. 支持真实世界级别的复杂性，支持多步工具调用

4. 丰富的API，可用于现实世界中的场景，如天气信息、搜索、股票更新和PowerPoint自动化

5. 所有的数据都是由OpenAI API自动生成并由开发团队进行过滤，数据的创建过程很容易扩展

不过需要注意的是，目前发布的数据还不是最终版本，研究人员仍然在对数据进行后处理来提高数据质量，并增加真实世界工具的覆盖范围。

ToolBench

ToolBench的总体思路是基于BMTools，在有监督数据中训练大型语言模型。

清华等开源「工具学习基准」ToolBench，微调模型ToolLLaMA性能超越ChatGPT

仓库中包含31.2万次真实API调用得到的9800条数据，涵盖单工具场景和多工具场景，下面是单工具的统计信息。

清华等开源「工具学习基准」ToolBench，微调模型ToolLLaMA性能超越ChatGPT

其中每行数据都是一个json dict，包含数据创建的提示模板、工具使用的人工指令（查询）、中间思维/工具执行循环和最终答案。

Tool Descrition:BMTools Tool_name: translationTool action: get_translationaction_input: {"text": target texts, "tgt_lang": target language}Generated Data:{"prompt": "Answer the following questions as best you can. Specifically, you have access to the following APIs:\n\nget_translation: . Your input should be a json (args json schema): {{\"text\" : string, \"tgt_lang\" : string, }} The Action to trigger this API should be get_translation and the input parameters should be a json dict string. Pay attention to the type of parameters.\n\nUse the following format:\n\nQuestion: the input question you must answer\nThought: you should always think about what to do\nAction: the action to take, should be one of [get_translation]\nAction Input: the input to the action\nObservation: the result of the action\n... (this Thought/Action/Action Input/Observation can repeat N times, max 7 times)\nThought: I now know the final answer\nFinal Answer: the final answer to the original input question\n\nBegin! Remember: (1) Follow the format, i.e,\nThought:\nAction:\nAction Input:\nObservation:\nFinal Answer:\n (2) Provide as much as useful information in your Final Answer. (3) Do not make up anything, and if your Observation has no link, DO NOT hallucihate one. (4) If you have enough information and want to stop the process, please use \nThought: I have got enough information\nFinal Answer: **your response. \n The Action: MUST be one of the following:get_translation\nQuestion: {input}\n Agent scratchpad (history actions):\n {agent_scratchpad}","query": "My intention is to convert the data provided in ما هي الأقسام الثلاثة للقوات المسلحة؟ into Arabic(ara).\n","chains": [{"thought": "I need to use the get_translation API to convert the text into Arabic.","action": "get_translation","action_input": "{\"text\": \"What are the three branches of the military?\", \"tgt_lang\": \"ara\"}","observation": "\"ما هي الفروع الثلاثة للجيش ؟\""}],"answer": "The translation of \"What are the three branches of the military?\" into Arabic is \"ما هي الفروع الثلاثة للجيش ؟\"."}