API Pack: A Massive Multi-Programming Language Dataset for API Call Generation
CoRR(2024)
摘要
We introduce API Pack, a massive multi-programming language dataset
containing more than 1 million instruction-API call pairs to improve the API
call generation capabilities of large language models. By fine-tuning
CodeLlama-13B on 20,000 Python instances from API Pack, we achieved around 10
and 5
generating unseen API calls. Fine-tuning on API Pack enables cross-programming
language generalization by leveraging a large amount of data in one language
and small amounts of data from other languages. Scaling the training data to 1
million instances further improves the model's generalization to new APIs not
encountered during training. We open-source the API Pack dataset, trained
models, and associated source code at https://github.com/zguo0525/API-Pack to
facilitate further research.
更多查看译文
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
数据免责声明
页面数据均来自互联网公开来源、合作出版商和通过AI技术自动分析结果,我们不对页面数据的有效性、准确性、正确性、可靠性、完整性和及时性做出任何承诺和保证。若有疑问,可以通过电子邮件方式联系我们:report@aminer.cn