NovelAPIBench: Diagnosing How A Code LLM Learns to Use Novel APIs
Abstract
NovelAPIBench evaluates how language models integrate external API knowledge through retrieval versus fine-tuning, finding that usage examples and signatures are key and that retrieval and tuning serve complementary roles.
Rapidly evolving software libraries require coding language models to use unfamiliar APIs. Yet failed solutions alone tell us little about what API knowledge models need or how to provide it. We present NovelAPIBench, an automated diagnostic framework that builds benchmarks around the API knowledge gaps of a target model. It identifies novel APIs, generates executable coding tasks with separately controllable knowledge components, and classifies failures into six categories. The framework can be rerun as models and libraries evolve. We study 21 libraries across five domains and six backbone models, with our primary benchmark comprising 1,670 tasks covering 856 APIs. Our experiments show that when supplied directly, usage examples provide the strongest standalone guidance for novel API acquisition, and combining them with signature descriptions approaches full-knowledge performance. Retrieving this knowledge from a pool of APIs yields lower accuracy than supplying only the target API's knowledge, even when retrieval finds the target. Fine-tuning on other APIs improves performance on unseen APIs mainly when external knowledge is supplied, helping models both select the target API and integrate it into the surrounding code. These findings highlight complementary roles for external API knowledge and the ability to apply it. By separating knowledge content, delivery, and use, NovelAPIBench helps diagnose what limits novel API use and guides adaptation to evolving libraries. Code and data are available at https://github.com/MAPS-research/NovelAPIBench.
Get this paper in your agent:
hf papers read 2606.03657 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper