A Study of LLMs' Preferences for Libraries and Programming Languages
Lukas Twist, Mark Harman, Don Syme, Joost Noppen, Helen Yannakoudakis, Detlef Nauck, Jie M. Zhang
Why It Matters
What makes this one worth your time
Understanding LLMs' preferences can inform better model training and evaluation, leading to more effective code generation tools.
LLMs show a strong bias towards popular libraries and Python, often at the expense of task suitability.
Summary
The paper conducts an empirical study on large language models' (LLMs) preferences for libraries and programming languages in code generation, revealing a tendency to favor popular libraries and Python, even when not optimal.
Key contributions
- Empirical analysis of LLMs' library and programming language preferences.
- Identification of the overuse of popular libraries like NumPy in code generation.
- Insights into the implications of LLMs' preferences for model fine-tuning and evaluation benchmarks.
Notable insights
- LLMs prioritize familiarity over optimality in library and language selection, which may lead to suboptimal coding practices.
- The study highlights a significant gap in existing evaluations of LLMs that typically overlook design choices in code generation.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2503.17181v4 Announce Type: cross Abstract: Despite the rapid progress of large language models (LLMs) in code generation, existing evaluations focus on functional correctness or syntactic validity, overlooking how LLMs make critical design choices such as which library or programming language to use. To fill this gap, we perform the first empirical study of LLMs' preferences for libraries and programming languages when generating code, covering eight diverse LLMs. We observe a strong tendency to overuse widely adopted libraries such as NumPy; in up to 45% of cases, this usage is not required and deviates from the ground-truth solutions. The LLMs we study also show a significant preference toward Python as their default language. For high-performance project initialisation tasks where Python is not the optimal language, it remains the dominant choice in 58% of cases, and Rust is not used once. These results highlight how LLMs prioritise familiarity and popularity over suitability and task-specific optimality; underscoring the need for targeted fine-tuning, data diversification, and evaluation benchmarks that explicitly measure language and library selection fidelity.