Back to today's list

LangChoiceBench: Measuring and Explaining Programming-Language Choice in LLMs

Lukas Twist, Twm Stone, Helen Yannakoudakis, Jie M. Zhang

Published Aug 8, 2026Featured #4In the daily list Aug 9, 2026
Daily score62.3
Editorial review6.8
Relevance0.485
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding language biases in LLMs is crucial for improving their applicability across diverse programming tasks and ensuring they meet project-specific requirements.

LangChoiceBench reveals LLMs' strong bias towards Python and introduces 'phantom evidence' as a failure mode.

Summary

The paper introduces LangChoiceBench, a benchmark designed to evaluate programming language preferences in large language models, particularly focusing on the over-selection of Python. It assesses 25 models across 28 projects and finds a strong bias towards Python, low recommendation-implementation consistency, and limited language diversity. The study also identifies a failure mode termed 'phantom evidence' where models fabricate support for Python selection.

Key contributions

  • Introduction of LangChoiceBench for evaluating language preference in LLMs.
  • Analysis of 9,826 reasoning traces to understand language choice motivations.
  • Identification of 'phantom evidence' as a failure mode in language selection.

Notable insights

  • Models often choose Python automatically or for ease rather than based on project needs.
  • The concept of 'phantom evidence' where models fabricate support for language choice.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2608.06041v1 Announce Type: cross Abstract: Large language models (LLMs) have been shown to exhibit strong Python preferences when generating project-level code, but there is currently no systematic way to measure this behaviour across new models. To bridge this gap, we introduce LangChoiceBench, a project-level code-generation benchmark for measuring Python preference, recommendation-implementation consistency, and language diversity. LangChoiceBench covers 28 projects across seven software areas where Python is often a poor default. We evaluate 25 diverse LLMs and find that Python remains heavily over-selected, recommendation-implementation consistency is low, and smaller open-weight models generally show stronger Python preference and lower language diversity. We further analyse 9,826 reasoning traces and find that most Python choices are automatic or driven primarily by ease, rather than explicit consideration of project requirements. In a smaller but important set of cases, models fabricate contextual support for choosing Python - a failure mode we call phantom evidence - or produce code that contradicts the language selected in their own reasoning.