IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages
Sakshi Joshi, Dhruv Subhash Rathi, Sanskar Singh, Eldho Ittan George, R J Hari, Kaushal Bhogale, Mitesh M. Khapra
Why It Matters
What makes this one worth your time
Understanding how AudioLLMs utilize context can improve their accuracy and adaptability in multilingual and domain-specific applications.
IndicContextEval assesses context utilization in AudioLLMs across multiple Indic languages.
Summary
The paper introduces IndicContextEval, a benchmark designed to evaluate the context utilization capabilities of Audio Large Language Models (AudioLLMs) across eight Indian languages and 23 professional domains. It features a 7-level prompting framework that includes various contextual signals to assess whether models genuinely use context or rely on pre-learned knowledge.
Key contributions
- Introduction of a multilingual benchmark for evaluating context utilization in AudioLLMs.
- Design of a 7-level prompting framework to assess context grounding.
- Evaluation of five models to reveal differences in context utilization behavior.
Notable insights
- The use of a 7-level prompting framework to progressively introduce contextual signals is a novel approach to evaluate context utilization.
- Inclusion of adversarial prompts with incorrect entities to test model robustness against misleading context.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2606.19157v2 Announce Type: replace-cross Abstract: AudioLLMs enable speech recognition conditioned on textual prompts such as domain descriptions or entity lists. However, it remains unclear whether these models genuinely utilise such context or rely on parametric knowledge learned during pretraining. Existing benchmarks cannot answer this question because they evaluate transcription under fixed prompting conditions and rarely include explicit contextual inputs. We introduce IndicContextEval, a 56-hour multilingual benchmark of natural speech from 555 speakers across 8 Indian languages and 23 professional domains. We design a 7-level prompting framework that progressively introduces contextual signals, including metadata, natural-language descriptions, entity lists in English and native script, and adversarial prompts with incorrect entities. Evaluating five models reveals substantial differences in context utilisation behaviour, highlighting the need for explicit evaluation of contextual grounding in AudioLLMs.