Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation
Kevin Miller, Arjun Chandra, Venkatesh Saligrama
Why It Matters
What makes this one worth your time
Understanding whether ALMs use paralinguistic cues is crucial for improving their reliability and effectiveness in speech-to-speech systems, impacting their deployment in real-world applications.
The paper proposes counterfactual audits to assess if audio-language models use paralinguistic evidence effectively.
Summary
The paper introduces counterfactual audits to evaluate whether audio-language models (ALMs) use paralinguistic evidence when judging speech-to-speech systems. By holding transcripts constant and varying paralinguistic features, the study assesses the models' reliance on audio cues over lexical content. The research identifies different failure modes in ALM judges and suggests that accuracy alone is insufficient for evaluation, advocating for comprehensive behavioral audits.
Key contributions
- Introduction of counterfactual audits for evaluating paralinguistic evidence use in ALMs.
- Identification of different failure modes in ALM judges through detailed diagnostic states.
Notable insights
- Counterfactual audits can reveal whether ALMs rely on paralinguistic cues rather than just lexical content.
- Similar aggregate accuracies in ALMs can mask different failure modes, necessitating more nuanced evaluation methods.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2608.06718v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paralinguistic evidence. We introduce counterfactual audits for paralinguistic response evaluation. Each audit item holds the transcript fixed while varying affect, prosody, or the timing of an affective shift, forcing a valid judge to track the audio cue rather than lexical content or response style. We evaluate ALM judges using a native one-context judgment protocol and a contrastive recoverability control, then further decompose each item into its constituent perception and response-mapping skills. This yields useful diagnostic states that identify different sources of judge failures. Across Gemini, GPT, and open audio models, we find that contrastive success often overstates native judge reliability, and that similar aggregate accuracies can hide different failure modes. These results suggest that ALM judges should not be evaluated by accuracy alone, instead requiring thorough behavioral audits before deployment.