Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models
Linkai Peng, Baorian Nuchged
Why It Matters
What makes this one worth your time
Understanding and improving the decision-making process in speech language models can lead to more accurate and reliable AI systems, particularly in paralinguistic tasks.
A diagnostic framework to identify and address errors in speech language models' decision rules and readout coverage.
Summary
The paper introduces a diagnostic framework to separate different types of errors in speech language models, specifically focusing on decision-rule misalignment and readout-coverage limitations. It uses a generation-aligned diagnostic ladder to analyze the gaps between emitted answers, option logits, and hidden states. The study finds that state decoding significantly outperforms generation and that a label-free logit correction can improve accuracy. The research also shows that emotion information can generalize across speakers and is not solely dependent on acoustic descriptors.
Key contributions
- Introduction of a diagnostic framework to separate decision-rule misalignment from readout-coverage limitations.
- Demonstration of state decoding outperforming generation by a significant margin.
- Proposal of a label-free logit correction to enhance accuracy.
Notable insights
- The use of a generation-aligned diagnostic ladder to dissect different error sources in speech language models.
- Label-free logit correction as a method to improve generated answer accuracy.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2608.06409v1 Announce Type: cross Abstract: Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. We introduce a generation-aligned diagnostic ladder that compares the emitted answer, the option logits, an affine readout of those logits, and a linear readout of the hidden state at the same answer token. Successive differences separate endpoint, decision-rule, and readout-coverage gaps. Across five systems and two emotion corpora, state decoding exceeds generation by 27.8 accuracy points on average, and both the decision-rule and readout-coverage gaps are positive in all ten conditions. A label-free logit correction improves generated accuracy in every condition, showing that part of the decision-rule gap is actionable. In rank-matched comparisons, emotion information outside the native readout generalizes to held-out speakers and survives controls for measured acoustic descriptors, but replacing the selected readout-external directions usually has little effect on emitted answers. These results distinguish information availability from behavioral use and localize performance losses across the decision rule and the state-to-answer readout.