Back to today's list

Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models

Linkai Peng, Baorian Nuchged

Published Aug 10, 2026
Editorial review6.8
Relevance0.481
Freshness0.000

Why It Matters

What makes this one worth your time

Understanding and improving the decision-making process in speech language models can lead to more accurate and reliable AI systems, particularly in paralinguistic tasks.

A diagnostic framework to identify and address errors in speech language models' decision rules and readout coverage.

Summary

The paper introduces a diagnostic framework to separate different types of errors in speech language models, specifically focusing on decision-rule misalignment and readout-coverage limitations. It uses a generation-aligned diagnostic ladder to analyze the gaps between emitted answers, option logits, and hidden states. The study finds that state decoding significantly outperforms generation and that a label-free logit correction can improve accuracy. The research also shows that emotion information can generalize across speakers and is not solely dependent on acoustic descriptors.

Key contributions

  • Introduction of a diagnostic framework to separate decision-rule misalignment from readout-coverage limitations.
  • Demonstration of state decoding outperforming generation by a significant margin.
  • Proposal of a label-free logit correction to enhance accuracy.

Notable insights

  • The use of a generation-aligned diagnostic ladder to dissect different error sources in speech language models.
  • Label-free logit correction as a method to improve generated answer accuracy.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2608.06409v1 Announce Type: cross Abstract: Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. We introduce a generation-aligned diagnostic ladder that compares the emitted answer, the option logits, an affine readout of those logits, and a linear readout of the hidden state at the same answer token. Successive differences separate endpoint, decision-rule, and readout-coverage gaps. Across five systems and two emotion corpora, state decoding exceeds generation by 27.8 accuracy points on average, and both the decision-rule and readout-coverage gaps are positive in all ten conditions. A label-free logit correction improves generated accuracy in every condition, showing that part of the decision-rule gap is actionable. In rank-matched comparisons, emotion information outside the native readout generalizes to held-out speakers and survives controls for measured acoustic descriptors, but replacing the selected readout-external directions usually has little effect on emitted answers. These results distinguish information availability from behavioral use and localize performance losses across the decision rule and the state-to-answer readout.