LLM Reasoning for Subjective Tasks: Failure Modes, Mitigation, and Dynamic Reasoning Routing
Juncheng Dong, Ding Tong, Ishan Gupta, Yuyan Wang
Why It Matters
What makes this one worth your time
Understanding and improving LLM performance in subjective tasks is crucial for their deployment in real-world applications where human preferences vary widely.
The paper addresses reasoning collapse in LLMs for subjective tasks by proposing a length-penalized algorithm and persona-based reasoning routing.
Summary
The paper investigates the challenges faced by Large Language Models (LLMs) in subjective verification tasks, identifying a phenomenon called 'reasoning collapse' where models resort to heuristic guessing. It proposes a conditional length-penalized post-training algorithm to improve verification accuracy and introduces a mid-training architecture that routes reasoning through contextually aligned personas.
Key contributions
- Identification of reasoning collapse in LLMs during subjective verification tasks.
- Development of a conditional length-penalized post-training algorithm to enhance verification accuracy.
- Introduction of a mid-training architecture for persona-based reasoning routing.
Notable insights
- Rigid, math-centric reasoning degrades LLM performance in subjective tasks.
- Verification accuracy is significantly influenced by the socio-linguistic framing of reasoning traces.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2608.08889v1 Announce Type: new Abstract: Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a matter of subjective human preference. As Large Language Models (LLMs) are deployed as autonomous verifiers of safety and quality guidelines, they face a distinctive challenge: context-aware preference alignment. Recent gains in Reinforcement Learning with Verifiable Rewards (RLVR) are indexed mostly on objective, mathematical tasks. Through a large-scale study spanning both proprietary and open-source models on four real-world verification tasks from a production recommender platform, we ask whether explicit reasoning generalizes to subjective, human-centric industry rubrics. We expose a fundamental vulnerability: rigid, math-centric reasoning traces actively degrade verification, and applying standard RLVR triggers a phenomenon we term reasoning collapse, in which the policy abandons deliberation in favor of rapid heuristic guessing. We introduce a conditional length-penalized post-training algorithm that intertwines verification accuracy with bounded reasoning length, halting collapse and recovering performance. Finally, we show that a reasoning trace's efficacy is tightly coupled with its socio-linguistic framing: across 1500 synthesized personas, verification accuracy swings by nearly 0.38 macro-F1 depending solely on the adopted reasoning persona---evidence that much subjective-verification error is really reasoning-style mismatch. This observation motivates a mid-training architecture that routes reasoning through contextually aligned personas. This work offers both a scalable algorithmic patch and a long-term architectural blueprint for aligning reasoning models with real-world subjective constraints.