Back to today's list

Evaluating LLM Usage for Efficient and Explainable Numerical and Classified Implicit Sentiment Analysis of Product Desirability

Sherri Weitl-Harms, John Hastings

Published Jun 25, 2026
Editorial review6.8
Relevance0.480
Freshness0.000

Why It Matters

What makes this one worth your time

Understanding implicit sentiment in product feedback can guide product development and marketing strategies, making this framework valuable for businesses seeking to leverage qualitative data.

A framework using LLMs for efficient sentiment analysis of product desirability from qualitative feedback.

Summary

The paper proposes a framework using large language models (LLMs) to analyze implicit sentiment in qualitative product feedback, aiming to quantify product desirability. It evaluates zero-shot numerical sentiment scoring and categorical sentiment classification, achieving high correlation and accuracy compared to expert labels. The framework also includes model confidence ratings and rationale explanations to enhance interpretability.

Key contributions

  • Development of a scalable framework for sentiment analysis using LLMs.
  • Demonstration of high correlation and accuracy in sentiment scoring and classification.
  • Inclusion of model confidence and rationale explanations for better interpretability.

Notable insights

  • LLMs can achieve high accuracy in sentiment analysis without explicit review scores.
  • Incorporating model confidence and rationale explanations enhances interpretability and trust.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2606.23701v1 Announce Type: cross Abstract: Qualitative product feedback can reveal nuanced user experiences, but its implicit sentiment is difficult to measure. This paper presents a scalable and interpretable framework that uses large language models (LLMs) to quantify product desirability from such data. Using two Product Desirability Toolkit (PDT) datasets from ZORQ and CARMA comprising 106 respondent term groupings with gold-standard human annotation, zero-shot continuous numerical sentiment scoring and categorical sentiment classification are evaluated without relying on explicit review scores. Across the datasets, LLMs generated numerical sentiment scores directly from qualitative responses and closely matched expert labels, achieving Pearson correlations up to 0.97 and classification accuracy up to 94%. LLMs maintained robustness even when handling data presented in multiple forms and consistently expressed high confidence. In contrast, lexicon-based and transformer baselines did not produce statistically significant results. Among the models tested, GPT-4o-mini achieved performance comparable to larger models at 94% lower cost, supporting scalable deployment. The framework also incorporates model confidence ratings and human-readable rationale explanations (xAI), improving interpretability, transparency, and trust while supporting practical use in product satisfaction assessment. In general, using the PDT tool as a survey method along with a cost efficient LLM for sentiment analysis has the potential to provide for product evaluation with results that are rich in terms of sentiment scores (both numerical and classified sentiment) and in terms of the high-level user impressions of the product that can be used to identify ideas for product development and improvement, as well as marketing ideas for target audiences.