So Many Opinions, So Many LLMs: Comparing Large Language Models to Traditional Machine Learning for Open- Ended Survey Analysis
Abdullah Akinde, Mariam Akinde, Rasheedat Emiola, Ahmed Akinsola
Why It Matters
What makes this one worth your time
Understanding the strengths and weaknesses of LLMs in qualitative analysis can guide researchers in choosing the right tools for large-scale survey data interpretation.
LLMs outperform traditional models in survey analysis but face challenges in consistency and explainability.
Summary
The paper compares the performance of large language models (LLMs) like OpenAI's GPT, Twitter-roBERTa-base, and Meta's LLaMA against traditional machine learning models for analyzing open-ended survey responses, focusing on tasks such as sentiment analysis and thematic classification.
Key contributions
- Comparison of LLMs with traditional machine learning models for survey analysis.
- Evaluation of model agreement, classification accuracy, and interpretability of reasoning.
Notable insights
- LLMs show superior accuracy in understanding complex mood and theme patterns.
- There are significant differences in how LLMs justify their predictions and apply category boundaries.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.11890v1 Announce Type: cross Abstract: Open-ended surveys offer valuable insights, but they are notoriously difficult to analyze at scale. Building on previous work that employed traditional machine learning to classify text ("So Many Responses, So Little Time: A Machine-Learning Approach to Analyzing Open-Ended Survey Data") [1], this study investigates how different large language models (LLMs) understand and analyze NSSE open-ended survey responses. We focus on several cutting-edge LLMSs-OpenAI's GPT series, Twitter-roBERTa-base model, and Meta's LLaMA-and compare their performance to the previous machine learning models in tasks like sentiment analysis and thematic classification. Our research analysis assesses model agreement, classification accuracy, and interpretability of reasoning. The findings reveal that current LLMs routinely beat classic machine learning models in classification accuracy, particularly in understanding complex mood and theme patterns in student replies. While LLMs have superior accuracy, they differ greatly in how explicitly and consistently they justify their predictions and apply category boundaries. These distinctions highlight crucial trade-offs when using LLMs for qualitative analysis: increased predictive strength comes with issues in consistency and explainability. Our findings illustrate the benefits and drawbacks of utilizing various LLMs for large-scale qualitative research, and we provide practical advice for researchers looking to balance automation and interpretive rigor.