Economic Evaluations of Language Models
Alexander Wan, Stephane Hatgis-Kessell, Tom\'as Aguirre, Percy Liang, Rishi Bommasani
Why It Matters
What makes this one worth your time
Understanding the economic impact of language models can guide their integration into the workforce, highlighting areas where they can save time and improve efficiency.
EconEvals measures the economic impact of language models on US labor tasks, revealing potential time savings.
Summary
The paper introduces EconEvals, an open-source evaluation suite designed to measure the economic impact of language models on tasks, work activities, and occupations in the US labor market. It claims to improve coverage over existing benchmarks at a lower cost and provides a simulation-based measure to estimate time savings across various tasks. The study highlights potential time savings but notes low usage of language models in practice due to privacy and proprietary constraints.
Key contributions
- Introduction of EconEvals, an open-source evaluation suite for economic impact assessment.
- Improved coverage over existing benchmarks at a significantly lower cost.
- Simulation-based measure for estimating time savings across tasks.
Notable insights
- The evaluation suite is grounded in real user queries and supplemented with synthetic data to improve coverage.
- A simulation-based exposure measure estimates potential time savings across US occupations.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.19375v1 Announce Type: cross Abstract: Language models perform economically valuable work, yet they are not currently assessed for how well they perform every economically valuable task. We introduce EconEvals as an open-source evaluation suite to measure capabilities relevant to tasks, work activities, and occupations in the US labor economy. We ground the evaluation suite in real user queries to language models where possible, and supplement these with synthetic data. Our evaluations improve coverage over OpenAI's GDPval benchmark, which is the existing state-of-the-art that covers 5% of US occupations, at 500x lower cost. Alongside benchmarks, we also introduce a simulation-based exposure measure to estimate how much time current language model capabilities could save across all tasks belonging to all US occupations, with detailed accounting for each estimate. Our estimates indicate that current models could save workers substantial time on at least half of their tasks in 47% of occupations. However, for 79% of tasks where we predict substantial time savings, observed Claude usage is low, suggesting that existing usage lags potential. Beyond inherent constraints of language model chatbots, our data identifies privacy and proprietary systems as the principal bottlenecks limiting further time savings from AI. Overall, we introduce adaptable infrastructure that grounds inferences about language models' labor-market impact in their current capabilities, which can be continually updated as capabilities improve.