PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows
Hongliang Li, Yijin Liu, Zhiwei Zhang, Zihe Liu, Xinyue Lou, Jinan Xu, Fandong Meng, Kaiyu Huang
Why It Matters
What makes this one worth your time
Understanding the impact of multilinguality on LLM agents is crucial for developing more robust and versatile AI systems capable of handling real-world multilingual workflows.
PolyWorkBench evaluates LLM agents on multilingual long-horizon tasks, revealing performance challenges.
Summary
The paper introduces PolyWorkBench, a benchmark designed to evaluate large language model agents on multilingual long-horizon tasks across various domains. It highlights the challenges of multilinguality in agentic execution and proposes a hybrid evaluation framework to assess functional correctness and linguistic consistency.
Key contributions
- Introduction of PolyWorkBench, a benchmark for multilingual long-horizon tasks.
- Empirical analysis showing performance degradation in multilingual settings.
- Proposal of a hybrid evaluation framework for comprehensive assessment.
Notable insights
- Multilinguality introduces compounding effects on reasoning and execution in LLM agents.
- A hybrid evaluation framework combining structural grading, executable verification, and semantic assessment is proposed.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.06008v3 Announce Type: replace Abstract: While Large Language Model (LLM) agents excel at monolingual long-horizon planning and tool use, enterprise workflows inherently require processing multilingual resources across extended trajectories. The interaction between multilinguality and long-horizon execution, however, remains underexplored. We introduce PolyWorkBench, a benchmark designed to evaluate LLM agents on multilingual, long-horizon workplace workflows. PolyWorkBench features 67 tasks across five core domains: commerce, knowledge work, legal analysis, localization, and manufacturing. Tasks are authored by the paper's authors from real-world data seeds and independently verified through a second-author audit. Agents must integrate heterogeneous multilingual inputs, execute iterative tool-use trajectories, and produce structured domain artifacts. To rigorously assess performance, we adopt Grade, a task-specific structural scoring rubric, as our primary ranking metric, and complement it with Pytest for executable state verification and LLM-as-Judge for semantic quality diagnostics. Benchmark evaluations reveal that agent performance varies substantially across languages and drops sharply on the harder cross-lingual tasks, and our analysis shows that multilingual execution exposes systematic failure modes across planning, tool interaction, and decision-making in long-horizon agents.