Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory
Jakub Rada (AI Center, Department of Computer Science, Faculty of Electrical Engineering, Czech Technical University in Prague), Viliam Lis\'y (AI Center, Department of Computer Science, Faculty of Electrical Engineering, Czech Technical University in Prague)
Why It Matters
What makes this one worth your time
Improving LLMs' decision-making capabilities in sequential tasks could enhance their applicability in complex real-world scenarios where strategic planning is crucial.
The paper proposes an experience memory framework to enhance LLMs' sequential decision-making in games.
Summary
The paper investigates the performance of large language models (LLMs) in sequential decision-making tasks, specifically in fully-observable two-player zero-sum games. It identifies a performance gap in LLMs when compared to MCTS opponents and introduces an agentic framework with experience memory to improve decision-making without altering model weights.
Key contributions
- Identification of performance gaps in LLMs for sequential decision-making tasks.
- Introduction of an experience memory framework to address these gaps.
- Demonstrated improvements in tic-tac-toe performance through post-game reflection and rule extraction.
Notable insights
- Post-game reflection and rule extraction can improve LLM performance without changing model weights.
- Obfuscations that preserve the game tree but alter its surface form do not significantly affect LLM performance, suggesting limitations in strategy recall.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2608.03420v1 Announce Type: new Abstract: Large language models have improved substantially on single-shot reasoning tasks, but their performance in sequential decision-making is less well understood. We study this on fully-observable two-player zero-sum games, which provide ground-truth evaluation: outcomes are determined by the rules, and optimality of individual moves can be computed or approximated, without relying on a judge model. Across model tiers, LLMs play suboptimally in simple games such as tic-tac-toe or Connect Four, and lose to MCTS opponents. Obfuscations that preserve the game tree but rewrite its surface form leave performance largely unchanged, indicating the gap is not fully explained by recall of memorized strategies. Motivated by this performance gap, we introduce an agentic framework enhanced with an experience memory designed for the sequential setting and addressing common challenges of sequential decision-making such as credit assignment. We show that post-game reflection and rule extraction yield measurable improvements on tic-tac-toe without modifying the model weights.