Back to today's list

Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory

Jakub Rada (AI Center, Department of Computer Science, Faculty of Electrical Engineering, Czech Technical University in Prague), Viliam Lis\'y (AI Center, Department of Computer Science, Faculty of Electrical Engineering, Czech Technical University in Prague)

Published Aug 5, 2026
Editorial review6.5
Relevance0.488
Freshness0.000

Why It Matters

What makes this one worth your time

Improving LLMs' decision-making capabilities in sequential tasks could enhance their applicability in complex real-world scenarios where strategic planning is crucial.

The paper proposes an experience memory framework to enhance LLMs' sequential decision-making in games.

Summary

The paper investigates the performance of large language models (LLMs) in sequential decision-making tasks, specifically in fully-observable two-player zero-sum games. It identifies a performance gap in LLMs when compared to MCTS opponents and introduces an agentic framework with experience memory to improve decision-making without altering model weights.

Key contributions

  • Identification of performance gaps in LLMs for sequential decision-making tasks.
  • Introduction of an experience memory framework to address these gaps.
  • Demonstrated improvements in tic-tac-toe performance through post-game reflection and rule extraction.

Notable insights

  • Post-game reflection and rule extraction can improve LLM performance without changing model weights.
  • Obfuscations that preserve the game tree but alter its surface form do not significantly affect LLM performance, suggesting limitations in strategy recall.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2608.03420v1 Announce Type: new Abstract: Large language models have improved substantially on single-shot reasoning tasks, but their performance in sequential decision-making is less well understood. We study this on fully-observable two-player zero-sum games, which provide ground-truth evaluation: outcomes are determined by the rules, and optimality of individual moves can be computed or approximated, without relying on a judge model. Across model tiers, LLMs play suboptimally in simple games such as tic-tac-toe or Connect Four, and lose to MCTS opponents. Obfuscations that preserve the game tree but rewrite its surface form leave performance largely unchanged, indicating the gap is not fully explained by recall of memorized strategies. Motivated by this performance gap, we introduce an agentic framework enhanced with an experience memory designed for the sequential setting and addressing common challenges of sequential decision-making such as credit assignment. We show that post-game reflection and rule extraction yield measurable improvements on tic-tac-toe without modifying the model weights.