Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents
Yuanhong Jiang, Jingjie Zou, Rui Jiang, Zhenghong Lin, Xusheng Yu, Qiqi Huang, Shuai Jia, Shijie Dai
Why It Matters
What makes this one worth your time
Understanding and improving the decision-making process of financial agents can lead to more reliable and personalized financial advice, which is crucial for investors with diverse goals and risk profiles.
InvestLogicBench evaluates LLMs on personalized investment logic using real-world investor data.
Summary
The paper introduces InvestLogicBench, a benchmark designed to evaluate the investment logic of large language models by using a dataset of documented decisions from real-world investors. The benchmark aims to assess the reasoning and decision-making process of financial agents rather than just their outcomes.
Key contributions
- Introduction of InvestLogicBench, a benchmark for evaluating investment logic in LLMs.
- Analysis of logical plausibility versus event grounding in financial decision-making by LLMs.
Notable insights
- The benchmark uses a P→E→R→D→O trace to capture the full decision-making process, emphasizing reasoning over mere outcomes.
- The study reveals a discrepancy between logical plausibility and event grounding in LLMs, highlighting the need for better grounding in real-world data.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2608.06108v2 Announce Type: replace Abstract: Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries. Yet financial LLMs are evaluated either by static question answering or by terminal profit and loss. The former omits agency; the latter cannot reveal whether a profitable action was grounded, profile-consistent, or merely lucky. We ask whether the community is using the wrong ruler for consequential agents. We introduce \textsc{InvestLogicBench}, a process-native benchmark containing 201,247 documented decisions from 151 real-world investors. Each episode instantiates a \textbf{P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O} trace: investor \textit{Profile}, observable market \textit{Events}, investment \textit{Reasoning}, executable \textit{Decision}, and delayed \textit{Outcome}. The release includes profile construction, point-in-time event binding, structured logic, horizons, outcomes, and post-mortems, and supports comprehension, profile-conditioned generation, and end-to-end replay. Across four leading LLMs, logical plausibility remains near 4/5 while event grounding is only 0.8--2.8/5; return and process quality also disagree. These results expose polished but weakly grounded reasoning that outcome-only evaluation hides. We further argue that P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O should be a data-system interface, requiring versioned profiles, temporal provenance, inspectable retrieval, decision ledgers, and replayable outcomes. Finance is our stress test for a broader class of personalized, consequential agents.