Back to today's list

How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots

Eric Nichols, Alva Markelius, Hatice Gunes

Published Aug 10, 2026
Editorial review6.5
Relevance0.482
Freshness0.000

Why It Matters

What makes this one worth your time

Selecting the right foundation model is crucial for developing effective social robots, and a structured evaluation framework can guide researchers in making informed choices.

Proposes a community evaluation framework for selecting foundation models in social robotics.

Summary

The paper proposes a community-driven evaluation framework for selecting foundation models for social robots, identifying five key evaluation dimensions and suggesting a three-tiered evaluation funnel to streamline model selection.

Key contributions

  • Proposes a structured evaluation framework for foundation models in social robotics.
  • Maps evaluation dimensions across a three-tiered funnel.

Notable insights

  • Identifies five specific evaluation dimensions critical for social robots.
  • Introduces a three-tiered evaluation funnel to optimize model selection.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2608.06898v1 Announce Type: cross Abstract: Researchers who seek to build social robot applications on foundation models are faced with a difficult question: how should we pick a model? Public leaderboards offer little guidance: the demands of real-time, embodied social interaction lie largely outside their focus. And direct evaluation is impractical at scale: each embodied study requires scarce participant, robot, and experimenter time. In this paper, we identify five evaluation dimensions for foundation models in social robots: (i) conversational competence, (ii) user safety, (iii) embodied character, (iv) target scene effectiveness, and (v) audience appropriateness. To make model selection cheaper and better informed, we propose a three-tiered evaluation funnel paradigm that first filters with general metrics, then extends to simulated interactions, and terminates in more expensive, robot-specific evaluation. We map all five dimensions across all three tiers, chart where applicable evaluation methods exist and are missing, and close with a call to action: let's build the evaluation framework together as a community.