Pailitao-MMSearch: Building Native E-Commerce Multimodal Search Foundation
Xiaohan Ye, Xu Chen, Zihan Gong, Jian Ding, Lianyu Du, Baicheng Chen, Yunmeng Shu, Jingqian Zhao, Zhixiang Zhao, Shuaiqi Jia, Chong Ma, Shuwen Xiao, Xiangheng Kong, Yuan Gao, Jun Song, Jinsong Lan, Xiaoyong Zhu, Bo Zheng
Why It Matters
What makes this one worth your time
This research is significant for improving user experience in e-commerce by enabling more effective and intuitive product searches, which can lead to increased sales and customer satisfaction.
Pailitao-MMSearch enhances e-commerce search by integrating multimodal inputs for better product retrieval.
Summary
The paper introduces Pailitao-MMSearch, a native e-commerce multimodal search foundation model that integrates text, images, and voice to improve product search, addressing limitations of existing single-modal and general-purpose models.
Key contributions
- Development of the Pailitao-MMSearch model for multimodal e-commerce search.
- Introduction of HybSID (Hybrid Semantic ID) for improved semantic understanding.
- Implementation of a two-stage continual pre-training strategy.
Notable insights
- The introduction of a two-stage continual pre-training strategy may allow for better adaptation to evolving user behaviors and product trends.
- The hybrid reasoning post-training pipeline could enhance the model's ability to understand complex user intents.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2607.17499v1 Announce Type: new Abstract: The evolution of e-commerce has fundamentally transformed how users search for products, shifting from simple text-based keyword queries to complex multimodal interactions that seamlessly combine product images, natural language descriptions, and mixed-intent instructions. However, existing approaches face a critical dilemma: single-modal specialist models, deployed independently for text retrieval, visual search, and voice recognition, operate in isolation and cannot handle cross-modal queries, while general-purpose vision-language models lack the domain-specific knowledge necessary for fine-grained product understanding, user behavior modeling, and commercial intent reasoning. In this work, we present Pailitao-MMSearch, one native e-commerce multimodal search foundation model designed to bridge this gap. Our approach introduces three key innovations: (1)HybSID (Hybrid Semantic ID);(2)a two-stage continual pre-training strategy; and (3)a hybrid reasoning post-training pipeline. Built upon Qwen and deployed on Taobao's Pailitao multimodal search platform, Pailitao-MMSearch achieves substantial improvements in online A/B testing, including up to +13.61\% in Gross Merchandise Volume (GMV) and +8.21\% in transaction volume compared to traditional multi-modal search pipeline, demonstrating the effectiveness of our native e-commerce multimodal search large language models.