FISHER: A Foundation Model for Multi-Modal Industrial Signal Comprehensive Representation
Pingyi Fan, Anbai Jiang, Shuwei Zhang, Xinhu Zheng, Zhiqiang Lv, Bing Han, Wenrui Liang, Junjie Li, Wei-Qiang Zhang, Yanmin Qian, Xie Chen, Jia Liu
Why It Matters
What makes this one worth your time
This work is significant for AI engineers and researchers focusing on industrial signal processing, offering a scalable and robust solution to data heterogeneity and sampling rate variability, potentially improving diagnostic accuracy and versatility in industrial applications.
FISHER is a foundation model that advances multi-modal industrial signal representation by addressing data heterogeneity and multi-sampling-rate challenges.
Summary
The paper introduces FISHER, a foundation model designed for comprehensive representation of multi-modal industrial signals, addressing the M5 problem of data heterogeneity. It employs a novel sub-band modeling approach to handle multi-sampling-rate issues and is pre-trained using teacher-student self-distillation on external audio and music data. The model is evaluated on the RMIS benchmark, outperforming existing state-of-the-art series encoders with smaller model sizes.
Key contributions
- Introduction of FISHER, a foundation model for multi-modal industrial signal representation.
- Development of a novel sub-band modeling approach for adaptive usage of full signal bandwidth.
- Establishment of the RMIS benchmark for evaluating multi-modal industrial signal models.
Notable insights
- The use of sub-band modeling to treat sampling rate increments as concatenated sub-band information is a clever approach to handle multi-sampling-rate issues.
- Leveraging audio and music data for pre-training provides better temporal variability, enhancing the model's generalization capabilities.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2507.16696v3 Announce Type: replace-cross Abstract: Industrial signal analysis is hindered by severe data heterogeneity, which we characterize as the M5 problem. Existing solutions rely on specialized models that lack robustness and scalability, while large-scale pre-training has rarely been investigated in this area. In this work, we derive a prioritized roadmap for the M5 problem and propose FISHER, a Foundation model for multi-modal Industrial Signal compreHEnsive Representation. To address the foremost multi-sampling-rate problem, FISHER utilizes a novel sub-band modeling approach that treats sampling rate increments as concatenated sub-band information, enabling the adaptive usage of full signal bandwidth without resampling. FISHER is pre-trained by teacher-student self-distillation over external audio and music data. We also establish the RMIS benchmark, comprising 19 datasets across four modalities. In the experiment, FISHER outperforms 24 state-of-the-art series encoders (up to 2B) with much smaller sizes (up to 16x), showcasing groundbreaking diagnostic accuracy and remarkable versatility. We further demonstrate that 1) seamless adaptation to variable sampling rates is the key to generalization 2) audio and music data provide better temporal variability, which is essential for pre-training. Both FISHER and RMIS are open-sourced.