FETS Benchmark: Foundation Models Enable Scalable and Generalizable Energy Time Series Forecasting
Marco Obermeier, Marco Pruckner, Florian Haselbeck, Andreas Zeiselmair
Why It Matters
What makes this one worth your time
This research highlights the potential of foundation models to provide scalable and generalizable solutions for energy forecasting, which is crucial for efficient energy system planning and operation.
Foundation models outperform traditional machine learning in energy time series forecasting.
Summary
The paper introduces the FETS benchmark to evaluate the performance of foundation models in energy time series forecasting, demonstrating their superiority over traditional dataset-specific machine learning models across various datasets and settings.
Key contributions
- Introduction of the FETS benchmark for energy time series forecasting.
- Demonstration of foundation models' superior performance over classical machine learning models.
- Analysis of performance factors such as spectral entropy and aggregation levels.
Notable insights
- Foundation models show strong performance even without full historic target data.
- Performance improves with higher aggregation levels and is correlated with spectral entropy.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2604.22328v2 Announce Type: replace-cross Abstract: Driven by the transition towards a climate-neutral energy system, accurate energy time series forecasting is critical for planning and operations. Yet, it remains a dataset-specific task, requiring comprehensive training data, limiting scalability, and resulting in high model development and maintenance effort. Recently, foundation models aiming to learn generalizable patterns via extensive pretraining have shown strong performance in multiple prediction tasks. Despite their success and strong potential in energy forecasting, a systematic, use-case-differentiated evaluation is still missing. We address this gap by presenting the Foundation Models in Energy Time Series Forecasting (FETS) benchmark. We (1) provide a structured overview of energy forecasting use cases along three main dimensions, i.e., stakeholders, attributes, and data categories, (2) curate 54 datasets across 9 data categories, guided by typical stakeholder interests, and (3) benchmark foundation models against task-specific machine learning across different forecasting settings. In our benchmark study, covariate-informed zero-shot foundation models perform best in aggregate, with Chronos-2 attaining the lowest overall median NRMSE (0.472), closely followed by TiRex-2 (0.474). Both perform better than XGBoost (0.611) and random forest (0.696), although they were trained task-specifically on the full historic target data. Further analysis reveals a strong correlation between predictive performance and spectral entropy. Performance saturates beyond a certain context length and improves with aggregation level, e.g., for national load, district heating, and power grid data. Overall, with the lowest median error, limited data requirements, and low inference and hardware demands, foundation models reduce development and maintenance effort, emerging as scalable and generalizable energy forecasting solutions.