RESCAST-100K
收藏资源简介:
RESCAST-100K是由肯塔基大学研究团队创建的大规模住宅能源预测基准数据集,旨在解决跨领域泛化研究的空白。该数据集包含约10万个美国家庭的15分钟分辨率时间序列,涵盖总负荷、HVAC负荷和室内温度三大耦合目标变量,并整合了天气数据、HVAC设定点及40多项静态建筑协变量,数据总量达百万级时间点。数据集通过ResStock统计模型与EnergyPlus物理模拟相结合构建,采用高性能计算集群生成全年仿真数据,并统一整合了五个真实住宅数据集以支持仿真到现实的评估。该数据集主要应用于住宅能源管理、电网需求响应和社区能效项目,通过配置驱动接口支持地理、气候区、建筑围护结构和供暖设备等多维度领域划分,为机器学习模型在数据稀缺场景下的跨领域适应和零样本泛化研究提供系统化评估平台。
RESCAST-100K is a large-scale residential energy forecasting benchmark dataset developed by the research team at the University of Kentucky, aiming to fill the gap in cross-domain generalization research. This dataset contains 15-minute resolution time series from approximately 100,000 U.S. households, covering three coupled target variables: total load, HVAC load, and indoor temperature. It also integrates weather data, HVAC setpoints, and over 40 static building covariates, with a total of millions of time points across the dataset. The dataset is constructed by combining the ResStock statistical model and EnergyPlus physics-based simulations, with full-year simulated data generated via high-performance computing clusters, and unifies five real residential datasets to support simulation-to-reality evaluation. This dataset is primarily applied in residential energy management, grid demand response, and community energy efficiency projects. It supports multi-dimensional domain partitioning across dimensions such as geography, climate zones, building envelopes, and heating equipment via configuration-driven interfaces, providing a systematic evaluation platform for cross-domain adaptation and zero-shot generalization research of machine learning models in data-scarce scenarios.
数据集概述:RESCAST-100K
名称
RESCAST-100K: A Comprehensive Dataset for Cross-Domain Residential Load and Indoor Temperature Forecasting
发布机构与论文
- arXiv 预印本 ID: arXiv:2606.02852v1
- 作者:Jainam Dhruva, Yousaf Raza, A.B. Siddique, Simone Silvestri
- 提交日期:2026年6月1日
- 所属领域:计算机科学 > 机器学习 (cs.LG)
背景与目的
- 精准的短期住宅能耗负荷与室内温度预测对家庭能源管理系统、电网需求响应及社区能效提升至关重要。
- 现有住宅数据集覆盖范围窄,且难以支持结构化的跨域评估。
- RESCAST-100K 旨在为跨域泛化研究提供大规模住宅预测基准,推动迁移学习、域适应和零样本泛化的系统评估。
数据集规模与来源
- 覆盖约 100,000 个美国住宅,基于 EnergyPlus 模拟数据,源自 ResStock 项目。
- 时间序列分辨率为 15 分钟,对于每个住宅包含三个耦合目标变量:总负荷、HVAC 负荷和室内温度。
- 配套数据包括天气通道、HVAC 设定点以及超过 40 个静态建筑协变量。
核心特性
- 提供配置驱动接口,可沿可解释轴(如地理区域、气候区、墙体构造、供暖设备)实例化源域与目标域。
- 支持在受控域偏移条件下,系统评估迁移学习、域适应和零样本泛化。
- 集成 5 个真实世界住宅数据集,采用统一 schema,支持 sim-to-real(模拟到真实)评估。
评估与基线
- 基准测试涵盖循环神经网络、注意力机制模型和 MLP-Mixer 架构,评估内容:跨域零样本性能、缺失输入条件及多预测任务。
- 实验表明:跨注意力模型和 MLP-Mixer 模型在域偏移下持续优于循环网络和经典 Transformer 基线。




