[论文解读] Analysis of Aggregated Functional Data from Mixed Populations with Application to Energy Consumption
本文提出一种函数混合效应模型,用于从汇总的、误报的数据中估计真实的消费者类别数量和类别特定的能耗模式。该方法采用B样条平滑与随机系数,并结合基于似然的推断方法,辅以一种新颖的数值技巧以处理报告误差,在真实和模拟数据中均表现出对真实消费者类型和能耗曲线的改进估计。
Understanding the energy consumption patterns of different types of consumers is essential in any planning of energy distribution. However, obtaining consumption information for single individuals is often either not possible or too expensive. Therefore, we consider data from aggregations of energy use, that is, from sums of individuals' energy use, where each individual falls into one of C consumer classes. Unfortunately, the exact number of individuals of each class may be unknown: consumers do not always report the appropriate class, due to various factors including differential energy rates for different consumer classes. We develop a methodology to estimate the expected energy use of each class as a function of time and the true number of consumers in each class. We also provide some measure of uncertainty of the resulting estimates. To accomplish this, we assume that the expected consumption is a function of time that can be well approximated by a linear combination of B-splines. Individual consumer perturbations from this baseline are modeled as B-splines with random coefficients. We treat the reported numbers of consumers in each category as random variables with distribution depending on the true number of consumers in each class and on the probabilities of a consumer in one class reporting as another class. We obtain maximum likelihood estimates of all parameters via a maximization algorithm. We introduce a special numerical trick for calculating the maximum likelihood estimates of the true number of consumers in each class. We apply our method to a data set and study our method via simulation.
研究动机与目标
- 当报告数字因差异化的能源费率而存在误报时,估计真实的消费者类别数量。
- 使用B样条基展开,将类别特定的能耗模式建模为时间的平滑函数。
- 通过样条系数的随机效应,考虑变压器内部的相关性与测量变异性。
- 开发一种最大似然估计框架,联合估计能耗曲线、真实数量和报告误差概率。
- 在不同数据结构下评估该方法的性能,包括变压器数量和重复观测次数。
提出的方法
- 将每类的期望能耗建模为B样条基函数的线性组合,以确保时间上的平滑性。
- 将个体消费者相对于均值曲线的偏离建模为具有随机系数的B样条,以捕捉类内变异性。
- 将报告的消费者数量视为随机变量,其分布类似多项分布,依赖于真实数量和误报概率。
- 使用专门的数值算法(定理1)高效计算真实消费者数量的最大似然估计。
- 采用轮廓似然方法估计方差分量和消费者层面的随机效应。
- 通过模拟研究和巴西变压器的真实数据验证模型在各种报告误差和数据可得性条件下的表现。
实验结果
研究问题
- RQ1当由于经济激励导致报告数字系统性误报时,如何估计每个类别中真实消费者数量?
- RQ2在使用B样条平滑的功能数据分析中,从汇总的变压器层级数据中,能多大程度准确恢复类别特定的能耗模式?
- RQ3在消费者层面偏差中引入随机效应,如何改善模型拟合度和不确定性量化?
- RQ4变压器数量和观测天数(重复次数)对估计的消费者数量和能耗曲线的准确性和精确度有何影响?
- RQ5尽管存在报告误差,该提出的基于似然的方法能否有效降低估计消费者类型时的偏差?
主要发现
- 该方法通过将估计值从报告数字向真实值调整,成功估计了真实消费者数量,尤其在使用多个变压器时效果更明显。
- 当每个变压器仅有一天空数据时,消费者层面方差分量的估计值常为零,表明此类估计的信息有限。
- 增加变压器数量可降低消费者数量估计的偏差,但增加重复观测次数并未显著减少偏差。
- 随着重复次数增加,估计消费者数量的变异性降低,但增加变压器数量并未显著降低变异性,表明重复次数在提高精度方面比减少偏差更有效。
- 模型对真实数据提供了合理的拟合,表现为观测到的汇总曲线与估计的类别特定均值曲线加权和之间高度一致。
- 在模拟中,真实消费者数量的最常见估计值通常介于报告数字与真实值之间,表明估计值向真实值进行了校正。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。