Skip to main content
QUICK REVIEW

[论文解读] Sample-Based Bounds for Coherent Risk Measures: Applications to Policy Synthesis and Verification

Prithvi Akella, Anushri Dixit|arXiv (Cornell University)|Apr 21, 2022
Health Systems, Economic Evaluations, Quality of Life被引用 4
一句话总结

本文提出了一种基于样本的 $g$-微分风险度量的边界方法——这是一类一致风险度量——以实现在不确定机器人系统中高置信度的风险感知验证与策略综合。通过利用集中不等式和基于样本的优化,该方法识别出在99%置信度下优于99%替代策略的策略,已在多智能体系统仿真中验证,其中风险感知控制器的表现优于基线方法。

ABSTRACT

The dramatic increase of autonomous systems subject to variable environments has given rise to the pressing need to consider risk in both the synthesis and verification of policies for these systems. This paper aims to address a few problems regarding risk-aware verification and policy synthesis, by first developing a sample-based method to bound the risk measure evaluation of a random variable whose distribution is unknown. These bounds permit us to generate high-confidence verification statements for a large class of robotic systems. Second, we develop a sample-based method to determine solutions to non-convex optimization problems that outperform a large fraction of the decision space of possible solutions. Both sample-based approaches then permit us to rapidly synthesize risk-aware policies that are guaranteed to achieve a minimum level of system performance. To showcase our approach in simulation, we verify a cooperative multi-agent system and develop a risk-aware controller that outperforms the system's baseline controller. We also mention how our approach can be extended to account for any $g$-entropic risk measure - the subset of coherent risk measures on which we focus.

研究动机与目标

  • 在底层分布未知的情况下,开发 $g$-微分风险度量的基于样本的边界,以实现风险感知验证与策略综合。
  • 建立生成不确定系统中风险度量高置信度上界所需的基本样本量要求。
  • 将风险感知验证重新表述为使用基于样本边界的度量识别问题。
  • 开发一种基于样本的优化方法,以识别在高置信度下优于决策空间中大部分方案的策略。
  • 在合作式多智能体系统上验证该方法,展示其在风险感知方面优于基线控制器的性能。

提出的方法

  • 利用独立同分布样本的经验证据风险估计,推导 $g$-微分风险度量的集中不等式。
  • 应用霍夫丁型不等式,以高概率界定风险度量评估的上尾部分。
  • 将风险感知验证重新表述为估计由系统轨迹导出的随机变量风险度量的问题。
  • 使用基于样本的优化,基于 $N \geq 459$ 个采样参数,识别出在99%置信度下位于性能第99百分位的策略。
  • 采用鲁棒性度量 $\rho(\mathbf{x}^{\theta,p})$ 评估在不确定性下的系统性能,其中 $p$ 参数化控制器。
  • 将基于样本的风险边界作为非凸优化问题中的目标函数,用于策略综合。

实验结果

研究问题

  • RQ1当分布未知时,生成 $g$-微分风险度量的高置信度上界,所需的最小样本数是多少?
  • RQ2基于样本的边界能否用于生成风险感知系统性能的高置信度验证陈述?
  • RQ3基于样本的方法如何识别在风险感知优化中优于决策空间大部分方案的策略?
  • RQ4所提出的方法能否扩展至合成在风险感知意义上优于大多数替代方案的控制器?
  • RQ5与现有验证方法相比,基于样本的边界在紧致性和保守性方面表现如何?

主要发现

  • 使用 $N_{\mathcal{R}} = 149$ 个样本,该方法对 $\alpha = 0.1$ 时鲁棒性条件风险价值(CVaR)生成了95%置信度的上界 $-0.1489$,确保了高置信度的性能验证。
  • 识别出一个控制器参数集 $p_i$,其上界为 $\operatorname{\mathcal{R}}(p_i, 0.95, 0.1) = -0.1489$,证实其在99%置信度下位于性能的第99百分位。
  • 如图10所示,风险感知控制器在鲁棒性分布上优于基线 Robotarium 控制器,证实了其优越的风险感知性能。
  • 根据推论9,为在99%置信度下识别出位于第99百分位的策略,该方法要求 $N \geq 459$ 个样本。
  • 理论边界通过数值验证,确认了基于样本的集中不等式和优化框架的正确性。
  • 该方法可在无需事先了解底层不确定性分布的情况下,实现系统化风险感知验证与策略综合。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。