Our ultimate goal is to build robust policies for robots that assist people. What makes this hard is that people can behave unexpectedly at test time, potentially interacting with the robot outside its training distribution and leading to failures. Even just measuring robustness is a challenge. Adversarial perturbations are the default, but they can paint the wrong picture: they can correspond to human motions that are unlikely to occur during natural interactions with people. A robot policy might fail under small adversarial perturbations but work under large natural perturbations. We propose that capturing robustness in these interactive settings requires constructing and analyzing the entire natural-adversarial frontier: the Pareto-frontier of human policies that are the best trade-offs between naturalness and low robot performance. We introduce RIGID, a method for constructing this frontier by training adversarial human policies that trade off between minimizing robot reward and acting human-like (as measured by a discriminator). On an Assistive Gym task, we use RIGID to analyze the performance of standard collaborative Reinforcement Learning, as well as the performance of existing methods meant to increase robustness. We also compare the frontier RIGID identifies with the failures identified in expert adversarial interaction, and with naturally-occurring failures during user interaction. Overall, we find evidence that RIGID can provide a meaningful measure of robustness predictive of deployment performance, and uncover failure cases in human-robot interaction that are difficult to find manually. https://ood-human.github.io.

构建机器人辅助人类的强大策略是我们的最终目标，而在测试时间，人类的行为可能出乎意料，并可能与机器人在其训练分布之外进行互动，导致失败。我们提出在这些交互环境中捕捉稳健性需要构建和分析整个自然-对抗前沿：人类策略的最佳权衡自然性和低机器人性能之间的帕累托前沿。我们引入了RIGID，一种通过训练对抗人类策略来构建这一前沿的方法，以在Assistive Gym任务中分析标准协作强化学习的性能，以及旨在提高稳健性的现有方法的性能。我们还将RIGID确定的前沿与专家对抗交互中确定的失败以及用户交互期间自然发生的失败进行比较。总的来说，我们发现RIGID能够提供有意义的稳健性测量，可以预测部署性能，并发现难以手动发现的人机交互故障案例。

通过自然-对抗边界量化辅助健壮性