While Convolutional Neural Networks (CNNs) have been widely successful in 2D human pose estimation, Vision Transformers (ViTs) have emerged as a promising alternative to CNNs, boosting state-of-the-art performance. However, the quadratic computational complexity of ViTs has limited their applicability for processing high-resolution images and long videos. To address this challenge, we propose a simple method for reducing ViT's computational complexity based on selecting and processing a small number of most informative patches while disregarding others. We leverage a lightweight pose estimation network to guide the patch selection process, ensuring that the selected patches contain the most important information. Our experimental results on three widely used 2D pose estimation benchmarks, namely COCO, MPII and OCHuman, demonstrate the effectiveness of our proposed methods in significantly improving speed and reducing computational complexity with a slight drop in performance.

提出了一种用于减少Vision Transformers计算复杂度的简单方法，通过选择和处理最有信息的小片段，我们将二维人体姿态估计网络的结果作为指导进行小片段的选择，实验结果表明这种方法在显著提高速度和减少计算复杂度方面非常有效，而且性能略微下降。

通过补丁选择实现人体姿势估计的高效视觉变换器