The attention mechanism is considered the backbone of the widely-used Transformer architecture. It contextualizes the input by computing input-specific attention matrices. We find that this mechanism, while powerful and elegant, is not as important as typically thought for pretrained language models. We introduce PAPA, a new probing method that replaces the input-dependent attention matrices with constant ones -- the average attention weights over multiple inputs. We use PAPA to analyze several established pretrained Transformers on six downstream tasks. We find that without any input-dependent attention, all models achieve competitive performance -- an average relative drop of only 8% from the probing baseline. Further, little or no performance drop is observed when replacing half of the input-dependent attention matrices with constant (input-independent) ones. Interestingly, we show that better-performing models lose more from applying our method than weaker models, suggesting that the utilization of the input-dependent attention mechanism might be a factor in their success. Our results motivate research on simpler alternatives to input-dependent attention, as well as on methods for better utilization of this mechanism in the Transformer architecture.

该研究介绍了一种新的探测方法 PAPA，它通过使用常量作为注意力权重值，取代了输入相关的注意力矩阵。该研究表明，当使用PAPA时，预训练Transformer模型在6个下游任务上仍然能够保持不错的性能表现，说明模型中的注意力机制并非如人们通常认为的那样重要。因此，该研究为探索更为简单的替代输入相关的注意力机制以及更好地利用这一机制提供了新的研究思路。

关注机制的实际作用是多少？质疑预训练Transformers模型中关注机制的重要性