Contextual word embeddings obtained from pre-trained language model (PLM) have proven effective for various natural language processing tasks at the word level. However, interpreting the hidden aspects within embeddings, such as syntax and semantics, remains challenging. Disentangled representation learning has emerged as a promising approach, which separates specific aspects into distinct embeddings. Furthermore, different linguistic knowledge is believed to be stored in different layers of PLM. This paper aims to disentangle semantic sense from BERT by applying a binary mask to middle outputs across the layers, without updating pre-trained parameters. The disentangled embeddings are evaluated through binary classification to determine if the target word in two different sentences has the same meaning. Experiments with cased BERT$_{\texttt{base}}$ show that leveraging layer-wise information is effective and disentangling semantic sense further improve performance.

该论文使用二进制掩码对预训练模型中不同层的输出进行切割，以解离BERT中的语义意义，而不更新预训练参数，从而产生解离的嵌入表示。使用二进制分类验证解离的嵌入的效果，判断两个不同句子中目标词的含义是否相同。实验结果表明，利用层次信息是有效的，而解离的语义意义进一步提高了性能。

通过逐层维度选择从预训练语言模型中解析单词语义