Non-parametric, k-nearest-neighbor algorithms have recently made inroads to assist generative models such as language models and machine translation decoders. We explore whether such non-parametric models can improve machine translation models at the fine-tuning stage by incorporating statistics from the kNN predictions to inform the gradient updates for a baseline translation model. There are multiple methods which could be used to incorporate kNN statistics and we investigate gradient scaling by a gating mechanism, the kNN's ground truth probability, and reinforcement learning. For four standard in-domain machine translation datasets, compared with classic fine-tuning, we report consistent improvements of all of the three methods by as much as 1.45 BLEU and 1.28 BLEU for German-English and English-German translations respectively. Through qualitative analysis, we found particular improvements when it comes to translating grammatical relations or function words, which results in increased fluency of our model.

研究探究了在微调阶段引入kNN预测的统计数据来提高基线翻译模型，发现通过引入gating机制，kNN的真实概率和强化学习三种方法，相比于传统的微调，可以在四个标准机器翻译数据集上实现一致的改进，尤其于翻译语法关系或功能词时表现出更大的提升。

非参数最近邻辅助微调神经机器翻译