Effectively aligning Large Language Models (LLMs) with human-centric values
while preventing the degradation of abilities acquired through Pre-training and
Supervised Fine-tuning (SFT) poses a central challenge in Reinforcement
Learning from Human Feedback (RLHF). In this paper, we first discover that
interpolating RLHF and SFT model parameters can adjust the trade-off between
human preference and basic capabilities, thereby reducing the alignment tax at
the cost of alignment reward. Inspired by this, we propose integrating the RL
policy and SFT models at each optimization step in RLHF to continuously
regulate the training direction, introducing the Online Merging Optimizer.
Specifically, we merge gradients with the parameter differences between SFT and
pretrained models, effectively steering the gradient towards maximizing rewards
in the direction of SFT optimization. We demonstrate that our optimizer works
well with different LLM families, such as Qwen and LLaMA, across various model
sizes ranging from 1.8B to 8B, various RLHF algorithms like DPO and KTO, and
existing model merging methods. It significantly enhances alignment reward
while mitigating alignment tax, achieving higher overall performance across 14
benchmarks.

通过在线合并优化器，在人类反馈强化学习中持续调节训练方向，实现大语言模型的高性能表现和对齐奖励的显著提升，同时减小对齐成本。

在线合并优化器用于提升回报和降低税额的对齐

Online Merging Optimizers for Boosting Rewards and Mitigating Tax in  Alignment

Supervised fine-tuning (SFT) on instruction-following corpus is a crucial
approach toward the alignment of large language models (LLMs). However, the
performance of LLMs on standard knowledge and reasoning benchmarks tends to
suffer from deterioration at the latter stage of the SFT process, echoing the
phenomenon of alignment tax. Through our pilot study, we put a hypothesis that
the data biases are probably one cause behind the phenomenon. To address the
issue, we introduce a simple disperse-then-merge framework. To be concrete, we
disperse the instruction-following data into portions and train multiple
sub-models using different data portions. Then we merge multiple models into a
single one via model merging techniques. Despite its simplicity, our framework
outperforms various sophisticated methods such as data curation and training
regularization on a series of standard knowledge and reasoning benchmarks.

通过我们的研究，我们提出一个假设：数据偏差可能是大型语言模型在细调过程的后期出现性能下降的原因之一。为了解决这个问题，我们引入了一个简单的分散然后合并的框架。尽管简单，我们的框架在一系列标准的知识和推理基准测试中优于各种复杂的方法。

分散 - 合并：通过减少对齐税来推动指令调优的极限

Disperse-Then-Merge: Pushing the Limits of Instruction Tuning via  Alignment Tax Reduction

Finetuning language models with reinforcement learning (RL), e.g. from human
feedback (HF), is a prominent method for alignment. But optimizing against a
reward model can improve on reward while degrading performance in other areas,
a phenomenon known as reward hacking, alignment tax, or language drift. First,
we argue that commonly-used test metrics are insufficient and instead measure
how different algorithms tradeoff between reward and drift. The standard method
modified the reward with a Kullback-Lieber (KL) penalty between the online and
initial model. We propose Elastic Reset, a new algorithm that achieves higher
reward with less drift without explicitly modifying the training objective. We
periodically reset the online model to an exponentially moving average (EMA) of
itself, then reset the EMA model to the initial model. Through the use of an
EMA, our model recovers quickly after resets and achieves higher reward with
less drift in the same number of steps. We demonstrate that fine-tuning
language models with Elastic Reset leads to state-of-the-art performance on a
small scale pivot-translation benchmark, outperforms all baselines in a
medium-scale RLHF-like IMDB mock sentiment task and leads to a more performant
and more aligned technical QA chatbot with LLaMA-7B. Code available at
github.com/mnoukhov/elastic-reset.

使用弹性复位算法对语言模型进行微调，以在获得更高奖励的同时减少语言漂移，达到最佳性能。