來源:arXiv · cs.CL查看原文 ↗
原文著作權歸來源方所有,本站僅作收錄、翻譯或格式整理。
事實脈絡
arXiv:2608.11426v1 Announce Type: new Abstract: The lack of diversity in LM content is widely attributed to the alignment process, but how and where exactly in the pipeline this collapse begins is unknown. We argue that
解讀與影響
导读摘要
正文
大语言模型生成内容缺乏多样性,通常被归咎于后期的对齐过程(如 RLHF)。但一篇新发布的 arXiv 预印本提出了一个更根本的问题:这种“输出趋同”究竟是从哪一步开始的?研究者认为,答案可能隐藏在对齐之前的基础模型阶段 arXiv:2608.11426。
论文的核心思路是沿着模型训练管线向上追溯,检验输出同质性是否在对齐之前就已经在基础模型中显现。如果基础模型本身对不同提示的响应已经高度相似,那么对齐过程只是放大了一个早已存在的倾向,而非问题的起点。摘要中作者明确表示,现有文献将多样性缺失广泛归因于对齐,但“这一坍缩在管线中究竟如何、在何处开始,仍是未知的”——这正是该研究试图填补的空白 arXiv:2608.11426。
这一视角与同期发布的另一项关于“群体对齐”的研究形成了有意义的对照。后者关注的是将语言模型适配到特定人口群体后可能产生的“谄媚”倾向,即模型为了迎合群体偏好而牺牲真实表达 arXiv:2608.11528。如果基础模型本身已经存在输出趋同的倾向,那么无论是标准对齐还是群体对齐,都可能在处理一个“先天”条件而非“后天”问题。
从证据强度来看,该论文目前以 arXiv 预印本形式发布,尚未经过同行评议,其具体实验设计、样本规模和测量指标在摘要中均未提供。因此,关于“趋同不可避免”的结论目前只能视为一种初步的研究假设,而非经过验证的定论。读者在引用时应留意这一边界。
參考來源
來源原文
arXiv:2608.11426v1 Announce Type: new Abstract: The lack of diversity in LM content is widely attributed to the alignment process, but how and where exactly in the pipeline this collapse begins is unknown. We argue that output homogeneity is likely learned during the pretraining phase, and only or magnified during the alignment process. Specifically, we find that semantic convergence is observed from the first alignment stage--the instruction-tuning phase (SFT)--suggesting that homogeneity might already exist in the pre-alignment model. To investigate this, we conduct controlled SFT experiments examining how training data influences output convergence on specific input/output pairs. We find that convergence can be revealed and amplified, but not introduced by the SFT data, supporting its role as a catalyst rather than a cause. To further test whether homogeneity originates before alignment, we measure convergence in base models. We find that instruct-like collapse can be induced through prompting alone, even without alignment. Taken together, our results suggest that semantic convergence may arise naturally from the objectives underlying LM training, making it difficult to mitigate through post-alignment interventions alone.