Debias-SparseGPT:面向大语言模型的偏见感知剪枝方法

内容摘要
Debias-SparseGPT是一种针对大语言模型的偏见感知剪枝方法。该方法通过在具有代表性差异的输入上定义的二阶项,结合表征去偏,对模型进行后训练剪枝。实验表明,Debias-SparseGPT在多种生成型大语言模型上,相较于SparseGPT,能够持续减少剪枝引起的偏见,同时保持模型困惑度和零样本准确率。在结构化2:4稀疏度模式下,通过增加长上下文、内容丰富的示例来扩充校准集,可以进一步提高下游性能和公平性。总体而言,Debias-SparseGPT在保持稀疏模型计算效率的同时,提升了偏见与性能之间的权衡。
Debias-SparseGPT是一种针对大语言模型的偏见感知剪枝方法。该方法通过在具有代表性差异的输入上定义的二阶项,结合表征去偏,对模型进行后训练剪枝。实验表明,Debias-SparseGPT在多种生成型大语言模型上,相较于SparseGPT,能够持续减少剪枝引起的偏见,同时保持模型困惑度和零样本准确率。在结构化2:4稀疏度模式下,通过增加长上下文、内容丰富的示例来扩充校准集,可以进一步提高下游性能和公平性。总体而言,Debias-SparseGPT在保持稀疏模型计算效率的同时,提升了偏见与性能之间的权衡。

Model compression techniques such as pruning and quantization facilitate the efficient deployment and acceleration of Large Language Models (LLMs). However, recent studies show that weight sparsification methods, such as SparseGPT, can amplify existing biases in models, with outputs varying significantly depending on persona cues in the prompt. In this paper, we introduce Debias-SparseGPT, a post-training pruning method incorporating representational debiasing using a second-order term defined over demographically contrasting inputs. We perform empirical validation of our method over a wide range of generative LLMs. Across models and sparsity regimes (25%, 50%, and structured 2:4 sparsity), Debias-SparseGPT consistently reduces pruning-induced bias compared to SparseGPT while preserving model perplexity and zero-shot accuracy. Under the most restrictive 2:4 structured sparsity pattern, which most aggressively degrades model quality, augmenting the calibration set with long-context, content-rich examples further improves both downstream performance and fairness. Overall, Debias-SparseGPT advances the bias-performance trade-off while preserving the computational efficiency of sparse models.

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-02 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值