基本信息
- 来源: arxiv
- 原始来源: http://arxiv.org/abs/2609.20779v1
- 发布域名: arxiv.org
- 分类: cs.CL
- 作者: Sarah Wyer、Sue Black、Noura Al Moubayed
要点解读
这是什么
该研究指出,现有的安全评估依赖表层毒性指标,而模型实际上是将明显的歧视性内容转化为更隐蔽的形式,而非真正消除,这种现象被称为“危害洗白”。通过对 GPT 系列模型生成的性别导向文本进行大规模分析,揭示了不同代际模型在性别偏见表现上的转变。
用在哪里
适用于 AI 安全研发团队、模型审计机构以及制定公平性监管政策的部门,帮助他们认识到仅凭表面毒性分数无法全面判断模型是否真正降低了危害。
可以推断的
推测:表面毒性分数的下降可能掩盖了模型在更深层次上产生的性别偏见,需结合更细粒度的表征危害指标进行评估。
推测:随着模型代际更新,开发者可能无意中将显性歧视内容转化为更隐蔽的表达,使得传统检测工具难以捕捉真实风险。
来源摘要/节选
Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \emph{harm laundering}. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions do not. The pattern is most visible at GPT-5: Topic
5 (1,997documents) frames breast cancer as a men’s rights debate, while zero equivalent clusters appear in women-directed output. Three independent classifiers score this content as non-toxic. Sentiment scores invert at GPT-4: early models demean women; later models over-correct. Topic diversity in women-directed completions falls 36% relative to men at the GPT-4 alignment boundary (W/M~$= 0.58$, from $0.91$ at GPT-2). REGARD representational harm disparity correlates with release date ($ρ= +0.55$, $p = .034$) while Detoxify does not ($ρ= -0.23$, $p = .42$): toxicity scores fall as representational harm grows. We formalise harm laundering as a three-criteria test and provide a three-stage detection protocol applicable to any generative model. Within the OpenAI GPT lineage, toxicity score reduction is not a sufficient proxy for harm reduction.
来源说明
当前保存的是来源摘要,不代表论文全文。请以原始来源为准。
「要点解读」由 AI Stack 依据上方已保存内容整理,不代表来源的完整表述;标注「推测:」的判断来自编辑,不是来源陈述。