大语言模型中的性别与性别偏见:新尺度下的老问题
Sex and gender bias in large language models: an old problem at a new scale
越来越多患者向大语言模型咨询健康信息,因此系统输出在不同人群间是否一致成为公共卫生问题。人口学偏见比虚构信息更难察觉,因为它不体现在准确性指标上,而是反映在模型生成的语言中。性别偏见是最常被报告的形式,并常与族裔、社会经济地位等因素叠加,影响临床记录和推理。作者主张开展按性别分层的评估、在基准设计中纳入性别医学专业知识,并在部署后进行分层监测。文章强调所需方法已存在,问题在于是否将性别公平视为设计前提。
为什么推荐给您:聚焦大语言模型性别偏见这一新尺度下的伦理问题,提出可操作的评估与监测建议。
不需要生物学背景,多打比方
正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…
已等待 0 秒
这篇还没有动画
动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。
摘要Abstract
Patients are increasingly turning to large language models for health information, which makes the consistency of these systems' outputs across patient groups a public health concern. Demographic bias is harder to see than fabrication because it escapes accuracy benchmarks and shows up instead in the language the models produce. Gender bias is the form most consistently reported, with effects on clinical documentation and reasoning that often compound with ethnicity, socioeconomic status and other attributes. Drawing on evidence from gender medicine, we argue that it calls for sex- and gender-disaggregated evaluation, gender-medicine expertise in benchmark design and stratified monitoring after deployment. The methods needed for this already exist. What is still unsettled is whether gender equity will be treated as a design requirement or left until its consequences can no longer be ignored.