基本信息
- 来源: arxiv
- 原始来源: http://arxiv.org/abs/2608.23507v1
- 发布域名: arxiv.org
- 分类: cs.CL
- 作者: Xiang Chen、Zeyu Zhang
要点解读
这是什么
该内容介绍一种针对蒙古世界中历史人物姓名的跨语言、跨文字对齐评估基准,提供基于来源证据的成对身份验证任务和仅基于姓名的对照测试集。
用在哪里
适用于开发或评估能够处理古代文献、跨语言、跨文字记载的自然语言处理系统,尤其是需要区分同名不同人的研究项目。
可以推断的
推测:该基准可以帮助判断生成系统在缺乏表面名称信息时,是否能够利用历史上下文进行正确的人物身份消歧。
推测:在实际历史研究中,若需将不同文字来源的同名人物区分开来,可借助该基准检验模型利用证据的能力。
来源摘要/节选
Historical people may appear under different languages, scripts, and transcription traditions, while distinct individuals may share highly similar or even identical names. This makes historical identity reconciliation more than a problem of string matching or transliteration. We introduce MHER, a provenance-controlled benchmark for pairwise reconciliation of person-name attestations from the Mongol world. MHER contains a balanced 396-pair Name-only core over 84 primary historical persons and a stricter 160-pair Source-grounded subset constructed from mention-by-source evidence, with entity-disjoint development and test splits. Across five generative systems, correctly Source-grounded evidence improves paired TEST accuracy by 12.96 to 94.44 percentage points relative to Name-only input. On five identical-surface different-person cases, all models fail under names alone (0/25 model-item decisions), whereas Source-grounded evidence yields 24/25 correct resolutions, with the remaining output an abstention. Context-only ablations show that historical descriptions often carry substantial identity information, while explicitly signaled misgrounding controls produce substantially lower performance. We also find that names are not uniformly beneficial: for Qwen3-8B, restoring surface forms converts ten otherwise correct Context-only distinctions into false identity merges. These results show that historical entity reconciliation depends not only on surface correspondence, but on whether identity judgments respond appropriately to provenance-controlled historical evidence. MHER therefore provides a controlled framework for studying evidence use, abstention, and failure modes in historical NLP.
来源说明
当前保存的是来源摘要,不代表论文全文。请以原始来源为准。
「要点解读」由 AI Stack 依据上方已保存内容整理,不代表来源的完整表述;标注「推测:」的判断来自编辑,不是来源陈述。