基本信息
- 来源: arxiv
- 原始来源: http://arxiv.org/abs/2608.12278v1
- 发布域名: arxiv.org
- 分类: cs.CL
- 作者: Avijit Roy、Proma Roy
要点解读
这是什么
文章分析了在资源不足语言社区中,AI 教育工具因底层基础设施的系统性缺陷而对使用者产生不利影响,包括网络内容稀缺、训练语料不平衡、分词导致的 token 费用以及网络接入差异,指出这些都是结构性障碍,并主张离线优先的设计是实现公平的途径。
用在哪里
适用于关注语言技术公平性、研究低资源语言处理、或在网络受限地区规划教育 AI 项目的政策制定者与技术研发人员。
可以推断的
推测:若要在类似环境取得成效,需要投入大量本地数据采集与离线模型适配工作。
推测:此类结构性问题的解决可能促使 AI 研发规范中加入对语言覆盖率的硬性要求。
来源摘要/节选
Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures, can systematically disadvantage speakers of underrepresented languages before a model is trained. This paper examines these structural barriers through Bengali, one of the world’s most widely spoken languages, focusing on AI-assisted education in low-connectivity environments. We identify four interlocking failures: a severe web presence gap, with Bengali accounting for less than 0.5% of global web content despite representing nearly 4% of the global population; a 67:1 training-token deficit between English and Bengali in major multilingual corpora; a tokenization penalty associated with Bengali’s alphasyllabary script that compounds the data deficit through higher token fertility; and connectivity exclusion, with individual internet penetration at 36.5% in rural areas compared with 71.4% in urban areas. These failures reflect longstanding resource-allocation decisions, institutional priorities, and design defaults that did not center underrepresented languages in mainstream AI development. We argue that dataset scarcity should be understood as a structural barrier rather than an isolated technical limitation, and that offline-first design should be treated as an equity-oriented infrastructure strategy. We conclude with directions for linguistics and AI research aimed at reducing these structural inequalities.
来源说明
当前保存的是来源摘要,不代表论文全文。请以原始来源为准。
「要点解读」由 AI Stack 依据上方已保存内容整理,不代表来源的完整表述;标注「推测:」的判断来自编辑,不是来源陈述。