基本信息
- 来源: arxiv
- 原始来源: https://arxiv.org/abs/2602.23239v1
- 作者: Radha Sarma
- 分类: cs.AI
- 论文时间: 2026-02-26T17:16:17Z
- 论文 PDF: https://arxiv.org/pdf/2602.23239v1.pdf
来源摘要/节选
AI systems are increasingly deployed in high-stakes contexts – medical diagnosis, legal research, financial analysis – under the assumption they can be governed by norms. This paper demonstrates that assumption is formally invalid for optimization-based systems, specifically Large Language Models trained via Reinforcement Learning from Human Feedback (RLHF). We establish that genuine agency requires two necessary and jointly sufficient architectural conditions: the capacity to maintain certain boundaries as non-negotiable constraints rather than tradeable weights (Incommensurability), and a non-inferential mechanism capable of suspending processing when those boundaries are threatened (Apophatic Responsiveness). These conditions apply across all normative domains. RLHF-based systems are constitutively incompatible with both conditions. The operations that make optimization powerful – unifying all values on a scalar metric and always selecting the highest-scoring output – are precisely the operations that preclude normative governance. This incompatibility is not a correctable training bug awaiting a technical fix; it is a formal constraint inherent to what optimization is. Consequently, documented failure modes - sycophancy, hallucination, and unfaithful reasoning - are not accidents but structural manifestations. Misaligned deployment triggers a second-order risk we term the Convergence Crisis: when humans are forced to verify AI outputs under metric pressure, they degrade from genuine agents into criteria-checking optimizers, eliminating the only component in the system capable of normative accountability. Beyond the incompatibility proof, the paper’s primary positive contribution is a substrate-neutral architectural specification defining what any system – biological, artificial, or institutional – must satisfy to qualify as an agent rather than a sophisticated instrument.
来源说明
当前只保存了官方论文摘要,不代表论文全文。请以原始来源为准。
本页只呈现已做哈希绑定的来源证据,不包含基于旧正文或缺失原文的扩展推断。