基本信息
- 来源: arxiv
- 原始来源: https://arxiv.org/abs/2602.23248v1
- 作者: Yegon Kim, Juho Lee
- 分类: cs.AI
- 论文时间: 2026-02-26T17:25:22Z
- 论文 PDF: https://arxiv.org/pdf/2602.23248v1.pdf
来源摘要/节选
As large language models become increasingly capable, it is critical that their outputs can be easily checked by less capable systems. Prover-verifier games can be used to improve checkability of model outputs, but display a degradation in accuracy compared to a baseline trained only to maximize correctness – a phenonemon named legibility tax. We propose a solution by decoupling the correctness from the checkability condition and instead training a “translator” model that turns a fixed solver model’s solution into a checkable form. This allows us to first train the solver to maximize correctness, and then train the translator to translate the solver into a checkable form while retaining the solver’s answer. To accommodate this new objective of translation, we formulate a decoupled prover-verifier game where the equilibria correspond to faithful and checkable translators.
来源说明
当前只保存了官方论文摘要,不代表论文全文。请以原始来源为准。
本页只呈现已做哈希绑定的来源证据,不包含基于旧正文或缺失原文的扩展推断。