基本信息
- 来源: arxiv
- 原始来源: https://arxiv.org/abs/2602.21204v1
- 作者: Junchen Liu, Sven Elflein, Or Litany, Zan Gojcic, Ruilong Li
- 分类: cs.LG
- 论文时间: 2026-02-24T18:59:30Z
- 论文 PDF: https://arxiv.org/pdf/2602.21204v1.pdf
来源摘要/节选
Test-time training (TTT) with KV binding as sequence modeling layer is commonly interpreted as a form of online meta-learning that memorizes a key-value mapping at test time. However, our analysis reveals multiple phenomena that contradict this memorization-based interpretation. Motivated by these findings, we revisit the formulation of TTT and show that a broad class of TTT architectures can be expressed as a form of learned linear attention operator. Beyond explaining previously puzzling model behaviors, this perspective yields multiple practical benefits: it enables principled architectural simplifications, admits fully parallel formulations that preserve performance while improving efficiency, and provides a systematic reduction of diverse TTT variants to a standard linear attention form. Overall, our results reframe TTT not as test-time memorization, but as learned linear attention with enhanced representational capacity.
来源说明
当前只保存了官方论文摘要,不代表论文全文。请以原始来源为准。
本页只呈现已做哈希绑定的来源证据,不包含基于旧正文或缺失原文的扩展推断。