基本信息

  • 来源: arxiv
  • 原始来源: http://arxiv.org/abs/2609.04190v1
  • 发布域名: arxiv.org
  • 分类: cs.CV
  • 作者: Adheesh Sunil Juvekar、Onkar Kishor Susladkar、Kiet A. Nguyen 等

要点解读

这是什么

EditVid 是一种无需训练的通用视频编辑框架,利用稀疏因果记忆保证局部连贯、对应驱动的后注意力令牌注入保持远距离身份,并采用软潜在混合实现编辑局部性。该框架同时支持文字指令和参考图像两种编辑方式。

用在哪里

适用于需要根据文字说明或参照图像对视频进行风格迁移、属性修改、对象插入、部件级编辑或主体替换等多样化编辑任务的场景。

可以推断的

推测:该框架不依赖大规模微调,适合资源受限或快速部署的环境。
推测:将多种编辑需求统一在同一模型中,可降低系统复杂度并简化实际产品的维护。

来源摘要/节选

Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits, including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On FiVE, EditVid achieves 78.16 FiVE-Acc, compared with 58.95 for the strongest evaluated training-free baseline, while obtaining competitive results on IVEBench. A user study further shows a 51.8% overall preference for EditVid over 7 competing methods.

来源说明

当前保存的是来源摘要,不代表论文全文。请以原始来源为准。

「要点解读」由 AI Stack 依据上方已保存内容整理,不代表来源的完整表述;标注「推测:」的判断来自编辑,不是来源陈述。