基本信息
- 来源: arxiv
- 原始来源: https://arxiv.org/abs/2602.24266v1
- 作者: Amir Asiaee
- 分类: cs.LG
- 论文时间: 2026-02-27T18:35:10Z
- 论文 PDF: https://arxiv.org/pdf/2602.24266v1.pdf
来源摘要/节选
Neural networks are hypothesized to implement interpretable causal mechanisms, yet verifying this requires finding a causal abstraction – a simpler, high-level Structural Causal Model (SCM) faithful to the network under interventions. Discovering such abstractions is hard: it typically demands brute-force interchange interventions or retraining. We reframe the problem by viewing structured pruning as a search over approximate abstractions. Treating a trained network as a deterministic SCM, we derive an Interventional Risk objective whose second-order expansion yields closed-form criteria for replacing units with constants or folding them into neighbors. Under uniform curvature, our score reduces to activation variance, recovering variance-based pruning as a special case while clarifying when it fails. The resulting procedure efficiently extracts sparse, intervention-faithful abstractions from pretrained networks, which we validate via interchange interventions.
来源说明
当前只保存了官方论文摘要,不代表论文全文。请以原始来源为准。
本页只呈现已做哈希绑定的来源证据,不包含基于旧正文或缺失原文的扩展推断。