[论文解读] LSAP: Rethinking Inversion Fidelity, Perception and Editability in GAN Latent Space
LSAP 提出了一种新颖的潜在空间对齐反演范式,通过将反演的潜在码与生成器潜在空间的合成分布对齐,以提升 GAN 反演中的感知质量和可编辑性。它引入了归一化风格空间(S^N)和归一化风格空间余弦距离(NSCD)作为可微分度量与优化损失,实现了在基于编码器和基于优化的两种方法中,保真度、感知质量和可编辑性方面的最先进性能。
As research on image inversion advances, the process is generally divided into two stages. The first step is Image Embedding, involves using an encoder or optimization procedure to embed an image and obtain its corresponding latent code. The second stage, referred to as Result Refinement, further improves the inversion and editing outcomes. Although this refinement stage substantially enhances reconstruction fidelity, perception and editability remain largely unchanged and are highly dependent on the latent codes derived from the first stage. Therefore, a key challenge lies in obtaining latent codes that preserve reconstruction fidelity while simultaneously improving perception and editability. In this work, we first reveal that these two properties are closely related to the degree of alignment (or disalignment) between the inverted latent codes and the synthetic distribution. Based on this insight, we propose the extbf{ Latent Space Alignment Inversion Paradigm (LSAP)}, which integrates both an evaluation metric and a unified inversion solution. Specifically, we introduce the extbf{Normalized Style Space ($\mathcal{S^N}$ space)} and extbf{Normalized Style Space Cosine Distance (NSCD)} to quantify the disalignment of inversion methods. Moreover, our paradigm can be optimized for both encoder-based and optimization-based embeddings, providing a consistent alignment framework. Extensive experiments across various domains demonstrate that NSCD effectively captures perceptual and editable characteristics, and that our alignment paradigm achieves state-of-the-art performance in both stages of inversion.
研究动机与目标
- 解决 GAN 反演中感知质量和可编辑性受限于反演潜在码与合成潜在分布对齐程度的问题。
- 开发一种统一的、可微分的度量,以在潜在码层面定量反映感知质量和可编辑性,且独立于编辑向量。
- 设计一种可泛化的对齐解决方案,适用于基于编码器和基于优化的反演方法,同时不损害重建保真度。
- 证明提升潜在空间与合成分布的对齐程度可显著增强 GAN 反演中的图像质量和可编辑性。
提出的方法
- 引入归一化风格空间(S^N),作为衡量失配程度的更高效、更有效的潜在空间,相较于 Z、W 或 S 空间。
- 提出归一化风格空间余弦距离(NSCD),作为可微分的、在潜在码层面的度量,用于量化反演码与合成分布之间的失配程度。
- 基于 NSCD 构建一种对齐损失,该损失可在基于编码器的训练和基于优化的反演过程中进行优化。
- 在基于编码器和基于优化的反演框架中统一应用基于 NSCD 的对齐损失,以提升感知质量和可编辑性。
- 将 LSAP 范式集成到现有的两阶段方法(如 HFGI、SAM、PTI)中,以进一步提升性能。
- 使用 NSCD 作为代理度量,验证其值越低,图像质量与编辑结果越优。
实验结果
研究问题
- RQ1反演潜在码与合成分布的对齐程度在多大程度上影响 GAN 反演中的感知质量和可编辑性?
- RQ2能否开发一种统一的、可微分的度量,以在不同反演方法的潜在码层面评估感知质量和可编辑性?
- RQ3提升潜在空间对齐程度在多大程度上能增强保真度和可编辑性,同时不降低重建质量?
- RQ4所提出的对齐解决方案能否有效应用于基于编码器和基于优化的反演方法?
- RQ5NSCD 度量在实践中是否能可靠地反映感知质量和可编辑性?
主要发现
- LSAP 在包括人脸、汽车、教堂和野生动物在内的多个领域中,均实现了保真度和可编辑性的最先进性能。
- NSCD 度量与感知质量及可编辑性具有有效相关性,NSCD 值越低,图像质量越好,编辑结果越自然。
- 在基于优化的反演中,LSAP 显著提升了 W 和 W+ 空间中的可编辑性和图像质量,而原始投影方法会产生不自然的细节。
- LSAP E(基于编码器的变体)在重建保真度上与 pSp 相当,但在感知质量和可编辑性上显著优于 e4e。
- 当与 HFGI、SAM 和 PTI 等两阶段方法结合时,LSAP 达到了最先进性能,证明了其泛化能力与有效性。
- 视觉结果表明,LSAP 在保留身份特征和细节(如反光、眼镜)方面优于基线方法,尤其在编辑任务中表现更优。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。