Skip to main content
QUICK REVIEW

[论文解读] A Structured Dictionary Perspective on Implicit Neural Representations

Gizem Yüce, Guillermo Ortiz-Jiménez|arXiv (Cornell University)|Dec 3, 2021
Neural Networks and Applications被引用 7
一句话总结

本文從結構化字典的視角出發,理論分析隱式神經表徵(INRs),揭示INRs的功能類似於信號字典,其原子為初始頻率的整數諧波,進而實現參數線性增長下頻率支持的指數級擴展。研究進一步表明,神經正切核(NTK)的特徵函數即為字典原子,而元學習透過從元訓練資料中學習更豐富的字典來提升性能,進而提高重構效率。

ABSTRACT

Implicit neural representations (INRs) have recently emerged as a promising alternative to classical discretized representations of signals. Nevertheless, despite their practical success, we still do not understand how INRs represent signals. We propose a novel unified perspective to theoretically analyse INRs. Leveraging results from harmonic analysis and deep learning theory, we show that most INR families are analogous to structured signal dictionaries whose atoms are integer harmonics of the set of initial mapping frequencies. This structure allows INRs to express signals with an exponentially increasing frequency support using a number of parameters that only grows linearly with depth. We also explore the inductive bias of INRs exploiting recent results about the empirical neural tangent kernel (NTK). Specifically, we show that the eigenfunctions of the NTK can be seen as dictionary atoms whose inner product with the target signal determines the final performance of their reconstruction. In this regard, we reveal that meta-learning has a reshaping effect on the NTK analogous to dictionary learning, building dictionary atoms as a combination of the examples seen during meta-training. Our results permit to design and tune novel INR architectures, but can also be of interest for the wider deep learning theory community.

研究动机与目标

  • 理解隱式神經表徵(INRs)成功與限制的理論機制。
  • 基於諧波分析與信號字典,統一多種INR架構於單一理論框架之下。
  • 利用經驗神經正切核(NTK)及其特徵函數,表徵INRs的偏好偏差。
  • 解釋元學習如何透過重塑NTK以形成更高效的信號字典,從而提升INR性能。

提出的方法

  • 利用諧波分析,顯示INR每一層將輸入分解為高階諧波,從而指數級擴展頻率支持。
  • 將INRs建模為結構化信號字典,其中原子對應於初始輸入頻率的整數諧波。
  • 透過初始化時經驗NTK的特徵函數分析INRs的偏好偏差,將其視為字典原子。
  • 利用NTK透過與特徵函數的內積來編碼目標信號的能力,量化學習難度。
  • 示範元學習作為字典學習演算法,透過結合元訓練樣本形成更高效的NTK原子。
  • 透過在CelebA上的實驗驗證結果,測量NTK特徵函數上的PSNR與不同初始化方式下的重構性能。

实验结果

研究问题

  • RQ1INRs的表達能力如何與其架構深度及頻率支持相關?
  • RQ2在多大程度上可透過其經驗NTK的特徵函數來表徵INRs的偏好偏差?
  • RQ3為何元學習能提升INR的收斂與泛化能力?它如何改變NTK結構?
  • RQ4INRs的失敗模式(如不完全重構與混疊)的成因為何?它們與字典表徵有何關聯?
  • RQ5在單一影像上微調預訓練模型相比元學習,性能為何下降?在此過程中NTK扮演何種角色?

主要发现

  • INRs隨著深度增加,其頻率支持呈指數級增長,僅需線性增加參數即可高效表徵寬頻信號。
  • 初始化時經驗NTK的特徵函數即為字典原子,這些原子與目標信號的內積決定了重構性能。
  • 元學習透過結合元訓練樣本,將NTK重塑為更豐富的信號字典,顯著提升編碼效率與泛化能力。
  • 在單一影像上預訓練會導致NTK特徵函數過度擬合該影像,進而導致壓縮性能更差且收斂速度慢於隨機初始化。
  • 重構性能與NTK特徵值譜強烈相關:與高特徵值NTK特徵函數對齊的信號更容易被學習。
  • 從預訓練權重微調通常導致負向遷移,原因在於NTK字典與目標信號分佈對齊不佳。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。