Skip to main content
QUICK REVIEW

[논문 리뷰] Brain-like representational straightening of natural movies in robust feedforward neural networks

Tahereh Toosi, Elias B. Issa|arXiv (Cornell University)|2023. 08. 26.
Image Processing Techniques and Applications인용 수 4
한 줄 요약

이 논문은 자연 영화를 처리할 때, 적대적 훈련 또는 랜덤 스무딩을 통해 훈련된 강건한 피드포워드 신경망이 입력 노이즈에 대한 강건성 덕분에 뇌 유사 표현 직선화를 특성 공간에서 자연스럽게 발현하며, 이로 인해 선형 보간을 통해 현실적인 중간 프레임을 생성할 수 있음을 보여준다. 주요 기여는 이러한 직선화가 시간 동적 또는 예측 목표에 대한 명시적 훈련 없이도 강건성에 기인해 자연스럽게 발생한다는 점이다.

ABSTRACT

Representational straightening refers to a decrease in curvature of visual feature representations of a sequence of frames taken from natural movies. Prior work established straightening in neural representations of the primate primary visual cortex (V1) and perceptual straightening in human behavior as a hallmark of biological vision in contrast to artificial feedforward neural networks which did not demonstrate this phenomenon as they were not explicitly optimized to produce temporally predictable movie representations. Here, we show robustness to noise in the input image can produce representational straightening in feedforward neural networks. Both adversarial training (AT) and base classifiers for Random Smoothing (RS) induced remarkably straightened feature codes. Demonstrating their utility within the domain of natural movies, these codes could be inverted to generate intervening movie frames by linear interpolation in the feature space even though they were not trained on these trajectories. Demonstrating their biological utility, we found that AT and RS training improved predictions of neural data in primate V1 over baseline models providing a parsimonious, bio-plausible mechanism -- noise in the sensory input stages -- for generating representations in early visual cortex. Finally, we compared the geometric properties of frame representations in these networks to better understand how they produced representations that mimicked the straightening phenomenon from biology. Overall, this work elucidating emergent properties of robust neural networks demonstrates that it is not necessary to utilize predictive objectives or train directly on natural movie statistics to achieve models supporting straightened movie representations similar to human perception that also predict V1 neural responses.

연구 동기 및 목표

  • 강건한 신경망이 자연 영화를 처리할 때 특성 공간에서 뇌 유사 표현 직선화를 어떻게 발현하는지 조사하기.
  • 적대적 훈련 또는 랜덤 스무딩을 통한 입력 노이즈에 대한 강건성이 자연 영화 시퀀스의 선형적이고 역행 가능한 표현을 생성할 수 있는지 판단하기.
  • 기본 피드포워드 신경망과 비교하여 이러한 강건한 네트워크가 영양성 피라미드의 주요 시각 피질(V1)에서의 신경 반응을 얼마나 잘 예측하는지 평가하기.
  • 강건한 네트워크의 특성 표현 기하학적 성질과 생물학적 직선화 및 차원 증가 간의 관계 탐색하기.

제안 방법

  • 입력 노이즈 수준(σ² = 0.1, 0.5, 1.0)을 다양하게 설정하여 적대적 훈련(AT) 및 랜덤 스무딩(RS)을 사용해 ResNet-50 모델을 훈련시켰다.
  • 자연 영화 프레임의 시퀀스에 대해 훈련된 네트워크의 최종 완전 연결 계층에서 특성 표현을 추출했다.
  • 연속된 프레임이 특성 공간에서 형성하는 궤적의 곡률을 계산하여 표현 직선화 정도를 측정했다.
  • 시작 프레임과 끝 프레임 특성 간에 선형 보간을 수행하고, 보간된 특성을 다시 이미지 공간으로 복원하여 역행성 평가를 수행했다.
  • 표준 ResNet50 기준 모델과 비교하여 전개 점수(영화 표현의 반경 크기) 및 곡률과 같은 기하학적 성질을 평가했다.
  • 정확도 기반 지표를 사용해 강건 및 비강건 모델 간의 프라임레이트 V1 신경 변동 예측 성능을 비교했다.
Figure 1: Perceptual straightening of movie frames can be viewed as invertibility of latent representations for static images. Left: straightening of representations refers to a decrease in the curvature of the trajectory in representation space such as a neural population in the brain or human perc
Figure 1: Perceptual straightening of movie frames can be viewed as invertibility of latent representations for static images. Left: straightening of representations refers to a decrease in the curvature of the trajectory in representation space such as a neural population in the brain or human perc

실험 결과

연구 질문

  • RQ1정적 이미지에서 훈련된 강건한 피드포워드 신경망이 시간적 지도 없이 자연 영화 시퀀스에 대해 표현 직선화를 발현할 수 있는가?
  • RQ2적대적 훈련과 랜덤 스무딩이 자연 영화의 특성 공간에서 선형적이고 역행 가능한 표현을 얼마나 잘 유도하는가?
  • RQ3강건한 네트워크의 특성 표현 기하학적 성질(곡률, 전개)이 생물학적 시각 체계와 비교해 어떻게 다른가?
  • RQ4입력 노이즈에 대한 강건성이 기존 분류기와 비교해 프라임레이트 V1의 신경 반응 예측 능력을 향상시키는가?
  • RQ5예를 들어 RS에서 σ² = 0.5일 때 최적의 입력 노이즈 수준이 직선화와 표현 전개를 가장 잘 균형 잡는가?

주요 결과

  • 적대적 훈련 및 랜덤 스무딩을 통해 훈련된 강건한 네트워크는 자연 영화의 특성 궤적에서 곡률이 크게 감소하여 영양성 피라미드 V1과 인간 인지에서 관찰되는 수준의 직선화를 달성했다.
  • 강건한 네트워크의 특성 공간에서 시작 프레임과 끝 프레임 간의 선형 보간을 통해 현실적인 중간 프레임을 생성할 수 있었으며, 이는 선형적이고 시간 예측 가능한 표현을 갖는다는 것을 보여주었다.
  • σ² = 0.5인 RS 모델은 높은 직선화와 최소한의 표현 수축을 동시에 달성하여 다른 노이즈 수준 및 기준 모델보다 V1 신경 변동을 설명하는 데 가장 뛰어난 성능을 보였다.
  • 특히 σ² = 0.5인 RS 모델과 최상의 AT 모델은 표준 ResNet50보다 프라임레이트 V1의 신경 반응 예측 능력에서 뚜렷이 뛰어나 생물학적 타당성이 높다는 것을 시사했다.
  • 최상의 RS 모델에서 영화 표현의 전개 점수가 표준 ResNet50 대비 가장 높았으며, 이는 강건성이 망막에서 V1로의 촉각 처리 과정과 일치하는 차원 증가를 지원한다는 것을 시사한다.
  • 강건한 네트워크에서 직선화된 표현이 시간적 시퀀스나 예측 목표에 대한 훈련 없이도 자연스럽게 발생했으며, 이는 노이즈에 대한 강건성이 뇌 유사 영화 표현을 생성하는 데 충분한 메커니즘이라는 것을 의미한다.
Figure 2: ANNs show straightening of representations when robustness to noise constraints (noise augmentation or adversarial attack) is added to their training. Measurements for straightening of movie sequences (from (Hénaff et al., 2019 ) , in each layer of ResNet50 architecture under different tra
Figure 2: ANNs show straightening of representations when robustness to noise constraints (noise augmentation or adversarial attack) is added to their training. Measurements for straightening of movie sequences (from (Hénaff et al., 2019 ) , in each layer of ResNet50 architecture under different tra

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.