Skip to main content
QUICK REVIEW

[論文レビュー] Squeeze, Recover and Relabel: Dataset Condensation at ImageNet Scale From A New Perspective

Zeyuan Yin, Eric P. Xing|arXiv (Cornell University)|Jun 22, 2023
Advanced Neural Network Applications被引用数 7
ひとこと要約

本稿では、大規模データセットの縮約のための新しい3段階フレームワーク、SRe2Lを提案する。このフレームワークは、モデル学習と合成データ最適化を分離することで、効率的かつ高解像度のImageNet-1Kのデータセット縮約を可能にする。Squeeze、Recover、Relabelの3段階による学習の分離により、バッチ正規化統計の一致と画素単位の再ラベル付けを実現し、1クラスあたり50枚の画像でImageNet-1Kで60.8%の検証精度を達成。これは先行手法よりも32.9%高い性能であり、学習時間を最大52倍短縮する。

ABSTRACT

We present a new dataset condensation framework termed Squeeze, Recover and Relabel (SRe$^2$L) that decouples the bilevel optimization of model and synthetic data during training, to handle varying scales of datasets, model architectures and image resolutions for efficient dataset condensation. The proposed method demonstrates flexibility across diverse dataset scales and exhibits multiple advantages in terms of arbitrary resolutions of synthesized images, low training cost and memory consumption with high-resolution synthesis, and the ability to scale up to arbitrary evaluation network architectures. Extensive experiments are conducted on Tiny-ImageNet and full ImageNet-1K datasets. Under 50 IPC, our approach achieves the highest 42.5% and 60.8% validation accuracy on Tiny-ImageNet and ImageNet-1K, outperforming all previous state-of-the-art methods by margins of 14.5% and 32.9%, respectively. Our approach also surpasses MTT in terms of speed by approximately 52$ imes$ (ConvNet-4) and 16$ imes$ (ResNet-18) faster with less memory consumption of 11.6$ imes$ and 6.4$ imes$ during data synthesis. Our code and condensed datasets of 50, 200 IPC with 4K recovery budget are available at https://github.com/VILA-Lab/SRe2L.

研究の動機と目的

  • 大規模データセット縮約において、特にフルImageNet解像度での計算的に非現実的な従来の二段階最適化手法の課題を解決すること。
  • モデルと合成データの最適化を統合的に最適化するのではなく、分離することで、メモリと計算コストを削減しながら高い性能を維持すること。
  • 多様なモデルアーキテクチャーやデータセット規模、特にフルImageNet-1Kを含む、スケーラブルで高解像度のデータセット縮約を実現すること。
  • 特に構造的およびクラス固有の詳細を保持する点で、従来手法と比較して画像の質と意味的忠実度を向上させること。
  • 継続的学習や事前学習済みモデルとの統合といった応用を通じて、実用的応用可能性を示すこと。

提案手法

  • 本手法は3段階のパラダイムを導入する:Squeeze(実データ上で事前学習し、グローバル統計を捉える)、Recover(BN統計の一致を用いてノイズから合成画像を生成する)、Relabel(クロップ単位の予測に基づき、正確なクラスラベルを割り当てる)。
  • 実データ入力と合成データの同時処理を回避することで、学習中に両方のデータを同時に処理する必要がなくなることにより、二段階最適化を分離する。
  • Recover段階では、実データのグローバルバッチ正規化統計を合成データに一致させることで、バッチ単位の特徴一致よりもより強固で包括的な分布一致を実現する。
  • Relabel段階では、クロップ単位の推論戦略を用いることで、合成画像への正確なラベル割り当てが可能になり、ラベルの一貫性と意味的正確性が向上する。
  • 任意のBN層を含む事前学習済みモデルと互換性があり、プラグアンドプレイで統合可能で、学習のオーバーヘッドを低減できる。
  • 二段階最適化におけるアンロールイテレーションを回避することで、メモリと計算コストを顕著に削減し、任意の入力解像度やモデルアーキテクチャに対応できる。
Figure 1 : Left is data synthesis time vs. accuracy on ImageNet-1K with 10 IPC (Images Per Class). Models include ConvNet-4, ResNet-{18, 50, 101}. † indicates ViT with 10M parameters [ 2 ] . Right is the comparison of widely-used bilevel optimization and our proposed decoupled training scheme.
Figure 1 : Left is data synthesis time vs. accuracy on ImageNet-1K with 10 IPC (Images Per Class). Models include ConvNet-4, ResNet-{18, 50, 101}. † indicates ViT with 10M parameters [ 2 ] . Right is the comparison of widely-used bilevel optimization and our proposed decoupled training scheme.

実験結果

リサーチクエスチョン

  • RQ1分離された学習戦略が、性能を損なわずに大規模データセット縮約における二段階最適化を代替できるか?
  • RQ2全データセットにわたるBN統計の一致が、バッチ単位の特徴一致よりも合成データ生成においてより良い一致を実現できるか?
  • RQ3提案手法が、計算コストを最小限に抑えながら、大規模(例:ImageNet-1K)かつ高解像度の画像合成を実現できるか?
  • RQ4継続的学習などの下流タスクにおいて、SRe2Lフレームワークは従来の縮約手法と比較してどの程度の性能を示すか?
  • RQ5本手法が、さまざまなモデルアーキテクチャーや画像解像度にどの程度一般化可能か?

主な発見

  • SRe2Lは、1クラスあたり50枚の画像でImageNet-1Kで60.8%の検証精度を達成し、従来のSOTAを32.9ポイントも上回った。
  • Tiny-ImageNetでは、1クラスあたり50枚の画像で42.5%の精度を達成し、従来手法を14.5ポイント上回った。
  • MTT(ConvNet-4)と比較して52倍速く、MTT(ResNet-18)と比較して16倍速く、それぞれ11.6倍および6.4倍の低いメモリ消費量を実現した。
  • SRe2Lによる合成画像は、MTTのぼやけた、色が支配的な出力と比較して、より明確な意味的特徴と輪郭を示す優れた視覚的品質を有している。
  • Tiny-ImageNetにおけるクラスインクリメンタル学習において、5ステップおよび10ステップの設定で、記憶容量を増加させてもSRe2Lは常にベースラインを上回った。
  • フレームワークにより、BN層を備えた市販の事前学習済みモデルを直接利用可能となり、学習のオーバーヘッドが低減され、実用的導入が促進された。
Figure 2 : Overview of our framework. It consists of three stages: in the first stage, a model is trained from scratch to accommodate most of the crucial information from the original dataset. In the second stage, a recovery process is performed to synthesize the target data from the Gaussian noise.
Figure 2 : Overview of our framework. It consists of three stages: in the first stage, a model is trained from scratch to accommodate most of the crucial information from the original dataset. In the second stage, a recovery process is performed to synthesize the target data from the Gaussian noise.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。