Skip to main content
QUICK REVIEW

[논문 리뷰] Why Tabular Foundation Models Should Be a Research Priority

Boris van Breugel, Mihaela van der Schaar|arXiv (Cornell University)|2024. 05. 02.
Geological Modeling and Analysis인용 수 4
한 줄 요약

이 논문은 표본 데이터를 위한 새로운 기초 모델인 대규모 표본 모델(Large Tabular Models, LTMs)에 대한 연구를 우선시해야 한다고 주장한다. 이는 다양한 분야에서 맥락 기반, 소수의 예제를 통한, 그리고 합성 데이터 생성을 가능하게 하여 데이터 과학을 혁신할 수 있다. 저자들은 LTMs를 과학적 발견, 개인정보 보호를 위한 데이터 공유, 강력하고 포용적인 기계 학습을 가능하게 하는 높은 영향력과 아직 탐색되지 않은 분야로 추천한다.

ABSTRACT

Recent text and image foundation models are incredibly impressive, and these models are attracting an ever-increasing portion of research resources. In this position piece we aim to shift the ML research community's priorities ever so slightly to a different modality: tabular data. Tabular data is the dominant modality in many fields, yet it is given hardly any research attention and significantly lags behind in terms of scale and power. We believe the time is now to start developing tabular foundation models, or what we coin a Large Tabular Model (LTM). LTMs could revolutionise the way science and ML use tabular data: not as single datasets that are analyzed in a vacuum, but contextualized with respect to related datasets. The potential impact is far-reaching: from few-shot tabular models to automating data science; from out-of-distribution synthetic data to empowering multidisciplinary scientific discovery. We intend to excite reflections on the modalities we study, and convince some researchers to study large tabular models.

연구 동기 및 목표

  • 실제 응용에서 지배적인 표본 데이터가 기초 모델 연구에서 다소 부족하게 다뤄지고 있음에도 불구하고, 이에 대한 부족한 대표성을 강조하기 위해.
  • 표본 기초 모델(LTMs)이 텍스트 및 시각 기초 모델과 유사한 높은 영향력과 아직 탐색되지 않은 분야를 제공하며, 이는 기초 모델 연구의 중요한 전망임을 주장하기 위해.
  • LTMs가 간과된 이유를 다루기 위해: 대규모 표본 메타데이터셋의 부족, 표본 기반 기계 학습의 본질적 어려움, 그리고 시각 및 텍스트를 선호하는 인간의 인지 편향.
  • LTMs가 소수의 예제를 통한 데이터 증강, 합성 데이터 생성, 그리고 다중 도메인 간 데이터 통합과 같은 변혁적 응용을 가능하게 할 수 있음을 제안하기 위해.
  • 연구자들이 LLM보다 접근성이 높고 공공의 복리에 기여할 잠재력이 큰 LTM에 주목하도록 장려하기 위해.

제안 방법

  • LLM과 시각 기초 모델(FMs)과 유사하게, 표본 데이터에 특화된 기초 모델인 대규모 표본 모델(LTMs)의 개념을 제안한다.
  • LTMs가 합성 표본 데이터 생성, 다양한 데이터셋 간 조인, 소수의 예제를 통한 부족한 컬럼 보완 등 기능을 수행할 수 있음을 프레임워크화한다.
  • 하류 작업에 LTM 임베딩을 사용하는 것을 제안하며, 이는 모델 미세조정이나 다른 기초 모델과의 통합에 활용될 수 있다.
  • ImageNet이 시각 기초 모델의 발전을 가능하게 했듯, 대규모, 다양하고 고도의 품질을 갖춘 표본 메타데이터셋이 LTMs를 훈련시키는 데 필수적임을 강조한다.
  • 다양한 도메인과 데이터 분포에서 일반화 능력, 편향, 신뢰성 등을 평가할 수 있는 평가 프로토콜의 필요성을 주장한다.
  • 표본 데이터의 추적 가능성과 신원 확인의 중요성을 강조하여, 생성된 이미지나 텍스트 모델에 비해 오용 위험을 줄일 수 있음을 설명한다.
Figure 1: Representation of different modalities in foundation model research across recent ML conferences, roughly estimated as the number of accepted papers with abstracts containing keywords (see Appendix A ). LLMs are booming and tabular data is heavily underrepresented.
Figure 1: Representation of different modalities in foundation model research across recent ML conferences, roughly estimated as the number of accepted papers with abstracts containing keywords (see Appendix A ). LLMs are booming and tabular data is heavily underrepresented.

실험 결과

연구 질문

  • RQ1과학 및 산업 분야에서 널리 사용되는 표본 데이터가 텍스트 및 시각 모델에 비해 기초 모델 연구에서 거의 다뤄지지 않은 이유는 무엇인가?
  • RQ2기초 모델 개발에 있어 표본 데이터는 어떤 고유한 과제를 안고 있으며, 이를 어떻게 해결할 수 있는가?
  • RQ3대규모 표본 모델(LTMs)은 어떻게 소수의 예제를 통한 데이터 증강, 합성 데이터 생성, 그리고 다중 데이터셋 간 추론을 가능하게 하는가?
  • RQ4LTMs는 데이터 민주화, 개인정보 보호, 과학적 발견, 모델의 강건성에 어떤 잠재적 영향을 미칠 수 있는가?
  • RQ5LTMs는 어떻게 신뢰성 있게 평가될 수 있으며, 편향과 오용을 방지하기 위해 어떤 보호 조치가 필요한가?

주요 결과

  • 표본 데이터는 과학, 헬스케어, 금융, 정부 분야에서 지배적인 데이터 모odal이며, 기초 모델 연구에서는 여전히 극히 부족하게 다뤄지고 있다.
  • XGBoost와 같은 강력한 기초 성능에도 불구하고, 표본 기반 기계 학습은 시각이나 NLP만큼 빠르게 발전하지 못했으며, 이는 기초 모델 분야에서 혁신의 여지가 크다는 것을 의미한다.
  • LTMs는 통계적 특성을 유지하면서도 고품질의 합성 표본 데이터를 생성할 수 있으며, 이는 개인정보 보호를 위한 데이터 공유를 가능하게 한다.
  • LTMs는 소수의 예제를 통한 힌트를 활용해, 예를 들어 부족한 컬럼을 추론하거나 다양한 도메인 간 데이터셋을 조인하는 등 다중 데이터셋 추론을 수행할 수 있다.
  • 도메인 전문 지식이 없이도 신뢰할 수 있는 표본 데이터를 위조하기 어려운 점을 감안할 때, LTMs의 오용 위험은 텍스트 및 시각 기초 모델보다 낮다.
  • LTM 연구는 LLM보다 계산 자원의 장벽이 낮아 더 넓은 참여와 더 빠른 진전이 가능하며, 높은 영향력과 접근성이 뛰어난 분야이다.
Figure 2: Sampling continuous distributions using LLMs autoregressively is inefficient . Assume we autoregressively sample tokens aiming to generate numbers that follow a standard Gaussian. What token probabilities should the LLM output at each sampling step? Let us consider a total vocabulary of ju
Figure 2: Sampling continuous distributions using LLMs autoregressively is inefficient . Assume we autoregressively sample tokens aiming to generate numbers that follow a standard Gaussian. What token probabilities should the LLM output at each sampling step? Let us consider a total vocabulary of ju

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.