[논문 리뷰] Comparing hundreds of machine learning classifiers and discrete choice models in predicting travel behavior: an empirical benchmark
이 연구는 모델 가족, 데이터셋, 표본 크기, 출력 유형의 4개 초차원을 포함하는 총 6,970개의 실험을 통해 105개의 기계학습(ML) 및 이산선택모형(DCM) 분류기들을 평가함으로써, 지금까지 가장 광범위한 실증 기준을 제공한다. 연구 결과, 앙상블 방법(예: 랜덤 포레스트, 부스팅)과 딥 네URAL 네트워크가 가장 높은 예측 정확도를 기록했으며, 랜덤 포레스트는 성능과 계산 효율성의 최적 균형을 제공함으로써 가장 우수한 균형을 이룬다. 반면 DCM은 다소 낮은 정확도를 보이나 대규모 적용에서는 계산적으로 매우 비효율적임을 확인했다.
Numerous studies have compared machine learning (ML) and discrete choice models (DCMs) in predicting travel demand. However, these studies often lack generalizability as they compare models deterministically without considering contextual variations. To address this limitation, our study develops an empirical benchmark by designing a tournament model, thus efficiently summarizing a large number of experiments, quantifying the randomness in model comparisons, and using formal statistical tests to differentiate between the model and contextual effects. This benchmark study compares two large-scale data sources: a database compiled from literature review summarizing 136 experiments from 35 studies, and our own experiment data, encompassing a total of 6,970 experiments from 105 models and 12 model families. This benchmark study yields two key findings. Firstly, many ML models, particularly the ensemble methods and deep learning, statistically outperform the DCM family (i.e., multinomial, nested, and mixed logit models). However, this study also highlights the crucial role of the contextual factors (i.e., data sources, inputs and choice categories), which can explain models' predictive performance more effectively than the differences in model types alone. Model performance varies significantly with data sources, improving with larger sample sizes and lower dimensional alternative sets. After controlling all the model and contextual factors, significant randomness still remains, implying inherent uncertainty in such model comparisons. Overall, we suggest that future researchers shift more focus from context-specific model comparisons towards examining model transferability across contexts and characterizing the inherent uncertainty in ML, thus creating more robust and generalizable next-generation travel demand models.
연구 동기 및 목표
- 이동 행동 예측에서 ML 및 DCM 분류기 간 비교를 위한 결정적인, 일반화 가능한 실증 기준을 수립하기 위해.
- 모델 성능이 데이터셋, 표본 크기, 출력 유형에 따라 어떻게 변하는지 조사하기 위해.
- 다양한 모델 가족 간 예측 정확도와 계산 비용 간 상충관계를 평가하기 위해.
- 향후 연구를 안내하기 위해 뛰어난 성능을 보이는 모델을 식별하고 DCM의 계산 효율성 향상 방안을 제안하기 위해.
- 공개 데이터셋과 표준화된 기준 프레임워크를 도입함으로써 방법론적 일관성을 촉진하기 위해.
제안 방법
- 연구는 12개의 모델 가족에서 온 105개의 분류기, 3개의 데이터셋(NHTS2017, LTDS2015, 및 추가 데이터셋), 3개의 표본 크기, 3개의 출력 유형(이진, 다항, 순서형 선택)을 포함하는 광범위한 실험 공간을 구성한다.
- 각 실험 포인트는 고정된 초차원을 가진 훈련된 모델을 의미하며, 총 6,970개의 고유한 실험을 생성한다.
- 예측 정확도는 표준 지표(예: 분류 정확도)를 사용해 측정하고, 계산 비용은 훈련 시간으로 기록한다.
- 결과의 타당성과 일반화 가능성을 확보하기 위해, 35개의 이전 연구에서 유래한 136개의 실험 포인트를 포함하는 메타데이터셋을 사용해 검증한다.
- 이 프레임워크는 향후 연구가 새로운 실험 포인트로 추가될 수 있도록 하여 지속적인 기준 설정과 지식 축적을 가능하게 한다.
실험 결과
연구 질문
- RQ1이동 행동 모델링에서 어떤 기계학습 및 이산선택모형 가족이 가장 높은 예측 정확도를 달성하는가?
- RQ2모델 성능은 다양한 데이터셋, 표본 크기, 출력 유형에 따라 어떻게 변화하는가?
- RQ3ML 및 DCM 분류기 간 예측 정확도와 계산 비용 간의 상충관계는 어떠한가?
- RQ4다양한 실험 조건에서 분류기 간 상대 순위는 얼마나 안정적인가?
- RQ5이산선택모형의 계산 효율성 향상은 어떤 방식으로 대규모 데이터 응용 분야에서의 실현 가능성을 높일 수 있는가?
주요 결과
- 앙상블 방법(랜덤 포레스트, 기울기 부스팅, 백킹 포함)은 평가된 모든 분류기 중에서 가장 높은 예측 정확도를 기록한다.
- 딥 네URAL 네트워크(DNNs) 역시 최상위 성능을 보이지만, 상당히 높은 계산 자원이 요구된다.
- 랜덤 포레스트는 예측 정확도와 계산 효율성의 최적 균형을 제공하므로 이상적인 기준 모델이다.
- 이산선택모형(DCMs)은 최고의 ML 모델보다 3~4%p 낮은 정확도를 보이나, 특히 대규모 데이터셋이나 고차원 입력에서 계산 속도가 매우 느리다.
- 분류기 간 상대 순위는 데이터셋과 조건에 관계없이 매우 안정적이지만, 절대 정확도와 계산 시간은 크게 다름을 확인했다.
- DCMs는 대규모 데이터 환경에서 심각한 계산 병목 현상을 겪고 있어, DCM 공동체가 모델 피팅보다 계산 효율성 향상을 우선순위로 삼을 필요가 있음이 시사된다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.