[Paper Review] Evaluating the Determinants of Mode Choice Using Statistical and Machine Learning Techniques in the Indian Megacity of Bengaluru
This study evaluates mode choice behavior in Bengaluru using statistical and machine learning models on a 1,350-household dataset, comparing multinomial logit with random forests, XGBoost, and SVM. The random forest model achieved the highest accuracy (60.5% on test data), and interpretability techniques revealed that a 10% increase in travel cost reduces bus usage probability by 0.34–0.66%, while a 10% time reduction increases metro preference by 0.16–0.42%.
The decision making involved behind the mode choice is critical for transportation planning. While statistical learning techniques like discrete choice models have been used traditionally, machine learning (ML) models have gained traction recently among the transportation planners due to their higher predictive performance. However, the black box nature of ML models pose significant interpretability challenges, limiting their practical application in decision and policy making. This study utilised a dataset of $1350$ households belonging to low and low-middle income bracket in the city of Bengaluru to investigate mode choice decision making behaviour using Multinomial logit model and ML classifiers like decision trees, random forests, extreme gradient boosting and support vector machines. In terms of accuracy, random forest model performed the best ($0.788$ on training data and $0.605$ on testing data) compared to all the other models. This research has adopted modern interpretability techniques like feature importance and individual conditional expectation plots to explain the decision making behaviour using ML models. A higher travel costs significantly reduce the predicted probability of bus usage compared to other modes (a $0.66\%$ and $0.34\%$ reduction using Random Forests and XGBoost model for $10\%$ increase in travel cost). However, reducing travel time by $10\%$ increases the preference for the metro ($0.16\%$ in Random Forests and 0.42% in XGBoost). This research augments the ongoing research on mode choice analysis using machine learning techniques, which would help in improving the understanding of the performance of these models with real-world data in terms of both accuracy and interpretability.
Motivation & Objective
- To analyze the determinants of mode choice in Bengaluru’s low- and low-middle-income households.
- To compare the predictive performance of traditional statistical models (e.g., multinomial logit) with modern machine learning classifiers.
- To evaluate the interpretability of machine learning models using feature importance and individual conditional expectation (ICE) plots.
- To quantify the impact of travel cost and time on mode preference using interpretable ML techniques.
- To support evidence-based transportation policy by combining high accuracy with model transparency in real-world urban data.
Proposed method
- Collected a dataset of 1,350 households from low and low-middle-income groups in Bengaluru.
- Applied multinomial logit model as a benchmark statistical method for mode choice analysis.
- Trained and compared four machine learning classifiers: decision trees, random forests, XGBoost, and support vector machines.
- Used feature importance and individual conditional expectation (ICE) plots to interpret model decisions and assess variable impacts.
- Evaluated model performance using accuracy on both training and test data splits.
- Quantified marginal effects of travel cost and time changes on mode choice probabilities using trained ML models.
Experimental results
Research questions
- RQ1Which model—statistical or machine learning—performs best in predicting mode choice in Bengaluru’s urban context?
- RQ2How do changes in travel cost and time influence the predicted probability of choosing specific transport modes?
- RQ3To what extent can interpretability techniques like feature importance and ICE plots enhance the transparency of ML models in transportation planning?
- RQ4What are the relative impacts of socioeconomic and travel-time variables on mode choice decisions?
- RQ5How do the predictions of ML models compare to traditional discrete choice models in terms of accuracy and interpretability?
Key findings
- The random forest model achieved the highest test accuracy of 60.5%, outperforming multinomial logit, decision trees, XGBoost, and SVM.
- A 10% increase in travel cost reduced the predicted probability of bus usage by 0.66% (random forest) and 0.34% (XGBoost).
- A 10% reduction in travel time increased the predicted probability of choosing the metro by 0.16% (random forest) and 0.42% (XGBoost).
- Feature importance and ICE plots revealed that travel cost and time were among the most influential predictors across all ML models.
- The study demonstrates that high-performing ML models like random forest and XGBoost can be made interpretable using modern explainability techniques.
- The integration of interpretability tools enables policymakers to understand and trust ML-based predictions for urban transport planning.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.