Skip to main content
QUICK REVIEW

[Paper Review] Dynamic Data-Driven Estimation of Non-Parametric Choice Models

Nam Ho-Nguyen, Fatma Kılınç-Karzan|arXiv (Cornell University)|Feb 19, 2017
Economic and Environmental Valuation29 references3 citations
TL;DR

This paper proposes a dynamic, convex optimization-based framework for non-parametric choice model estimation using a primal-dual online learning approach, enabling provable convergence and explicit sparsity guarantees. It introduces a non-smooth dual mirror descent algorithm that outperforms existing methods by achieving sparser, more accurate models faster in both static and dynamic settings.

ABSTRACT

We study non-parametric estimation of choice models, which were introduced to alleviate unreasonable assumptions in traditional parametric models, and are prevalent in several application areas. Existing literature focuses only on the static observational setting where all of the observations are given upfront, they are not equipped with explicit convergence rate guarantees, and consequently they cannot provide an a priori analysis for the model accuracy vs sparsity trade-off on the actual estimated model returned by their algorithms. As opposed to this, we focus on estimating a non-parametric choice model from observational data in a \emph{dynamic} setting, where observations are obtained over time. We show that choice model estimation can be cast as a convex-concave saddle-point (SP) joint estimation and optimization (JEO) problem, and we provide a primal-dual framework for deriving algorithms to solve this based on online convex optimization. By tailoring our framework carefully to the choice model estimation problem, we obtain tractable algorithms with provable convergence guarantees and explicit bounds on the sparsity of the estimated model. Our numerical experiments confirm the effectiveness of the algorithms derived from our framework.

Motivation & Objective

  • Address the lack of convergence rate guarantees and a priori sparsity-accuracy trade-off analysis in static, non-parametric choice model estimation.
  • Develop a dynamic estimation framework where data arrive sequentially, enabling real-time model updates and adaptability.
  • Provide a unified convex-concave saddle-point formulation to model joint estimation and optimization in choice modeling.
  • Derive first-order algorithms with explicit convergence and sparsity bounds, overcoming limitations of prior column generation and Frank-Wolfe methods.
  • Enable a priori analysis of model accuracy versus sparsity by providing theoretical bounds on the number of rankings (support size) in the estimated model.

Proposed method

  • Formulate non-parametric choice model estimation as a convex-concave saddle-point (SP) problem, embedding both model estimation and optimization.
  • Apply online convex optimization (OCO) principles to derive a primal-dual algorithm that updates the model incrementally as new data arrive.
  • Design a non-smooth dual mirror descent (MD) algorithm with constant step size, leveraging known problem-specific constants (ΩY, G) for convergence guarantees.
  • Introduce smoothing techniques for the dual MD method to handle non-smooth norms, enabling comparison with smoothed and Frank-Wolfe variants.
  • Use a subproblem oracle to identify the most violated constraint (i.e., the most informative ranking) at each iteration, ensuring model sparsity.
  • Establish theoretical bounds on convergence rate and sparsity, where the number of iterations directly controls the model’s support size (sparsity).

Experimental results

Research questions

  • RQ1Can a dynamic, online algorithm for non-parametric choice model estimation achieve provable convergence and sparsity guarantees?
  • RQ2How does the performance of the proposed non-smooth dual mirror descent algorithm compare to Frank-Wolfe and smoothed mirror descent in terms of convergence speed, sparsity, and accuracy?
  • RQ3To what extent can the number of iterations serve as a proxy for model sparsity in non-parametric choice models?
  • RQ4Can the proposed framework enable a priori analysis of the trade-off between model accuracy and sparsity in the estimated model?
  • RQ5Does the dynamic setting allow for more efficient and scalable estimation compared to static, batch methods?

Key findings

  • The non-smooth dual mirror descent (MD_ns) algorithm achieves the sparsest models with the fewest iterations and lowest solution time across all tested configurations.
  • The average number of iterations strongly correlates with model sparsity (Spearman correlation ≈ 0.922), validating iteration count as a reliable proxy for sparsity.
  • The MD_ns method maintains comparable test mean absolute error (MAE) to other methods while achieving significantly better sparsity and faster convergence.
  • Increasing the number of training subsets (m) reduces test MAE but increases model sparsity and iteration count, confirming the expected trade-off between data richness and model complexity.
  • The non-smooth MD algorithm outperforms both Frank-Wolfe naïve and smoothed MD methods in overall solution time and sparsity, despite requiring more parameter tuning.
  • Theoretical bounds on convergence and sparsity are explicitly tied to the number of iterations, enabling a priori control over model complexity and accuracy.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.