[Paper Review] Computational Optimal Transport
A comprehensive survey of optimal transport theory with a focus on computational methods, including Kantorovich relaxation, entropic regularization, and scalable algorithms for discrete and continuous measures.
Optimal transport (OT) theory can be informally described using the words of the French mathematician Gaspard Monge (1746-1818): A worker with a shovel in hand has to move a large pile of sand lying on a construction site. The goal of the worker is to erect with all that sand a target pile with a prescribed shape (for example, that of a giant sand castle). Naturally, the worker wishes to minimize her total effort, quantified for instance as the total distance or time spent carrying shovelfuls of sand. Mathematicians interested in OT cast that problem as that of comparing two probability distributions, two different piles of sand of the same volume. They consider all of the many possible ways to morph, transport or reshape the first pile into the second, and associate a "global" cost to every such transport, using the "local" consideration of how much it costs to move a grain of sand from one place to another. Recent years have witnessed the spread of OT in several fields, thanks to the emergence of approximate solvers that can scale to sizes and dimensions that are relevant to data sciences. Thanks to this newfound scalability, OT is being increasingly used to unlock various problems in imaging sciences (such as color or texture processing), computer vision and graphics (for shape manipulation) or machine learning (for regression, classification and density fitting). This short book reviews OT with a bias toward numerical methods and their applications in data sciences, and sheds lights on the theoretical properties of OT that make it particularly useful for some of these applications.
Motivation & Objective
- Explain the theoretical foundations of optimal transport and its linkage to geometry on probability spaces.
- Present computational frameworks and algorithms for solving OT problems, with emphasis on scalability to large problems.
- Bridge discrete and continuous OT and discuss practical computation for data-science applications.
- Survey generalizations and extensions of OT and connect them to related statistical and information-theoretic approaches.
Proposed method
- Introduce histograms and measures and define the OT problem in Monge and Kantorovich forms.
- Derive the Kantorovich relaxation as a linear program over couplings, enabling mass splitting.
- Discuss dual formulations, network-based and auction-type algorithms, and the structure of transport plans.
- Describe entropic regularization and Sinkhorn’s algorithm, including stability and log-domain implementation.
- Cover semidiscrete OT, W1, dynamic formulations, and extensions such as Gromov–Wasserstein and sliced transports.
Experimental results
Research questions
- RQ1How can optimal transport be formulated and solved efficiently for large-scale discrete and continuous measures?
- RQ2What are the key algorithmic strategies (e.g., Kantorovich relaxation, entropic regularization) that enable scalable OT computations?
- RQ3How do discrete and continuous OT relate, and what are the practical bridges between them for data-science applications?
- RQ4What are important extensions and generalizations of OT, and how do they connect to related inference and information-theoretic concepts?
Key findings
- OT can be formulated via Monge maps or Kantorovich couplings, with the latter enabling mass splitting and convex optimization.
- Kantorovich’s formulation leads to a linear program over the Birkhoff polytope, ensuring computational tractability and existence of solutions.
- Entropic regularization yields faster, more scalable algorithms (e.g., Sinkhorn) with stable implementations in the log-domain.
- The book/monograph connects OT to semidiscrete settings, Wp distances, dynamic formulations, and extensions such as Gromov–Wasserstein and sliced OT.
- The work highlights the interplay between OT theory and numerical methods, kernel methods, and information theory for practical data-science problems.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.