Skip to main content
QUICK REVIEW

[Paper Review] Robust inversion via semistochastic dimensionality reduction

Aleksandr Y. Aravkin, Michael P. Friedlander|arXiv (Cornell University)|Oct 5, 2011
Sparse and Compressive Sensing Techniques21 references17 citations
TL;DR

This paper proposes a robust inversion framework using the Student’s t-distribution penalty combined with semistochastic dimensionality reduction via randomized sampling to handle large-scale seismic inverse problems with heavy-tailed noise and up to 50% corrupted data. The method achieves superior convergence and accuracy over least-squares and Huber penalties by allowing flexible residual distributions, validated through full-waveform inversion experiments with limited-memory BFGS and adaptive sampling strategies.

ABSTRACT

We consider a class of inverse problems where it is possible to aggregate the results of multiple experiments. This class includes problems where the forward model is the solution operator to linear ODEs or PDEs. The tremendous size of such problems motivates dimensionality reduction techniques based on randomly mixing experiments. These techniques break down, however, when robust data-fitting formulations are used, which are essential in cases of missing data, unusually large errors, and systematic features in the data unexplained by the forward model. We survey robust methods within a statistical framework, and propose a semistochastic optimization approach that allows dimensionality reduction. The efficacy of the methods are demonstrated for a large-scale seismic inverse problem using the robust Student's t-distribution, where a useful synthetic velocity model is recovered in the extreme scenario of 60% data missing at random. The semistochastic approach achieves this recovery using 20% of the effort required by a direct robust approach.

Motivation & Objective

  • Address the limitations of least-squares and Huber penalties in handling heavy-tailed noise and outliers in large-scale inverse problems.
  • Develop a robust estimation framework based on the Student’s t-distribution that better models real-world data artifacts and outliers.
  • Integrate dimensionality reduction via randomized sampling to reduce computational cost while preserving convergence properties.
  • Enable scalable, efficient inversion for full-waveform seismic imaging with massive datasets using stochastic optimization techniques.
  • Demonstrate the superiority of the proposed method over classical approaches in terms of model accuracy and robustness under data corruption.

Proposed method

  • Formulate the inverse problem using a robust penalty function derived from the Student’s t-distribution to model residuals without assuming Gaussian or Laplace noise.
  • Apply a semistochastic dimensionality reduction strategy by forming weighted averages of data groups (meta-experiments) to reduce problem size while preserving statistical properties.
  • Use randomized sampling to approximate the full gradient in stochastic optimization, where each iteration samples a small, dynamically increasing batch of experiments.
  • Implement a limited-memory BFGS quasi-Newton method with adaptive Hessian approximation to accelerate convergence while maintaining low memory usage.
  • Employ an Armijo backtracking linesearch on the sampled objective to ensure sufficient decrease in the objective function at each step.
  • Leverage theoretical bounds on the second moment of gradient error to ensure convergence in expectation, based on the expected distance to the optimal solution.

Experimental results

Research questions

  • RQ1Can a robust penalty based on the Student’s t-distribution outperform classical least-squares and Huber penalties in the presence of outliers and corrupted data?
  • RQ2How effective is randomized sampling combined with dimensionality reduction in maintaining convergence speed and solution accuracy for large-scale inverse problems?
  • RQ3To what extent does the proposed semistochastic method preserve the convergence rate of full-gradient methods while reducing per-iteration cost?
  • RQ4How does the residual distribution evolve during optimization under different penalty functions, and does the Student’s t penalty allow for more realistic residual shapes?
  • RQ5What is the impact of adaptive batch size and stochastic Hessian approximation on the convergence behavior of the optimization algorithm?

Key findings

  • The Student’s t-penalty achieved the best model reconstruction accuracy, with the relative model error decreasing steadily over iterations and remaining lowest among all tested methods.
  • Least-squares optimization failed to reduce model error, indicating poor robustness to 50% data corruption and non-Gaussian noise.
  • Huber penalty showed initial improvement but suffered from increasing error after ~20 iterations, suggesting sensitivity to outlier distribution.
  • The sampling-based optimization method achieved convergence rates comparable to the full-gradient method while maintaining low per-iteration computational cost.
  • The evolution of the sampled data size increased gradually by one element per iteration, demonstrating a controlled and adaptive sampling strategy.
  • Histograms of residuals after 50 iterations confirmed that only the Student’s t penalty allowed the residual distribution to evolve toward the true, heavy-tailed distribution, unlike the forced shapes imposed by Gaussian or Laplace priors.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.