[Paper Review] Student-t Processes as Alternatives to Gaussian Processes
This paper proposes Student-t processes (TPs) as a flexible alternative to Gaussian processes (GPs), deriving them by marginalizing over an inverse Wishart process prior on the GP kernel. The TP retains analytical marginal and predictive distributions, enables nonparametric modeling with kernel-based flexibility, and uniquely exhibits predictive covariances that depend on training data values—improving robustness in Bayesian optimization and regression with no computational overhead over GPs.
We investigate the Student-t process as an alternative to the Gaussian process as a nonparametric prior over functions. We derive closed form expressions for the marginal likelihood and predictive distribution of a Student-t process, by integrating away an inverse Wishart process prior over the covariance kernel of a Gaussian process model. We show surprising equivalences between different hierarchical Gaussian process models leading to Student-t processes, and derive a new sampling scheme for the inverse Wishart process, which helps elucidate these equivalences. Overall, we show that a Student-t process can retain the attractive properties of a Gaussian process -- a nonparametric representation, analytic marginal and predictive distributions, and easy model selection through covariance kernels -- but has enhanced flexibility, and predictive covariances that, unlike a Gaussian process, explicitly depend on the values of training observations. We verify empirically that a Student-t process is especially useful in situations where there are changes in covariance structure, or in applications like Bayesian optimization, where accurate predictive covariances are critical for good performance. These advantages come at no additional computational cost over Gaussian processes.
Motivation & Objective
- To address the limitations of Gaussian processes in modeling uncertainty and handling model misspecification or structural changes in covariance.
- To formalize the inverse Wishart process as a nonparametric prior over covariance matrices for use in hierarchical GP models.
- To derive a Student-t process with closed-form marginal and predictive distributions, enabling practical use in regression and optimization.
- To demonstrate that predictive covariances in TPs depend on training observations—unlike in GPs—offering improved robustness and tail dependence.
- To show that TPs can be used in place of GPs with no computational penalty, while providing superior performance in key applications like Bayesian optimization.
Proposed method
- Derive the Student-t process by placing an inverse Wishart process prior over the covariance kernel of a Gaussian process, then integrating it out analytically.
- Use the inverse Wishart process as a nonparametric prior over covariance matrices of arbitrary size, ensuring consistency under marginalization.
- Derive closed-form expressions for the marginal likelihood and predictive distribution of the TP, including analytic derivatives for hyperparameter optimization.
- Propose a novel sampling scheme for the inverse Wishart process to clarify equivalences in hierarchical GP models and improve interpretability.
- Implement a marginalized expected improvement acquisition function in Bayesian optimization by integrating over hyperparameters using slice sampling.
- Compare TP and GP performance on synthetic and benchmark functions using identical kernels and hyperparameter inference methods.
Experimental results
Research questions
- RQ1Can a Student-t process be derived as a hierarchical generalization of a Gaussian process with analytically tractable marginals and conditionals?
- RQ2How do the predictive covariances of a Student-t process differ from those of a Gaussian process, particularly in their dependence on training data values?
- RQ3What is the role of the inverse Wishart process as a nonparametric prior over covariance matrices in constructing a Student-t process?
- RQ4In what scenarios does a Student-t process outperform a Gaussian process, particularly in Bayesian optimization and regression with structural changes?
- RQ5Can the Student-t process be used as a drop-in replacement for Gaussian processes with no additional computational cost?
Key findings
- The Student-t process is the most general elliptically symmetric process with analytically tractable marginal and predictive distributions, extending beyond the Gaussian process.
- Predictive covariances in the Student-t process explicitly depend on the values of training observations, unlike in Gaussian processes, enabling better modeling of uncertainty and tail dependence.
- In Bayesian optimization, the Student-t process outperforms the Gaussian process, finding the minimum of a 1D sinusoidal function 25% faster on average (8.1±0.4 vs. 10.7±0.6 iterations).
- On the 2D Branin-Hoo and 6D Hartmann functions, the TP shows more thorough exploration of local minima, behaving like a step function, while the GP shows more uniform improvement.
- The TP demonstrates improved robustness to model misspecification and changes in covariance structure, particularly in high-dimensional settings.
- The Student-t process can be used with an analytic noise model that separates signal and noise, unlike previous formulations, and supports nonparametric kernel learning with no computational overhead.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.