東京大学 · 数学
Sugasawa教授の研究室は、小領域推定やクラスタリングされたデータに対する統計的推論に注力しており、特に不確実性や外れ値に強いモデル構築と推定手法の開発を主眼としています。特に、密度パワー損失やγ損失に基づくロバストなベイズ推定、勾配ブースティングを用いた個別治療効果の推定など、機械学習的手法と統計的推論を融合した革新的なアプローチを展開しています。また、実用的で解釈可能なモデル構築と、情報量基準を用いたモデル選択手法の開発も進んでいます。
Figures are computed from collected data and may differ slightly.
Abstract For small area estimation of area‐level data, the Fay–Herriot model is extensively used as a model‐based method. In the Fay–Herriot model, it is conventionally assumed that the sampling variances are known, whereas estimators of sampling variances are used in practice. Thus, the settings of knowing sampling variances are unrealistic, and several methods are proposed to overcome this problem. In this paper, we assume the situation where the direct estimators of the sampling variances are
The development of molecular diagnostic tools to achieve individualized medicine requires accurate estimation of individual treatment effects (ITEs). Although several effective data analytic strategies have been proposed for this purpose, they have limitations when it comes to flexibly capturing the complex relationships between clinical outcome and possibly high-dimensional covariates. In this article, we propose an effective machine learning method to estimate ITEs using the gradient boosting
Although robust divergence, such as density power divergence and γ-divergence, is helpful for robust statistical inference in the presence of outliers, the tuning parameter that controls the degree of robustness is chosen in a rule-of-thumb, which may lead to an inefficient inference. We here propose a selection criterion based on an asymptotic approximation of the Hyvarinen score applied to an unnormalized model defined by robust divergence. The proposed selection criterion only requires first
Abstract We introduce a novel methodology for robust Bayesian estimation with robust divergence (e.g., density power divergence or $$\gamma $$ <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"> <mml:mi>γ</mml:mi> </mml:math> -divergence), indexed by tuning parameters. It is well known that the posterior density induced by robust divergence gives highly robust estimators against outliers if the tuning parameter is appropriately and carefully chosen. In a Bayesian framework, one way to find
Clustered data are ubiquitous in a variety of scientific fields. In this article, we propose a flexible and interpretable modeling approach, called grouped heterogeneous mixture modeling, for clustered data, which models cluster-wise conditional distributions by mixtures of latent conditional distributions common to all the clusters. In the model, we assume that clusters are divided into a finite number of groups and mixing proportions are the same within the same group. We provide a simple gene
Summary A two-stage normal hierarchical model called the Fay–Herriot model and the empirical Bayes estimator are widely used to obtain indirect and model-based estimates of means in small areas. However, the performance of the empirical Bayes estimator can be poor when the assumed normal distribution is misspecified. This article presents a simple modification that makes use of density power divergence and proposes a new robust empirical Bayes small area estimator. The mean squared error and est
Abstract For estimating area‐specific parameters (quantities) in a finite population, a mixed‐model prediction approach is attractive. However, this approach strongly depends on the normality assumption of the response values, although we often encounter a non‐normal case in practice. In such a case, transforming observations to make them suitable for normality assumption is a useful tool, but the problem of selecting a suitable transformation still remains open. To overcome the difficulty, we h
Random effects meta-analyses have been widely applied in evidence synthesis for various types of medical studies. However, standard inference methods (e.g. restricted maximum likelihood estimation) usually underestimate statistical errors and possibly provide highly overconfident results under realistic situations; for instance, coverage probabilities of confidence intervals can be substantially below the nominal level. The main reason is that these inference methods rely on large sample approxi
The article considers a nested error regression model with heteroscedastic variance functions for analyzing clustered data, where the normality for the underlying distributions is not assumed. Classical methods in normal nested error regression models with homogenous variances are extended in the two directions: heterogeneous variance functions for error terms and non-normal distributions for random effects and error terms. Consistent estimators for model parameters are suggested, and second-ord
Open papers in the app to read, cite, and organize with AI.