Jun-Hyuk Jang
Hanyang University · Computer Science
About the Lab
Professor Jun-Hyuk Jang's research lab specializes in speech and audio signal processing, with a focus on robust voice activity detection (VAD), speech enhancement, and non-invasive sleep stage classification. The lab develops advanced statistical models—such as complex Laplacian, Gamma, and generalized gamma distributions—for analyzing speech and image transform coefficients, particularly in the DCT and DFT domains. A key innovation is the use of warped transforms (e.g., WDCT) to better model perceptually relevant frequency characteristics under noisy conditions. The lab also pioneers multi-modal, non-contact sensing techniques using radar and audio for clinical-grade sleep monitoring, especially for patients with sleep disorders.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15One of the key issues in practical speech processing is to achieve robust voice activity detection (VAD) against the background noise. Most of the statistical model-based approaches have tried to employ the Gaussian assumption in the discrete Fourier transform (DFT) domain, which, however, deviates from the real observation. In this paper, we propose a class of VAD algorithms based on several statistical models. In addition to the Gaussian model, we also incorporate the complex Laplacian and Gam
A voice activity detector (VAD) based on the complex Laplacian model is proposed. The likelihood ratio based on the Laplacian model is computed and then applied to the VAD operation. According to experimental results, it is found that the Laplacian statistical model is more efficient for the VAD algorithm compared to the Gaussian model.
We propose an approach based on the warped discrete cosine transform (WDCT) to enhance degraded speech under background noise environments. To develop an effective expression for the frequency characteristics of the input speech, we apply a variable frequency warping filter to the conventional discrete cosine transform (DCT). The frequency warping control parameter is adjusted according to an analysis of the spectral distribution in each frame. For a more accurate analysis of spectral characteri
In this letter, we propose results of distribution tests that indicate that for many natural images, the statistics of the discrete cosine transform (DCT) coefficients are best approximated by a generalized gamma function (G/spl Gamma/F), which includes the conventional Gaussian, Laplacian, and gamma probability density functions. The major parameter of the G/spl Gamma/F is estimated according to the maximum likelihood (ML) principle. Experimental results on a number of /spl chi//sup 2/ tests in
In this paper, a warped discrete cosine transform (WDCT)-based approach to enhance the degraded speech under background noise environments is proposed. For developing an effective expression of the frequency characteristics of the input speech, the variable frequency warping filter is applied to the conventional discrete cosine transform (DCT). The frequency warping control parameter is adjusted according to the analysis of spectral distribution in each frame. For a more accurate analysis of spe
Polysomnography (PSG) is considered as the gold standard for determining sleep stages, but due to the obtrusiveness of its sensor attachments, sleep stage classification algorithms using noninvasive sensors have been developed throughout the years. However, the previous studies have not yet been proven reliable. In addition, most of the products are designed for healthy customers rather than for patients with sleep disorder. We present a novel approach to classify sleep stages via low cost and n
We propose a line-of-sight (LOS)/non-line-of-sight (NLOS) mixture source localization algorithm that utilizes the weighted least squares (WLS) method in LOS/NLOS mixture environments, where the weight matrix is determined in the algebraic form. Unless the contamination ratio exceeds 50 %, the asymptotic variance of the sample median can be approximately related to that of the sample mean. Based on this observation, we use the error covariance matrix for the sample mean and median to minimize the
A novel approach to a voice activity detector (VAD) in noisy environments is presented. The generalised Gaussian distribution (GGD) is employed as a parametric model for noisy speech, which enables tuning to the actual data. According to the experimental results, it was discovered that the proposed GGD model is more effective for the VAD algorithm compared to the conventional Laplacian model.
This paper proposes a voice activity detector (VAD) based on the complex Laplacian model. With the use of a goodness-of-fit (GOF) test, it is discovered that the Laplacian model is more suitable to describe noisy speech distribution than the conventional Gaussian model. The likelihood ratio (LR) based on the Laplacian model is computed and then applied to the VAD operation. According to the experimental results, we can find that the Laplacian statistical model is more suitable for the VAD algori
Research Areas
Dive deeper into Jun-Hyuk Jang's research on Nubint
Open this lab's papers in the app to read with AI, summarize, and cite in your writing.