장준혁 교수
Jun-Hyuk Jang
한양대학교 융합전자공학부 · 컴퓨터과학
연구실 소개
장준혁 교수의 연구실은 음성 처리와 생체 신호 분석을 중심으로, 소음 환경에서의 음성 인식 및 신뢰성 있는 수면 단계 분류 기술을 개발하고 있습니다. 특히, DFT 및 DCT 기반의 통계적 모델링을 활용해 소음 간섭에 강건한 음성 활성도 검출(VAD) 알고리즘과 주파수 왜곡을 적용한 WDCT 기반 음성 강화 기법을 연구하고 있으며, 레이더와 음성 기반 Context-aware 기술을 융합한 비침습적 수면 단계 분류 시스템도 개발 중입니다. 이는 임상적 적용이 가능한 정확도와 실용성을 갖춘 스마트 헬스 기술의 기반을 마련하고자 합니다.
연구 현황
연구 성과 추이
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
주요 논문
15One of the key issues in practical speech processing is to achieve robust voice activity detection (VAD) against the background noise. Most of the statistical model-based approaches have tried to employ the Gaussian assumption in the discrete Fourier transform (DFT) domain, which, however, deviates from the real observation. In this paper, we propose a class of VAD algorithms based on several statistical models. In addition to the Gaussian model, we also incorporate the complex Laplacian and Gam
A voice activity detector (VAD) based on the complex Laplacian model is proposed. The likelihood ratio based on the Laplacian model is computed and then applied to the VAD operation. According to experimental results, it is found that the Laplacian statistical model is more efficient for the VAD algorithm compared to the Gaussian model.
We propose an approach based on the warped discrete cosine transform (WDCT) to enhance degraded speech under background noise environments. To develop an effective expression for the frequency characteristics of the input speech, we apply a variable frequency warping filter to the conventional discrete cosine transform (DCT). The frequency warping control parameter is adjusted according to an analysis of the spectral distribution in each frame. For a more accurate analysis of spectral characteri
In this letter, we propose results of distribution tests that indicate that for many natural images, the statistics of the discrete cosine transform (DCT) coefficients are best approximated by a generalized gamma function (G/spl Gamma/F), which includes the conventional Gaussian, Laplacian, and gamma probability density functions. The major parameter of the G/spl Gamma/F is estimated according to the maximum likelihood (ML) principle. Experimental results on a number of /spl chi//sup 2/ tests in
In this paper, a warped discrete cosine transform (WDCT)-based approach to enhance the degraded speech under background noise environments is proposed. For developing an effective expression of the frequency characteristics of the input speech, the variable frequency warping filter is applied to the conventional discrete cosine transform (DCT). The frequency warping control parameter is adjusted according to the analysis of spectral distribution in each frame. For a more accurate analysis of spe
Polysomnography (PSG) is considered as the gold standard for determining sleep stages, but due to the obtrusiveness of its sensor attachments, sleep stage classification algorithms using noninvasive sensors have been developed throughout the years. However, the previous studies have not yet been proven reliable. In addition, most of the products are designed for healthy customers rather than for patients with sleep disorder. We present a novel approach to classify sleep stages via low cost and n
We propose a line-of-sight (LOS)/non-line-of-sight (NLOS) mixture source localization algorithm that utilizes the weighted least squares (WLS) method in LOS/NLOS mixture environments, where the weight matrix is determined in the algebraic form. Unless the contamination ratio exceeds 50 %, the asymptotic variance of the sample median can be approximately related to that of the sample mean. Based on this observation, we use the error covariance matrix for the sample mean and median to minimize the
A novel approach to a voice activity detector (VAD) in noisy environments is presented. The generalised Gaussian distribution (GGD) is employed as a parametric model for noisy speech, which enables tuning to the actual data. According to the experimental results, it was discovered that the proposed GGD model is more effective for the VAD algorithm compared to the conventional Laplacian model.
This paper proposes a voice activity detector (VAD) based on the complex Laplacian model. With the use of a goodness-of-fit (GOF) test, it is discovered that the Laplacian model is more suitable to describe noisy speech distribution than the conventional Gaussian model. The likelihood ratio (LR) based on the Laplacian model is computed and then applied to the VAD operation. According to the experimental results, we can find that the Laplacian statistical model is more suitable for the VAD algori
대표 연구 분야
장준혁 교수의 연구를 Nubint에서 더 깊이 살펴보세요
이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.