Skip to main content

Joo-Han Noh

Korea Advanced Institute of Science and Technology · Computer Science

About the Lab

Professor Joo-Han Noh's research lab specializes in deep learning for audio and multimodal signal processing, with a strong focus on end-to-end learning from raw waveforms. The lab explores innovative convolutional neural network architectures—such as SampleCNN and its advanced variants—that use sample-level filters to extract hierarchical representations directly from audio signals, achieving state-of-the-art performance in music auto-tagging and audio classification. A key research direction involves multimodal representation learning, where features from multiple modalities (e.g., audio and video) are jointly learned to improve model generalization and performance. The lab also investigates the integration of deep learning with real-world music streaming and intelligent audio systems, including applications in smart speakers and personalized music recommendation platforms.

end-to-end learningraw audiosample-level CNNmultimodal representationmusic auto-tagging

Research Overview

Papers
217
Total Citations
4,110
Papers (5y)
105
Primary Field
Computer Science

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
105total
2022
2023
2024
2025
2026
Citations per year (5y)
183total
20222023202420252026

Selected Papers

15
1
Article|2,289 citations·2011
Multimodal Deep Learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, Andrew Y. Ng

Deep networks have been successfully applied to unsupervised feature learning for single modalities (e.g., text, images or audio). In this work, we propose a novel application of deep networks to learn features over multiple modalities. We present a series of tasks for multimodal learning and show how to train deep networks that learn features to address these tasks. In particular, we demonstrate cross modality feature learning, where better features for one modality (e.g., video) can be learned

Signal ProcessingComputer Science
2
Article|129 citations·2018
SampleCNN: End-to-End Deep Convolutional Neural Networks Using Very Small Filters for Music Classification
Jongpil Lee, Jiyoung Park, Keunhyoung Kim, Juhan Nam
SJR Q2Applied SciencesOA

Convolutional Neural Networks (CNN) have been applied to diverse machine learning tasks for different modalities of raw data in an end-to-end fashion. In the audio domain, a raw waveform-based approach has been explored to directly learn hierarchical characteristics of audio. However, the majority of previous studies have limited their model capacity by taking a frame-level structure similar to short-time Fourier transforms. We previously proposed a CNN architecture which learns representations

Signal ProcessingComputer Science
3
Article|115 citations·2018
Deep Learning for Audio-Based Music Classification and Tagging: Teaching Computers to Distinguish Rock from Bach
Juhan Nam, Keunwoo Choi, Jongpil Lee, Szu-Yu Chou, Yi‐Hsuan Yang
SJR Q1IEEE Signal Processing Magazine

Over the last decade, music-streaming services have grown dramatically. Pandora, one company in the field, has pioneered and popularized streaming music by successfully deploying the Music Genome Project [1] (https://www.pandora.com/about/mgp) based on human-annotated content analysis. Another company, Spotify, has a catalog of over 40 million songs and over 180 million users as of mid-2018 (https://press.spotify.com/us/about/), making it a leading music service provider worldwide. Giant technol

Signal ProcessingComputer Science
4
Preprint|104 citations·2017
Sample-level Deep Convolutional Neural Networks for Music Auto-tagging Using Raw Waveforms
Jongpil Lee, Ji Young Park, Keunhyoung Luke Kim, Juhan Nam
arXiv (Cornell University)OA

Recently, the end-to-end approach that learns hierarchical representations from raw data using deep convolutional neural networks has been successfully explored in the image, text and speech domains. This approach was applied to musical signals as well but has been not fully explored yet. To this end, we propose sample-level deep convolutional neural networks which learn representations from very small grains of waveforms (e.g. 2 or 3 samples) beyond typical frame-level input representations. Ou

Signal ProcessingComputer Science
5
Article|91 citations·2018
Sample-Level CNN Architectures for Music Auto-Tagging Using Raw Waveforms
Tae‐Jun Kim, Jongpil Lee, Juhan Nam

Recent work has shown that the end-to-end approach using convolutional neural network (CNN) is effective in various types of machine learning tasks. For audio signals, the approach takes raw waveforms as input using an 1-D convolution layer. In this paper, we improve the 1-D CNN architecture for music auto-tagging by adopting building blocks from state-of-the-art image classification models, ResNets and SENets, and adding multi-level feature aggregation to it. We compare different combinations o

Signal ProcessingComputer Science
6
Article|88 citations·2019
Comparison and Analysis of SampleCNN Architectures for Audio Classification
Taejun Kim, Jongpil Lee, Juhan Nam
SJR Q1IEEE Journal of Selected Topics in Signal Processing

End-to-end learning with convolutional neural networks (CNNs) has become a standard approach in image classification. However, in audio classification, CNN-based models that use time-frequency representations as input are still popular. A recently proposed CNN architecture called SampleCNN takes raw waveforms directly and has very small sizes of filters. The architecture has proven to be effective in music classification tasks. In this paper, we scrutinize SampleCNN further by comparing it with

Signal ProcessingComputer Science
7
Article|81 citations·2019
Joint Detection and Classification of Singing Voice Melody Using Convolutional Recurrent Neural Networks
Sangeun Kum, Juhan Nam
SJR Q2Applied SciencesOA

Singing melody extraction essentially involves two tasks: one is detecting the activity of a singing voice in polyphonic music, and the other is estimating the pitch of a singing voice in the detected voiced segments. In this paper, we present a joint detection and classification (JDC) network that conducts the singing voice detection and the pitch estimation simultaneously. The JDC network is composed of the main network that predicts the pitch contours of the singing melody and an auxiliary ne

Signal ProcessingComputer Science
8
Article|75 citations·2011
A Classification-Based Polyphonic Piano Transcription Approach Using Learned Feature Representations.
Juhan Nam, Jiquan Ngiam, Honglak Lee, Malcolm Slaney
OA

[TODO] Add abstract here.

Signal ProcessingComputer Science
9
Article|73 citations·2016
Melody Extraction On Vocal Segments Using Multi-Column Deep Neural Networks.
Sangeun Kum, Changheun Oh, Juhan Nam
Zenodo (CERN European Organization for Nuclear Research)OA

[TODO] Add abstract here.

Signal ProcessingComputer Science
10
Article|64 citations·2012
Learning Sparse Feature Representations For Music Annotation And Retrieval.
Juhan Nam, Jorge Herrera, Malcolm Slaney, Julius O. Smith
OA

[TODO] Add abstract here.

Signal ProcessingComputer Science
11
Article|53 citations·2020
Disentangled Multidimensional Metric Learning for Music Similarity
Jongpil Lee, Nicholas J. Bryan, Justin Salamon, Zeyu Jin, Juhan Nam

Music similarity search is useful for a variety of creative tasks such as replacing one music recording with another recording with a similar "feel", a common task in video editing. For this task, it is typically necessary to define a similarity metric to compare one recording to another. Music similarity, however, is hard to define and depends on multiple simultaneous notions of similarity (i.e. genre, mood, instrument, tempo). While prior work ignore this issue, we embrace this idea and introd

Signal ProcessingComputer Science
12
Preprint|45 citations·2017
Raw Waveform-based Audio Classification Using Sample-level CNN Architectures
Jongpil Lee, Tae‐Jun Kim, Ji Young Park, Juhan Nam
arXiv (Cornell University)OA

Music, speech, and acoustic scene sound are often handled separately in the audio domain because of their different signal characteristics. However, as the image domain grows rapidly by versatile image classification models, it is necessary to study extensible classification models in the audio domain as well. In this study, we approach this problem using two types of sample-level deep convolutional neural networks that take raw waveforms as input and uses filters with small granularity. One is

Signal ProcessingComputer Science
13
Article|19 citations·2010
A super-resolution spectrogram using coupled PLCA
Juhan Nam, Gautham J. Mysore, Joachim Ganseman, Kyogu Lee, Jonathan S. Abel

The short-time Fourier transform (STFT) based spectrogram is commonly used to analyze the time-frequency content of a signal. Depending on window size, the STFT provides a trade-off between time and frequency resolutions. This paper presents a novel method that achieves high resolution simultaneously in both time and frequency. We extend Probabilistic Latent Component Analysis (PLCA) to jointly decompose two spectrograms, one with a high time resolution and one with a high frequency resolution.

Signal ProcessingComputer Science
14
Article|16 citations·2009
Efficient Antialiasing Oscillator Algorithms Using Low-Order Fractional Delay Filters
Juhan Nam, Vesa Välimäki, Jonathan S. Abel, Julius O. Smith
IEEE Transactions on Audio Speech and Language Processing

One of the challenges in virtual analog synthesis is avoiding aliasing when generating classic waveforms such as sawtooth and square wave which have theoretically infinite bandwidth in their ideal forms. The human auditory system renders a certain amount of aliasing inaudible, which allows room for finding cost-effective algorithms. This paper suggests efficient algorithms to reduce the aliasing using low-order fractional delay filters in the framework of bandlimited impulse train (BLIT) synthes

Signal ProcessingComputer Science
15
Article|12 citations·2008
On the Minimum-Phase Nature of Head-Related Transfer Functions
Juhan Nam, Miriam A. Kolar, Jonathan S. Abel
SJR Q1Journal of the Audio Engineering Society
Biomedical EngineeringEngineering

Research Areas

Signal ProcessingComputer Vision and Pattern RecognitionExperimental and Cognitive PsychologyArtificial IntelligenceCognitive NeuroscienceHuman-Computer Interaction

Dive deeper into Joo-Han Noh's research on Nubint

Open this lab's papers in the app to read with AI, summarize, and cite in your writing.