심규홍 교수
Gyu-Hong Shim
성균관대학교 소프트웨어학과 · 컴퓨터과학
연구실 소개
심규홍 교수의 연구실은 인공지능 기반의 첨단 신호 처리 기술을 핵심으로 삼고 있으며, 특히 컴퓨터 비전과 딥러닝을 융합한 혁신적 연구를 선도하고 있습니다. mmWave 통신의 비용 효율적인 빔 매니지먼트 기술, 일반화된 제로샷 러닝을 활용한 이미지 인식, 그리고 약한 감독 신호를 활용한 세분화된 세그멘테이션 기술 등 다양한 분야에서 고성능 AI 모델을 설계하고 있습니다. 특히, 비전 트랜스포머 기반의 특징 추출, 상대적 깊이 정보를 활용한 단일 영상 깊이 추정 등 첨단 기술을 접목한 연구로, 실생활 응용에 적합한 지능형 비전 시스템을 개발하고 있습니다.
연구 현황
연구 성과 추이
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
주요 논문
15Beamforming technique realized by the multipleinput-multiple-output (MIMO) antenna arrays has been widely used to compensate for the severe path loss in the millimeter wave (mmWave) bands. In 5G NR system, the beam sweeping and beam refinement are employed to find out the best beam codeword aligned to the mobile. Due to the complicated handshaking and finite resolution of the codebook, today's 5G-based beam management strategy is ineffective in various scenarios in terms of the data rate, energy
Generalized zero-shot learning (GZSL) is a technique to train a deep learning model to identify unseen classes using the attribute. In this paper, we put forth a new GZSL technique that improves the GZSL classification performance greatly. Key idea of the proposed approach, henceforth referred to as semantic feature extraction-based GZSL (SE-GZSL), is to use the semantic feature containing only attribute-related information in learning the relationship between the image and the attribute. In doi
We propose a fast approximation method of a softmax function with a very large vocabulary using singular value decomposition (SVD). SVD-softmax targets fast and accurate probability estimation of the topmost probable words during inference of neural network language models. The proposed method transforms the weight matrix used in the calculation of the output vector by using SVD. The approximate probability of each word can be estimated with only a small part of the weight matrix by using a few
Generalized zero-shot learning (GZSL) is a technique to train a deep learning model to identify unseen classes using the image attribute. In this paper, we put forth a new GZSL technique exploiting Vision Transformer (ViT) to maximize the attribute-related information contained in the image feature. In ViT, the entire image region is processed without the degradation of the image resolution and the local image information is preserved in patch features. To fully enjoy the benefits of ViT, we exp
Weakly-supervised semantic segmentation (WSSS) aims to train a semantic segmentation network using weak labels. Recent approaches generate the pseudo-label from the image-level label and then exploit it as a pixel-level supervision in the segmentation network training. A potential drawback of the conventional WSSS approaches is that the pseudo-label cannot accurately express the object regions and their classes, causing a degradation of the segmentation performance. In this paper, we propose a n
Recently, we have observed that Large Multimodal Models (LMMs) are revolutionizing the way machines interact with the world, unlocking new possibilities across various multimodal applications.To adapt LMMs for downstream tasks, parameter-efficient fine-tuning (PEFT) which only trains additional prefix tokens or modules, has gained popularity.Nevertheless, there has been little analysis of how PEFT works in LMMs.In this paper, we delve into the strengths and weaknesses of each tuning strategy, sh
Monocular depth estimation is very challenging because clues to the exact depth are incomplete in a single RGB image. To overcome the limitation, deep neural networks rely on various visual hints such as size, shade, and texture extracted from RGB information. However, we observe that if such hints are overly exploited, the network can be biased on RGB information without considering the comprehensive view. We propose a novel depth estimation model named RElative Depth Transformer (RED-T) that u
Image-text retrieval is a task to search for the proper textual descriptions of the visual world and vice versa. One challenge of this task is the vulnerability to input image/text corruptions. Such corruptions are often unobserved during the training, and degrade the retrieval model’s decision quality substantially. In this paper, we propose a novel image-text retrieval technique, referred to as robust visual semantic embedding (RVSE), which consists of novel image-based and text-based augmenta
In few-shot open-set recognition (FSOSR), a network learns to recognize closed-set samples with a few support samples while rejecting open-set samples with no class cue. Unlike conventional OSR, the FSOSR considers more practical open worlds where a closed-set class can be selected as an open-set class in another testing (task) and vice versa. Existing FSOSR methods have commonly represented the open set with task-dependent extra modules. These modules decently handle the varied closed and open
Efficient implementation of deep neural networks on CPU-based systems is very critical because applications proliferate to embedded and Internet of Things (IoT) systems. Many CPUs for personal computers and embedded systems equip Single Instruction Multiple Data (SIMD) instructions, which can be utilized to implement an efficient GEneral Matrix Multiply (GEMM) library that is very necessary for efficient deep neural network implementation. While many deep neural networks show quite good performa
Unsupervised semantic segmentation (USS) aims to discover and recognize meaningful categories without any labels. For a successful USS, two key abilities are required: 1) information compression and 2) clustering capability. Previous methods have relied on feature dimension reduction for information compression, however, this approach may hinder the process of clustering. In this paper, we propose a novel USS framework called Expand-and-Quantize Unsupervised Semantic Segmentation (EQUSS), which
Transformer-based networks have achieved great success in image captioning because of the attention mechanism that finds relevant image locations for each word. However, the current cross-attention process, which aligns word-to-image, does not consider the spatial relationships existing in patch-to-patch. This lack of spatial information may cause incorrect descriptions that fail at generating words that correctly describe the positional relationships. In this paper, we introduce a novel cross-a
The customization of large language models (LLMs) for user-specified tasks gets important.However, maintaining all the customized LLMs on cloud servers incurs substantial memory and computational overheads, and uploading user data can also lead to privacy concerns.Ondevice LLMs can offer a promising solution by mitigating these issues.Yet, the performance of on-device LLMs is inherently constrained by the limitations of small-scaled models.To overcome these restrictions, we first propose Crayon,
Phoneme recognition is a very important part of speech recognition that requires the ability to extract phonetic features from multiple frames. In this paper, we compare and analyze CNN, RNN, Transformer, and Conformer models using phoneme recognition. For CNN, the ContextNet model is used for the experiments. First, we compare the accuracy of various architectures under different constraints, such as the receptive field length, parameter size, and layer depth. Second, we interpret the performan
대표 연구 분야
심규홍 교수의 연구를 Nubint에서 더 깊이 살펴보세요
이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.