Junyeop Lee
Sungkyunkwan University · 工学
研究室紹介
Professor Junyeop Lee's research lab specializes in computer vision and deep learning, with a strong focus on scene understanding, text recognition, and image super-resolution. The lab develops advanced neural network architectures for challenging tasks such as recognizing arbitrary-shaped text, enhancing low-resolution images, and editing scene text while preserving style. Their work emphasizes model generalization, efficiency, and real-world applicability, particularly in autonomous driving and mobile computing environments.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15Many new proposals for scene text recognition (STR) models have been introduced in recent years. While each claim to have pushed the boundary of the technology, a holistic and fair comparison has been largely missing in the field due to the inconsistent choices of training and evaluation datasets. This paper addresses this difficulty with three major contributions. First, we examine the inconsistencies of training and evaluation datasets, and the performance gap results from inconsistencies. Sec
Single image extreme Super Resolution (SR) is a difficult task as scale factor in the order of 10X or greater is typically attempted. For instance, in the case of 16x upscale of an image, a single pixel from a low resolution image gets expanded to a 16x16 image patch. Such attempts often result fuzzy quality and loss in details in reconstructed images. To handle these difficulties, we propose a network architecture composed of a series of connected blocks in recurrent and feedback fashions for e
Scene text recognition (STR) is the task of recognizing character sequences in natural scenes. While there have been great advances in STR methods, current methods which convert two-dimensional (2D) image to one-dimensional (1D) feature map still fail to recognize texts in arbitrary shapes, such as heavily curved, rotated or vertically aligned texts, which are abundant in daily life (e.g. restaurant signs, product labels, company logos, etc). This paper introduces an architecture to recognize te
Scene text editing (STE), which converts a text in a scene image into the desired text while preserving an original style, is a challenging task due to a complex intervention between text and style. To address this challenge, we propose a novel representational learning-based STE model, referred to as RewriteNet that employs textual information as well as visual information. We assume that the scene text image can be decomposed into content and style features where the former represents the text
본 논문은 자율주행을 위한 실시간 의미론적 분할 방법으로 최적화된 심층 신경망 구조인 Wide Inception ResNet (WIR Net)을 제안한다. 신경망 구조는 Residual connection과 Inception module을 적용하여 특징을 추출하는 인코더와 Transposed convolution과 낮은 층의 특징 맵을 사용하여 해상도를 높이는 디코더로 구성하였고 ELU 활성화 함수를 적용함으로써 성능을 올렸다. 또한 신경망의 전체 층수를 줄이고 필터 수를 늘리는 방법을 통해 성능을 최적화하였다. 성능평가는 NVIDIA Geforce gtx 1080과 TX1 보드를 사용하여 주행환경의 Cityscapes 데이터에 대해 클래스와 카테고리별 IoU를 평가하였다. 실험 결과를 통해 클래스 IoU 53.4, 카테고리 IoU 81.8의 정확도와 TX1 보드에서 640×360, 720×480 해상도 영상처리에 17.8fps, 13.0fps의 실행속도를 보여주는 것을 확인하였다.
Delay tolerant networks generally use the epidemic routing protocol, because stable routing paths can hardly be maintained. However, the epidemic scheme leads to much network overheads because the messages are delivered to all the devices. In this paper, we propose an efficient protocol based on the human mobility patterns. In the proposed protocol, a message is delivered to several devices that are expected to deliver well the message to the destination. The proposed protocol relies on the huma
Abstract Vanadium oxides (VOx) are representative materials with a high temperature coefficient of resistance (TCR); however, VOx films can have complex phase structures that are dependent on their fabrication method. While past research has focused on the TCR behavior of VOx thin films, this study investigates the TCR of VOx thin films annealed at different temperatures as well as focuses on the relation between the VOx phase, surface morphology, sheet resistance, and TCR. VOx thin films were d
Weakly supervised video anomaly detection is a methodology that assesses anomaly levels in individual frames based on labeled video data. Anomaly scores are computed by evaluating the deviation of distances derived from frames in an unbiased state. Weakly supervised video anomaly detection encounters the formidable challenge of false alarms, stemming from various sources, with a major contributor being the inadequate reflection of frame labels during the learning process. Multiple instance learn
Abstract The relationship between the transmittance and FWHM of a Fabry–Perot filter for a nondispersive carbon dioxide (CO 2 ) sensor was investigated as a function of the number of distributed Bragg reflector (DBR) pairs consisting poly-Si and SiO 2 thin films. Given the significant prior research on the fabrication of high-performance Fabry–Perot filters, this study is focused on the relationship between the transmittance and FWHM that can be achieved by controlling the reflectance of the DBR
Hydrogen sulfide(H 2 S) is a highly toxic, corrosive, flammable gas. It is produced from the microbial breakdown of organic matter in the absence of oxygen. H 2 S also occurs naturally in volcanic gases, hot springs, and industrially in waste water treatment, gas drilling, tanneries, etc. H 2 S is colorless, and it reacts with metal ions to form metal sulfides, which are insoluble. For example, lead(II) acetate, soluble white crystalline, is converted lead(II) sulfide(PbS) which is black colored
Scene text editing (STE), which converts a text in a scene image into the desired text while preserving an original style, is a challenging task due to a complex intervention between text and style. In this paper, we propose a novel STE model, referred to as RewriteNet, that decomposes text images into content and style features and re-writes a text in the original image. Specifically, RewriteNet implicitly distinguishes the content from the style by introducing scene text recognition. Additiona
In this study, we propose a novel framework for time-series representation learning that integrates a learnable masking-augmentation strategy into a contrastive learning framework. Time-series data pose challenges due to their temporal dependencies and feature-extraction complexities. To address these challenges, we introduce a masking-based reconstruction approach within a contrastive learning context, aiming to enhance the model's ability to learn discriminative temporal features. Our method l