이준엽 교수
Junyeop Lee
성균관대학교 화학공학부 · 공학
연구실 소개
이준엽 교수의 연구실은 컴퓨터 비전과 영상처리 분야에서 실시간 및 정밀한 시각 인식 기술을 연구하고 있습니다. 특히 장면 내 텍스트 인식, 초해상도 복원, 텍스트 편집, 의미론적 분할 등 다양한 영상 이해 기술을 개발하며, 실제 응용 환경(예: 자율주행)에 최적화된 효율적인 신경망 아키텍처 설계에 중점을 두고 있습니다. 연구는 모델의 정확도 향상과 동시에 실시간 처리 성능을 동시에 확보하는 데 기여하고 있습니다. 특히, 복잡한 텍스트 형태나 낮은 해상도 이미지에서도 뛰어난 성능을 내는 알고리즘 개발에 주력하고 있습니다.
연구 현황
연구 성과 추이
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
주요 논문
15Many new proposals for scene text recognition (STR) models have been introduced in recent years. While each claim to have pushed the boundary of the technology, a holistic and fair comparison has been largely missing in the field due to the inconsistent choices of training and evaluation datasets. This paper addresses this difficulty with three major contributions. First, we examine the inconsistencies of training and evaluation datasets, and the performance gap results from inconsistencies. Sec
Single image extreme Super Resolution (SR) is a difficult task as scale factor in the order of 10X or greater is typically attempted. For instance, in the case of 16x upscale of an image, a single pixel from a low resolution image gets expanded to a 16x16 image patch. Such attempts often result fuzzy quality and loss in details in reconstructed images. To handle these difficulties, we propose a network architecture composed of a series of connected blocks in recurrent and feedback fashions for e
Scene text recognition (STR) is the task of recognizing character sequences in natural scenes. While there have been great advances in STR methods, current methods which convert two-dimensional (2D) image to one-dimensional (1D) feature map still fail to recognize texts in arbitrary shapes, such as heavily curved, rotated or vertically aligned texts, which are abundant in daily life (e.g. restaurant signs, product labels, company logos, etc). This paper introduces an architecture to recognize te
Scene text editing (STE), which converts a text in a scene image into the desired text while preserving an original style, is a challenging task due to a complex intervention between text and style. To address this challenge, we propose a novel representational learning-based STE model, referred to as RewriteNet that employs textual information as well as visual information. We assume that the scene text image can be decomposed into content and style features where the former represents the text
본 논문은 자율주행을 위한 실시간 의미론적 분할 방법으로 최적화된 심층 신경망 구조인 Wide Inception ResNet (WIR Net)을 제안한다. 신경망 구조는 Residual connection과 Inception module을 적용하여 특징을 추출하는 인코더와 Transposed convolution과 낮은 층의 특징 맵을 사용하여 해상도를 높이는 디코더로 구성하였고 ELU 활성화 함수를 적용함으로써 성능을 올렸다. 또한 신경망의 전체 층수를 줄이고 필터 수를 늘리는 방법을 통해 성능을 최적화하였다. 성능평가는 NVIDIA Geforce gtx 1080과 TX1 보드를 사용하여 주행환경의 Cityscapes 데이터에 대해 클래스와 카테고리별 IoU를 평가하였다. 실험 결과를 통해 클래스 IoU 53.4, 카테고리 IoU 81.8의 정확도와 TX1 보드에서 640×360, 720×480 해상도 영상처리에 17.8fps, 13.0fps의 실행속도를 보여주는 것을 확인하였다.
Delay tolerant networks generally use the epidemic routing protocol, because stable routing paths can hardly be maintained. However, the epidemic scheme leads to much network overheads because the messages are delivered to all the devices. In this paper, we propose an efficient protocol based on the human mobility patterns. In the proposed protocol, a message is delivered to several devices that are expected to deliver well the message to the destination. The proposed protocol relies on the huma
Abstract Vanadium oxides (VOx) are representative materials with a high temperature coefficient of resistance (TCR); however, VOx films can have complex phase structures that are dependent on their fabrication method. While past research has focused on the TCR behavior of VOx thin films, this study investigates the TCR of VOx thin films annealed at different temperatures as well as focuses on the relation between the VOx phase, surface morphology, sheet resistance, and TCR. VOx thin films were d
Weakly supervised video anomaly detection is a methodology that assesses anomaly levels in individual frames based on labeled video data. Anomaly scores are computed by evaluating the deviation of distances derived from frames in an unbiased state. Weakly supervised video anomaly detection encounters the formidable challenge of false alarms, stemming from various sources, with a major contributor being the inadequate reflection of frame labels during the learning process. Multiple instance learn
Hydrogen sulfide(H 2 S) is a highly toxic, corrosive, flammable gas. It is produced from the microbial breakdown of organic matter in the absence of oxygen. H 2 S also occurs naturally in volcanic gases, hot springs, and industrially in waste water treatment, gas drilling, tanneries, etc. H 2 S is colorless, and it reacts with metal ions to form metal sulfides, which are insoluble. For example, lead(II) acetate, soluble white crystalline, is converted lead(II) sulfide(PbS) which is black colored
Abstract The relationship between the transmittance and FWHM of a Fabry–Perot filter for a nondispersive carbon dioxide (CO 2 ) sensor was investigated as a function of the number of distributed Bragg reflector (DBR) pairs consisting poly-Si and SiO 2 thin films. Given the significant prior research on the fabrication of high-performance Fabry–Perot filters, this study is focused on the relationship between the transmittance and FWHM that can be achieved by controlling the reflectance of the DBR
Scene text editing (STE), which converts a text in a scene image into the desired text while preserving an original style, is a challenging task due to a complex intervention between text and style. In this paper, we propose a novel STE model, referred to as RewriteNet, that decomposes text images into content and style features and re-writes a text in the original image. Specifically, RewriteNet implicitly distinguishes the content from the style by introducing scene text recognition. Additiona
대표 연구 분야
이준엽 교수의 연구를 Nubint에서 더 깊이 살펴보세요
이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.