Jinwook Seo
Seoul National University · Computer Science
About the Lab
Professor Jinwook Seo's research lab specializes in visual analytics and interactive data exploration, with a focus on high-dimensional biological data such as microarray and gene expression datasets. The lab develops innovative visualization and clustering tools—like the Hierarchical Clustering Explorer (HCE) and the rank-by-feature framework—to help researchers identify patterns, clusters, outliers, and meaningful biological insights in complex multivariate data. Their work bridges the gap between statistical analysis and intuitive visual interaction, empowering biologists and data scientists to explore and interpret large-scale 'omics' data more effectively. The lab emphasizes user-centered design and evaluation of analytical tools to ensure practical utility in real-world biological research.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15To date, work in microarrays, sequenced genomes and bioinformatics has focused largely on algorithmic methods for processing and manipulating vast biological data sets. Future improvements will likely provide users with guidance in selecting the most appropriate algorithms and metrics for identifying meaningful clusters-interesting patterns in large data sets, such as groups of genes with similar profiles. Hierarchical clustering has been shown to be effective in microarray data analysis for ide
Interactive exploration of multidimensional data sets is challenging because: (1) it is difficult to comprehend patterns in more than three dimensions, and (2) current systems often are a patchwork of graphical and statistical methods leaving many researchers uncertain about how to explore their data in an orderly manner. We offer a set of principles and a novel rank-by-feature framework that could enable users to better understand distributions in one (1D) or two dimensions (2D), and then disco
MOTIVATION: Human clinical projects typically require a priori statistical power analyses. Towards this end, we sought to build a flexible and interactive power analysis tool for microarray studies integrated into our public domain HCE 3.5 software package. We then sought to determine if probe set algorithms or organism type strongly influenced power analysis results. RESULTS: The HCE 3.5 power analysis tool was designed to import any pre-existing Affymetrix microarray project, and interactively
MOTIVATION: The most commonly utilized microarrays for mRNA profiling (Affymetrix) include 'probe sets' of a series of perfect match and mismatch probes (typically 22 oligonucleotides per probe set). There are an increasing number of reported 'probe set algorithms' that differ in their interpretation of a probe set to derive a single normalized 'signal' representative of expression of each mRNA. These algorithms are known to differ in accuracy and sensitivity, and optimization has been done usin
Knowledge discovery in high-dimensional data is a challenging enterprise, but new visual analytic tools appear to offer users remarkable powers if they are ready to learn new concepts and interfaces. Our three-year effort to develop versions of the Hierarchical Clustering Explorer (HCE) began with building an interactive tool for exploring clustering results. It expanded, based on user needs, to include other potent analytic and visualization tools for multivariate data, especially the rank-by-f
Affymetrix microarrays have become a standard experimental platform for studies of mRNA expression profiling. Their success is due, in part, to the multiple oligonucleotide features (probes) against each transcript (probe set). This multiple testing allows for more robust background assessments and gene expression measures, and has permitted the development of many computational methods to translate image data into a single normalized "signal" for mRNA transcript abundance. There are now many pr
BACKGROUND: Though cluster analysis has become a routine analytic task for bioinformatics research, it is still arduous for researchers to assess the quality of a clustering result. To select the best clustering method and its parameters for a dataset, researchers have to run multiple clustering algorithms and compare them. However, such a comparison task with multiple clustering results is cognitively demanding and laborious. RESULTS: In this paper, we present XCluSim, a visual analytics tool t
Exploratory analysis of multidimensional data sets is challenging because of the difficulty in comprehending more than three dimensions. Two fundamental statistical principles for the exploratory analysis are (1) to examine each dimension first and then find relationships among dimensions, and (2) to try graphical displays first and then find numerical summaries [1]. We implement these principles in a novel conceptual framework called the rank-by-feature framework. In the framework, users can ch
BACKGROUND: MicroRNAs (miRNA) are short nucleotides that down-regulate its target genes. Various miRNA target prediction algorithms have used sequence complementarity between miRNA and its targets. Recently, other algorithms tried to improve sequence-based miRNA target prediction by exploiting miRNA-mRNA expression profile data. Some web-based tools are also introduced to help researchers predict targets of miRNAs from miRNA-mRNA expression profile data. A demand for a miRNA-mRNA visual analysis
Computer-based objective measurement of the ocular cyclotorsion using digital fundus photograph was developed. Color digital fundus photographs acquired with the field angle of 60 degrees , 1520 x 1080 in resolution were analyzed. Optic disc and macula were segmented by the program developed on MATLAB, which executed the serial analysis of the Otsu threshold, labeling, Canny edge. The angle between the horizontal line that bisects the optic disc and the line connecting the center of optic disc a
Data analysis and visualization is strongly influenced by noise and noise filters. There are multiple sources of "noise" in microarray data analysis, but signal/noise ratios are rarely optimized, or even considered. Here, we report a noise analysis of a novel 13 million oligonucleotide dataset - 25 human U133A (/spl sim/500,000 features) profiles of patient muscle biopsies. We use our recently described interactive visualization tool, the hierarchical clustering explorer (HCE) to systemically ad
Multidimensional data sets often include categorical information. When most dimensions have categorical information, clustering the data set as a whole can reveal interesting patterns in the data set. However, the categorical information is often more useful as a way to partition the data set: gene expression data for healthy versus diseased samples or stock performance for common, preferred, or convertible shares. We present novel ways to utilize categorical information in exploratory data anal
Research Areas
Dive deeper into Jinwook Seo's research on Nubint
Open this lab's papers in the app to read with AI, summarize, and cite in your writing.