The University of Tokyo · Medicine
Professor Yoshimasa Kawazoe's research lab specializes in medical artificial intelligence, focusing on the development and application of deep learning and natural language processing techniques for clinical data. The lab pioneers domain-specific pre-trained language models for Japanese medical text, enhances pathology image analysis using convolutional neural networks, and explores semantic interoperability in healthcare data through standards like HL7 and RDF. Their work bridges clinical informatics, AI, and electronic health records to improve diagnostic accuracy, clinical decision support, and data usability in hospital systems.
Figures are computed from collected data and may differ slightly.
The detection of objects of interest in high-resolution digital pathological images is a key part of diagnosis and is a labor-intensive task for pathologists. In this paper, we describe a Faster R-CNN-based approach for the detection of glomeruli in multistained whole slide images (WSIs) of human renal tissue sections. Faster R-CNN is a state-of-the-art general object detection method based on a convolutional neural network, which simultaneously proposes object bounds and objectness scores at ea
Generalized language models that are pre-trained with a large corpus have achieved great performance on natural language tasks. While many pre-trained transformers for English are published, few models are available for Japanese text, especially in clinical medicine. In this work, we demonstrate the development of a clinical specific BERT model with a huge amount of Japanese clinical text and evaluate it on the NTCIR-13 MedWeb that has fake Twitter messages regarding medical concerns with eight
The differences in the texture or frequency of the co-occurrence between the different features affected the CNN performance; thus, to improve the classification accuracy, methods such as segmentation are required.
The proposed methods enabled query expressions that separate knowledge resources and clinical data, thereby suggesting the feasibility for improving the usability of clinical data by enhancing the knowledge resources. We also demonstrate that when HL7 v2.5 messages are automatically converted into RDF, searches are still possible through SPARQL without modifying the structure. As such, the proposed method benefits not only our hospitals, but also numerous hospitals that handle HL7 v2.5 messages.
Abstract Generalized language models that pre-trained with a large corpus have achieved great performance on natural language tasks. While many pre-trained transformers for English are published, few models are available for Japanese text, especially in clinical medicine. In this work, we demonstrate a development of a clinical specific BERT model with a huge size of Japanese clinical narrative and evaluated it on the NTCIR-13 MedWeb that has pseudo-Twitter messages about medical concerns with e
The accuracy of the prediction model using clinical text is considered to be higher than the prediction accuracy of conventional assessments. However, our model's precision remained low at 9.3%. This may be due, in part, to the inclusion of cases in which falls did not occur because of preventative interventions during hospitalization. Nonetheless, it is estimated that interventions for cases when falls were predicted will reduce medical costs by 886 Yen/day (~US $6.50/day) of intervention, even
The histopathological findings of the glomeruli from whole slide images (WSIs) of a renal biopsy play an important role in diagnosing and grading kidney disease. This study aimed to develop an automated computational pipeline to detect glomeruli and to segment the histopathological regions inside of the glomerulus in a WSI. In order to assess the significance of this pipeline, we conducted a multivariate regression analysis to determine whether the quantified regions were associated with the pro
GBDT models for predicting BPD and mortality, designed for use within 6 h postpartum, demonstrated superior prognostic performance. SHAP value-based clustering, a data-driven approach, formed clusters of clinical relevance. These findings suggest the efficacy of a GBDT algorithm for the early postnatal prediction of BPD.
This study presents a prediction-based approach to determine thresholds for a medication alert in a computerized physician order entry. Traditional static thresholds can sometimes lead to physician's alert fatigue or overlook potentially excessive medication even if the doses are belowthe configured threshold. To address this problem, we applied a random forest algorithm to develop a prediction model for medication doses, and applied a boxplot to determine the thresholds based on the prediction
Important pieces of information related to patient symptoms and diagnosis are often written in free-text form in clinical texts. To utilize these texts, information extraction using natural language processing is required. This study evaluated the performance of named entity recognition (NER) and relation extraction (RE) using machine-learning methods. The Japanese case report corpus was used for this study, which had 113 types of entities and 36 types of relations that were manually annotated.
This study demonstrates that adverse events (AEs) extracted using natural language processing (NLP) from clinical texts reflect the known frequencies of AEs associated with anticancer drugs. Using data from 44,502 cancer patients at a single hospital, we identified cases prescribed anticancer drugs (platinum, PLT; taxane, TAX; pyrimidine, PYA) and compared them to non-treatment (NTx) group using propensity score matching. Over 365 days, AEs (peripheral neuropathy, PN; oral mucositis, OM; taste a
In this retrospective observational study, we aimed to investigate the potential of natural language processing (NLP) for drug repositioning by analyzing the preventive effects of cardioprotective drugs against anthracycline-induced cardiotoxicity (AIC) using electronic medical records. We evaluated the effects of angiotensin II receptor blockers/angiotensin-converting enzyme inhibitors (ARB/ACEIs), beta-blockers (BBs), statins, and calcium channel blockers (CCBs) on AIC using signals extracted
Using quantitative data, we were able to reveal how counterrumors influence the vaccination status of the COVID-19 vaccine. We think that our findings would be a foundation for considering countermeasures of vaccination.
• In medical texts, 18.6 % of patient condition phrases include unrelated information. • The accuracy of traditional NER and EL methods is limited. • We identified the types of entities that current extraction techniques often miss. We investigated the limitations of conventional named entity recognition (NER) and entity linking (EL) methods in accurately extracting patient condition information from medical texts, focusing on the challenges posed by non-contiguous spans and the potential inform
Open papers in the app to read, cite, and organize with AI.