名古屋大学 · 工学
Ana Davila教授の研究室は、医療画像と機械学習の融合を柱に、手術支援ロボットの制御技術や、医療画像分野における転移学習の高度な応用を研究しています。特に、手術画像のドメインギャップに起因する性能低下を解消するための新規微調整手法や、音声インターフェースを活用した医療ロボットのインタラクティブ制御技術の開発が特徴です。リアルタイム性と安全性を両立するインバースキネマティクスの高速解法や、進化的最適化を用いた柔軟なファインチューニング戦略の構築にも取り組んでいます。
Figures are computed from collected data and may differ slightly.
In the context of medical imaging and machine learning, one of the most pressing challenges is the effective adaptation of pre-trained models to specialized medical contexts. Despite the availability of advanced pre-trained models, their direct application to the highly specialized and diverse field of medical imaging often falls short due to the unique characteristics of medical data. This study provides a comprehensive analysis on the performance of various fine-tuning methods applied to pre-t
Supplementary data are available at <i>Bioinformatics Advances</i> online.
Abstract Robotic manipulation in surgical applications often demands the surgical instrument to pivot around a fixed point, known as remote center of motion (RCM). The RCM constraint ensures that the pivot point of the surgical tool remains stationary at the incision port, preventing tissue damage and bleeding. Precisely and efficiently controlling tool positioning and orientation under this constraint poses a complex Inverse Kinematics (IK) problem that must be solved in real-time to ensure pat
Traditional control interfaces for robotic-assisted minimally invasive surgery impose a significant cognitive load on surgeons. To improve surgical efficiency, surgeon-robot collaboration capabilities, and reduce surgeon burden, we present a novel voice control interface for surgical robotic assistants. Our system integrates Whisper, state-of-the-art speech recognition, within the ROS framework to enable real-time interpretation and execution of voice commands for surgical manipulator control. T
Transfer learning is a widely used technique to leverage pre-trained models on new tasks, but it often suffers from out-of-distribution shifts when the source and target domains are different. This is especially common in surgical images, where the appearance and context of the images vary significantly across different procedures and instruments. To address this problem, we propose a novel gradient-based fine-tuning strategy that selectively freezes layers of a pre-trained model based on their
Ambiguity in natural language instructions poses significant risks in safety-critical human-robot interaction, particularly in domains such as surgery. To address this, we propose a framework that uses Large Language Models (LLMs) for ambiguity detection specifically designed for collaborative surgical scenarios. Our method employs an ensemble of LLM evaluators, each configured with distinct prompting techniques to identify linguistic, contextual, procedural, and critical ambiguities. A chain-of
Deep learning has significantly advanced image analysis across diverse domains but often depends on large, annotated datasets for success. Transfer learning addresses this challenge by utilizing pre-trained models to tackle new tasks with limited labeled data. However, discrepancies between source and target domains can hinder effective transfer learning. We introduce BioTune, a novel adaptive fine-tuning technique utilizing evolutionary optimization. BioTune enhances transfer learning by optima
ABSTRACT Minimally invasive surgery can benefit significantly from automated surgical tool detection, enabling advanced analysis and assistance. However, the limited availability of annotated data in surgical settings poses a challenge for training robust deep learning models. This paper introduces a novel staged adaptive fine‐tuning approach consisting of two steps: a linear probing stage to condition additional classification layers on a pre‐trained CNN‐based architecture and a gradual freezin
Large Language Models (LLMs) have demonstrated significant capabilities in understanding and generating human language, contributing to more natural interactions with complex systems. However, they face challenges such as ambiguity in user requests processed by LLMs. To address these challenges, this paper introduces and evaluates a multi-agent debate framework designed to enhance detection and resolution capabilities beyond single models. The framework consists of three LLM architectures (Llama
Traditional control interfaces for robotic-assisted minimally invasive surgery impose a significant cognitive load on surgeons. To improve surgical efficiency, surgeon-robot collaboration capabilities, and reduce surgeon burden, we present a novel voice control interface for surgical robotic assistants. Our system integrates Whisper, state-of-the-art speech recognition, within the ROS framework to enable real-time interpretation and execution of voice commands for surgical manipulator control. T
Ambiguity in natural language instructions poses significant risks in safety-critical human-robot interaction, particularly in domains such as surgery. To address this, we propose a framework that uses Large Language Models (LLMs) for ambiguity detection specifically designed for collaborative surgical scenarios. Our method employs an ensemble of LLM evaluators, each configured with distinct prompting techniques to identify linguistic, contextual, procedural, and critical ambiguities. A chain-of
Effective human-robot collaboration in surgery is affected by the inherent ambiguity of verbal communication. This paper presents a framework for a robotic surgical assistant that interprets and disambiguates verbal instructions from a surgeon by grounding them in the visual context of the operating field. The system employs a two-level affordance-based reasoning process that first analyzes the surgical scene using a multimodal vision-language model and then reasons about the instruction using a
Large Language Models (LLMs) have demonstrated significant capabilities in understanding and generating human language, contributing to more natural interactions with complex systems. However, they face challenges such as ambiguity in user requests processed by LLMs. To address these challenges, this paper introduces and evaluates a multi-agent debate framework designed to enhance detection and resolution capabilities beyond single models. The framework consists of three LLM architectures (Llama
Open papers in the app to read, cite, and organize with AI.