The University of Tokyo · 컴퓨터과학
스테파노 마사롤리 교수의 연구실은 지속적 깊이(differential equation 기반) 신경망과 동역학 시스템 이론을 융합한 연속적 딥러닝 아키텍처를 중심으로 연구를 전개하고 있습니다. 특히 신경미분방정식(Neural ODEs)의 안정성 보장, 에너지 기반 제어 이론과의 융합, 최적 제어 프레임워크를 통한 파라미터 최적화 등, 이론적 안정성과 실용적 성능을 동시에 확보하는 데 초점을 맞추고 있습니다. 연구는 수학적 이론과 물리 기반 모델링을 기반으로 하여, 딥러닝의 '흑자'를 투명하게 드러내고자 합니다.
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
We introduce optimal energy shaping as an enhancement of classical passivity-based control methods. A promising feature of passivity theory, alongside stability, has traditionally claimed to be intuitive performance tuning along the execution of a given task. However, a systematic approach for adjusting performance within a passive control framework has yet to be developed, as each method relies on few and problem-specific practical insights. Here, we cast the classic energy-shaping control desi
Continuous deep learning architectures have recently re-emerged as Neural Ordinary Differential Equations (Neural ODEs). This infinite-depth approach theoretically bridges the gap between deep learning and dynamical systems, offering a novel perspective. However, deciphering the inner working of these models is still an challenge, as most applications apply them as generic black-box modules. In this work we open the box, further developing the continuous-depth formulation with the aim of clarify
We introduce a provably stable variant of neural ordinary differential equations (neural ODEs) whose trajectories evolve on an energy functional parametrised by a neural network. Stable neural flows provide an implicit guarantee on asymptotic stability of the depth-flows, leading to robustness against input perturbations and low computational burden for the numerical solver. The learning procedure is cast as an optimal control problem, and an approximate solution is proposed based on adjoint sen
Neural networks are discrete entities: subdivided into discrete layers and parametrized by weights which are iteratively optimized via difference equations. Recent work proposes networks with layer outputs which are no longer quantized but are solutions of an ordinary differential equation (ODE); however, these networks are still optimized via discrete methods (e.g. gradient descent). In this paper, we explore a different direction: namely, we propose a novel framework for learning in which the
This paper presents a novel identification procedure for a class of hybrid dynamical systems. In particular, we consider hybrid dynamical systems which are single flowed and single jumped and whose flow and jump maps linearly depend on two sets of unknown parameters. A systematic way to determine whether the system is flowing or jumping is introduced and used to identify the unknown parameters by employing a linear recursive estimator. Simulations have been performed to prove the validity of the
Recent advances in attention-free sequence models rely on convolutions as alternatives to the attention operator at the core of Transformers. In particular, long convolution sequence models have achieved state-of-the-art performance in many domains, but incur a significant cost during auto-regressive inference workloads -- naively requiring a full pass (or caching of activations) over the input sequence for each generated token -- similarly to attention-based models. In this paper, we seek to en
Neural networks are discrete entities: subdivided into discrete layers and parametrized by weights which are iteratively optimized via difference equations. Recent work proposes networks with layer outputs which are no longer quantized but are solutions of an ordinary differential equation (ODE); however, these networks are still optimized via discrete methods (e.g. gradient descent). In this paper, we explore a different direction: namely, we propose a novel framework for learning in which the
This paper presents a novel identification procedure. The proposed method consists in a recursive formulation of the algebraic Frisch scheme, which is an estimation procedure based on mild a priori assumptions in the context of error-in-variables (EIV) schemes. Simulations have been performed to show the validity of the new methodology.
This paper proposes an innovative identification scheme to estimate parameters constituting linear relations in time-invariant systems: the bounding box recursive Frisch scheme. A novel recursive version of the Frisch scheme, a linear estimator characterised by mild prior assumptions in the error-in-variables (EIV) framework, has been derived. The fast computational time and convergence in the identification of linear systems are the most relevant feature of this recursive version of the scheme.
Scientists have long been attracted to mechanisms surrounding the predator-prey system. The Lotka-Volterra (LV) model is the most popular formalism used to investigate the dynamics of this system. LV equations present non-linear dynamics that exhibit periodic oscillations in both prey and predator populations. In practical situations, it is useful to stabilise the system asymptotically to a desired set point (population) wherein the two species coexist by fashioning specific control actions. Thi
We detail a novel class of implicit neural models. Leveraging time-parallel methods for differential equations, Multiple Shooting Layers (MSLs) seek solutions of initial value problems via parallelizable root-finding algorithms. MSLs broadly serve as drop-in replacements for neural ordinary differential equations (Neural ODEs) with improved efficiency in number of function evaluations (NFEs) and wall-clock inference time. We develop the algorithmic framework of MSLs, analyzing the different choi
We detail a novel class of implicit neural models. Leveraging time-parallel methods for differential equations, Multiple Shooting Layers (MSLs) seek solutions of initial value problems via parallelizable root-finding algorithms. MSLs broadly serve as drop-in replacements for neural ordinary differential equations (Neural ODEs) with improved efficiency in number of function evaluations (NFEs) and wall-clock inference time. We develop the algorithmic framework of MSLs, analyzing the different choi
This paper presents a novel control strategy for stable linear time–invariant systems operating with a finite number of set points. Inspired by the theory of passivity-based control, the proposed method aims at simultaneously and asymptotically stabilize all the desired working modes by means of a static nonlinear state feedback law. An asynchronous external signal is then employed to trigger a hybrid controller in order to switch between the different working modes. The proposed approach is val