東京大学 · 情報科学
Stefano Massaroli教授の研究室では、深層学習と力学系理論の融合を軸に、連続的深層ネットワーク(Neural ODE)の理論的基盤と実装の最適化に取り組んでいます。特に、エネルギー関数に基づく制御理論を応用し、安定性と性能チューニングを一元的に扱える最適制御フレームワークの構築を進めています。また、学習プロセス自体を微分方程式として定式化することで、計算効率と数値的安定性を両立する新規なモデル設計を提案しています。
Figures are computed from collected data and may differ slightly.
We introduce optimal energy shaping as an enhancement of classical passivity-based control methods. A promising feature of passivity theory, alongside stability, has traditionally claimed to be intuitive performance tuning along the execution of a given task. However, a systematic approach for adjusting performance within a passive control framework has yet to be developed, as each method relies on few and problem-specific practical insights. Here, we cast the classic energy-shaping control desi
Continuous deep learning architectures have recently re-emerged as Neural Ordinary Differential Equations (Neural ODEs). This infinite-depth approach theoretically bridges the gap between deep learning and dynamical systems, offering a novel perspective. However, deciphering the inner working of these models is still an challenge, as most applications apply them as generic black-box modules. In this work we open the box, further developing the continuous-depth formulation with the aim of clarify
We introduce a provably stable variant of neural ordinary differential equations (neural ODEs) whose trajectories evolve on an energy functional parametrised by a neural network. Stable neural flows provide an implicit guarantee on asymptotic stability of the depth-flows, leading to robustness against input perturbations and low computational burden for the numerical solver. The learning procedure is cast as an optimal control problem, and an approximate solution is proposed based on adjoint sen
Neural networks are discrete entities: subdivided into discrete layers and parametrized by weights which are iteratively optimized via difference equations. Recent work proposes networks with layer outputs which are no longer quantized but are solutions of an ordinary differential equation (ODE); however, these networks are still optimized via discrete methods (e.g. gradient descent). In this paper, we explore a different direction: namely, we propose a novel framework for learning in which the
This paper presents a novel identification procedure for a class of hybrid dynamical systems. In particular, we consider hybrid dynamical systems which are single flowed and single jumped and whose flow and jump maps linearly depend on two sets of unknown parameters. A systematic way to determine whether the system is flowing or jumping is introduced and used to identify the unknown parameters by employing a linear recursive estimator. Simulations have been performed to prove the validity of the
Recent advances in attention-free sequence models rely on convolutions as alternatives to the attention operator at the core of Transformers. In particular, long convolution sequence models have achieved state-of-the-art performance in many domains, but incur a significant cost during auto-regressive inference workloads -- naively requiring a full pass (or caching of activations) over the input sequence for each generated token -- similarly to attention-based models. In this paper, we seek to en
Neural networks are discrete entities: subdivided into discrete layers and parametrized by weights which are iteratively optimized via difference equations. Recent work proposes networks with layer outputs which are no longer quantized but are solutions of an ordinary differential equation (ODE); however, these networks are still optimized via discrete methods (e.g. gradient descent). In this paper, we explore a different direction: namely, we propose a novel framework for learning in which the
This paper presents a novel identification procedure. The proposed method consists in a recursive formulation of the algebraic Frisch scheme, which is an estimation procedure based on mild a priori assumptions in the context of error-in-variables (EIV) schemes. Simulations have been performed to show the validity of the new methodology.
This paper proposes an innovative identification scheme to estimate parameters constituting linear relations in time-invariant systems: the bounding box recursive Frisch scheme. A novel recursive version of the Frisch scheme, a linear estimator characterised by mild prior assumptions in the error-in-variables (EIV) framework, has been derived. The fast computational time and convergence in the identification of linear systems are the most relevant feature of this recursive version of the scheme.
Scientists have long been attracted to mechanisms surrounding the predator-prey system. The Lotka-Volterra (LV) model is the most popular formalism used to investigate the dynamics of this system. LV equations present non-linear dynamics that exhibit periodic oscillations in both prey and predator populations. In practical situations, it is useful to stabilise the system asymptotically to a desired set point (population) wherein the two species coexist by fashioning specific control actions. Thi
We detail a novel class of implicit neural models. Leveraging time-parallel methods for differential equations, Multiple Shooting Layers (MSLs) seek solutions of initial value problems via parallelizable root-finding algorithms. MSLs broadly serve as drop-in replacements for neural ordinary differential equations (Neural ODEs) with improved efficiency in number of function evaluations (NFEs) and wall-clock inference time. We develop the algorithmic framework of MSLs, analyzing the different choi
We detail a novel class of implicit neural models. Leveraging time-parallel methods for differential equations, Multiple Shooting Layers (MSLs) seek solutions of initial value problems via parallelizable root-finding algorithms. MSLs broadly serve as drop-in replacements for neural ordinary differential equations (Neural ODEs) with improved efficiency in number of function evaluations (NFEs) and wall-clock inference time. We develop the algorithmic framework of MSLs, analyzing the different choi
This paper presents a novel control strategy for stable linear time–invariant systems operating with a finite number of set points. Inspired by the theory of passivity-based control, the proposed method aims at simultaneously and asymptotically stabilize all the desired working modes by means of a static nonlinear state feedback law. An asynchronous external signal is then employed to trigger a hybrid controller in order to switch between the different working modes. The proposed approach is val
Open papers in the app to read, cite, and organize with AI.