[Submitted on 26 Nov 2023 (v1), last revised 22 Jul 2026 (this version, v3)]
Abstract:We study an idealized training process for deep neural networks in a continuous-depth, mean-field model in which each layer is parameterized by a probability measure on a Euclidean parameter space. The training dynamics are formulated as a Wasserstein-type gradient flow of an objective with a fixed $L^2$-regularization. Under suitable analyticity and growth assumptions, together with a coercivity assumption and sufficient regularity of the initial data, we prove that every curve of maximal slope converges to a single critical point of the objective as the training time tends to infinity. The proof combines compactness of the curve with a Łojasiewicz--Simon inequality for the metric slope. To establish the inequality, we lift the objective to a Hilbert space of random variables and use the analyticity of the lifted gradient in a stronger $L^\infty$ topology to overcome its lack of continuous differentiability in the Hilbert-space topology. Our convergence result does not require global displacement convexity, a Polyak--Łojasiewicz-type condition, or initialization near a minimizer; the objective may remain genuinely nonconvex.
Submission history
From: Noboru Isobe [view email]
[v1]
Sun, 26 Nov 2023 17:44:29 UTC (113 KB)
[v2]
Sun, 14 Apr 2024 05:39:11 UTC (113 KB)
[v3]
Wed, 22 Jul 2026 02:39:17 UTC (105 KB)
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.