Computer Science > Machine Learning

[Submitted on 11 Oct 2021]

Imitating Deep Learning Dynamics via Locally Elastic Stochastic Differential Equations

Understanding the training dynamics of deep learning models is perhaps a necessary step toward demystifying the effectiveness of these models. In particular, how do data from different classes gradually become separable in their feature spaces when training neural networks using stochastic gradient descent? In this study, we model the evolution of features during deep learning training using a set of stochastic differential equations (SDEs) that each corresponds to a training sample. As a crucial ingredient in our modeling strategy, each SDE contains a drift term that reflects the impact of backpropagation at an input on the features of all samples. Our main finding uncovers a sharp phase transition phenomenon regarding the {intra-class impact: if the SDEs are locally elastic in the sense that the impact is more significant on samples from the same class as the input, the features of the training data become linearly separable, meaning vanishing training loss; otherwise, the features are not separable, regardless of how long the training time is. Moreover, in the presence of local elasticity, an analysis of our SDEs shows that the emergence of a simple geometric structure called the neural collapse of the features. Taken together, our results shed light on the decisive role of local elasticity in the training dynamics of neural networks. We corroborate our theoretical analysis with experiments on a synthesized dataset of geometric shapes and CIFAR-10.

Comments: Accepted to NeurIPS 2021 Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML) Cite as: arXiv:2110.05960 [cs.LG] (or arXiv:2110.05960v1 [cs.LG] for this version)

Submission history

From: Jiayao Zhang [view email]
[v1] Mon, 11 Oct 2021 17:17:20 UTC (29,186 KB)

[2110.05960] Imitating Deep Learning Dynamics via Locally Elastic Stochastic Dif...

Computer Science > Machine Learning

Imitating Deep Learning Dynamics via Locally Elastic Stochastic Differential Equations

Submission history

Recommend

低复杂度 - 服务网格的下一站

Open Source J2EE Frameworks

Ben Gottlieb - Swift Mentor - Codementor

Improve Application Performance with These Advanced GC Techniques

修炼内功，敏捷应变 | IDCF·云智慧DevOps黑客马拉松挑战赛（第六期）精彩回顾

Domain Driven Design Cheat Sheet | OverAPI.com

SpringBoot 整合 Spring Data JPA

All About CSS BEM!

Hosting & Servers — Heroku, Now, Galaxy, Digital Ocean, Linode, Docker, Netl...

Android入门教程 | 广播机制 Broadcast

About Joyk