Stability and Generalization Analysis of Gradient Methods for Shallow Neural Networks

Lei, Yunwen; Jin, Rong; Ying, Yiming

File Download

There are no files associated with this item.

Links for fulltext

(May Require Subscription)

Scopus: eid_2-s2.0-85148766126
Find via

Supplementary

Citations:
- Scopus: 0
Appears in Collections:
- Mathematics: Conference papers

Conference Paper: Stability and Generalization Analysis of Gradient Methods for Shallow Neural Networks

Title	Stability and Generalization Analysis of Gradient Methods for Shallow Neural Networks
Authors	Lei, Yunwen Jin, Rong Ying, Yiming
Issue Date	2022
Citation	Advances in Neural Information Processing Systems, 2022, v. 35 How to Cite?
Abstract	While significant theoretical progress has been achieved, unveiling the generalization mystery of overparameterized neural networks still remains largely elusive. In this paper, we study the generalization behavior of shallow neural networks (SNNs) by leveraging the concept of algorithmic stability. We consider gradient descent (GD) and stochastic gradient descent (SGD) to train SNNs, for both of which we develop consistent excess risk bounds by balancing the optimization and generalization via early-stopping. As compared to existing analysis on GD, our new analysis requires a relaxed overparameterization assumption and also applies to SGD. The key for the improvement is a better estimation of the smallest eigenvalues of the Hessian matrices of the empirical risks and the loss function along the trajectories of GD and SGD by providing a refined estimation of their iterates.
Persistent Identifier	http://hdl.handle.net/10722/329927
ISSN	1049-5258 2020 SCImago Journal Rankings: 1.399

DC Field	Value	Language
dc.contributor.author	Lei, Yunwen	-
dc.contributor.author	Jin, Rong	-
dc.contributor.author	Ying, Yiming	-
dc.date.accessioned	2023-08-09T03:36:30Z	-
dc.date.available	2023-08-09T03:36:30Z	-
dc.date.issued	2022	-
dc.identifier.citation	Advances in Neural Information Processing Systems, 2022, v. 35	-
dc.identifier.issn	1049-5258	-
dc.identifier.uri	http://hdl.handle.net/10722/329927	-
dc.description.abstract	While significant theoretical progress has been achieved, unveiling the generalization mystery of overparameterized neural networks still remains largely elusive. In this paper, we study the generalization behavior of shallow neural networks (SNNs) by leveraging the concept of algorithmic stability. We consider gradient descent (GD) and stochastic gradient descent (SGD) to train SNNs, for both of which we develop consistent excess risk bounds by balancing the optimization and generalization via early-stopping. As compared to existing analysis on GD, our new analysis requires a relaxed overparameterization assumption and also applies to SGD. The key for the improvement is a better estimation of the smallest eigenvalues of the Hessian matrices of the empirical risks and the loss function along the trajectories of GD and SGD by providing a refined estimation of their iterates.	-
dc.language	eng	-
dc.relation.ispartof	Advances in Neural Information Processing Systems	-
dc.title	Stability and Generalization Analysis of Gradient Methods for Shallow Neural Networks	-
dc.type	Conference_Paper	-
dc.description.nature	link_to_subscribed_fulltext	-
dc.identifier.scopus	eid_2-s2.0-85148766126	-
dc.identifier.volume	35	-

File Download

Links for fulltext

(May Require Subscription)

Supplementary

Conference Paper: Stability and Generalization Analysis of Gradient Methods for Shallow Neural Networks

Export via OAI-PMH Interface in XML Formats

OR

Export to Other Non-XML Formats