2022
Dimension independent excess risk by stochastic gradient descent
Abstract: One classical canon of statistics is that large models are prone to overfitting, and model selection procedures are necessary for high dimensional data. However, many overparameterized models, such as neural networks, perform very well in practice, although they are often trained with simple online methods and regularization. The empirical success of overparameterized models, which is often known as benign overfitting, motivates us to have a new look at the statistical generalization theory for online optimiza…
Search citation statements
Paper Sections
Select...
5
1
0
0
Citation Types
0
2
0
0
Year Published
Range
2024
2026
Publication Types
Select...
4
2
Relationship
0
6
Authors
Journals
Cited by 6 publications
(2 citation statements)
References 53 publications
0
2
0
0
Statistical Inference for Linear Functionals of Online Least-squares SGD when $t \gtrsim d^{1+δ}$
Preprint
Abstract
Smart CitationsHow this paper cites the one you are viewing
“…Developing an inferential theory for SGD becomes more challenging in particular in the growing-dimensional setting, when the number of parameters can grow with the number of iterations (or equivalently the number of observations used in online SGD). Such growing-dimensional settings are common in modern statistical machine learning problems and it well-known that online SGD has implicit regularization properties, as examined in several recent works including [1,68,74,70,19].…”
Section: Introduction
mentioning
confidence: 99%