2022
DOI: 10.1214/22-ejs2055
|Get access via publisher |Summarize |Cite
|
Sign up to set email alerts

Dimension independent excess risk by stochastic gradient descent

Abstract: One classical canon of statistics is that large models are prone to overfitting, and model selection procedures are necessary for high dimensional data. However, many overparameterized models, such as neural networks, perform very well in practice, although they are often trained with simple online methods and regularization. The empirical success of overparameterized models, which is often known as benign overfitting, motivates us to have a new look at the statistical generalization theory for online optimiza… Show more

Search citation statements

Order By: Relevance

Paper Sections

Select...
5
1
0
0

Citation Types

0
2
0
0

Year Published

Range
2024
2024
2026
2026

Publication Types

Select...
4
2

Relationship

0
6

Authors

Journals

citations

Cited by 6 publications

(2 citation statements)
references

References 53 publications

0
2
0
0
Order By: Relevance
How this paper cites the one you are viewing
“…Developing an inferential theory for SGD becomes more challenging in particular in the growing-dimensional setting, when the number of parameters can grow with the number of iterations (or equivalently the number of observations used in online SGD). Such growing-dimensional settings are common in modern statistical machine learning problems and it well-known that online SGD has implicit regularization properties, as examined in several recent works including [1,68,74,70,19].…”
Section: Introduction
mentioning
confidence: 99%
“…Beyond providing a theoretical framework for growing-dimensional inference, our results have practical implications for constructing algorithmic prediction intervals in linear regression. For a new test point a, independent of the training data used by SGD, choosing a in (2) directly yields a predictive confidence interval, complementing prior works on implicit regularization and benign overfitting [1,68,74,70,19]. Furthermore, our results can be used to develop algorithmic Wald-type tests for feature significance in highdimensional linear models-an essential tool in empirical sciences such as biology, social science, economics, and medicine [27,64,13].…”
Section: Introduction
mentioning
confidence: 99%
See 1 more Smart Citation