Pierre Wolinski scite author profile

Pierre Wolinski

3Publications

10Citation Statements Received

23Citation Statements Given

How they've been cited

How they cite others

Affiliations

University of Paris-Sud, Institut Polytechnique de Paris

Publications

Order By: Most citations

Learning with Random Learning Rates

Blier

Wolinski

Ollivier³

2020

View full text Add to dashboard Cite

In neural networks, the learning rate of the gradient descent strongly affects performance. This prevents reliable out-of-the-box training of a model on a new problem. We propose the All Learning Rates At Once (Alrao) algorithm: each unit or feature in the network gets its own learning rate sampled from a random distribution spanning several orders of magnitude, in the hope that enough units will get a close-to-optimal learning rate. Perhaps surprisingly, stochastic gradient descent (SGD) with Alrao performs close to SGD with an optimally tuned learning rate, for various network architectures and problems. In our experiments, all Alrao runs were able to learn well without any tuning.

show abstract

Learning with Random Learning Rates

Blier¹,

Wolinski²,

Ollivier³

2018

Preprint

View full text Add to dashboard Cite

Imposing Gaussian Pre-Activations in a Neural Network

Wolinski¹,

Arbel²

2022

Preprint

View full text Add to dashboard Cite

The goal of the present work is to propose a way to modify both the initialization distribution of the weights of a neural network and its activation function, such that all pre-activations are Gaussian. We propose a family of pairs initialization/activation, where the activation functions span a continuum from bounded functions (such as Heaviside or tanh) to the identity function.This work is motivated by the contradiction between existing works dealing with Gaussian pre-activations: on one side, the works in the line of the Neural Tangent Kernels and the Edge of Chaos are assuming it, while on the other side, theoretical and experimental results challenge this hypothesis.The family of pairs initialization/activation we are proposing will help us to answer this hot question: is it desirable to have Gaussian pre-activations in a neural network?

show abstract

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Pierre Wolinski

Learning with Random Learning Rates

Learning with Random Learning Rates

Imposing Gaussian Pre-Activations in a Neural Network

Contact Info

Product

Resources

About