Beat Gfeller scite author profile

We explore self-supervision as a way to learn general purpose audio representations. Specifically, we propose two self-supervised tasks: Audio2Vec, which aims at reconstructing a spectrogram slice from past and future slices and Temporal-Gap, which estimates the distance between two short audio segments extracted at random from the same audio clip. We evaluate how the representations learned via self-supervision transfer to different downstream tasks, either training a task-specific linear classifier on top of the pretrained embeddings, or fine-tuning a model end-to-end for each downstream task. Our results show that the representations learned with Audio2Vec transfer better than those learned by fully-supervised training on Audioset. In addition, by fine-tuning Audio2Vec representations it is possible to outperform fully-supervised models trained from scratch on each task, when limited data is available, thus improving label efficiency.

show abstract

A randomized distributed algorithm for the maximal independent set problem in growth-bounded graphs

Gfeller

Vicari

2007

View full text Add to dashboard Cite

SPICE: Self-Supervised Pitch Estimation

Gfeller

Frank

Roblek

et al. 2020

IEEE/ACM Trans. Audio Speech Lang. Process.

View full text Add to dashboard Cite

Towards optimal range medians

Brodal¹,

Gfeller

Jørgensen³

et al. 2011

Theoretical Computer Science

View full text Add to dashboard Cite

One-Shot Conditional Audio Filtering of Arbitrary Sounds

Gfeller

Roblek

Tagliasacchi

2021

View full text Add to dashboard Cite

We consider the problem of separating a particular sound source from a single-channel mixture, based on only a short sample of the target source (from the same recording). Using SoundFilter, a waveto-wave neural network architecture, we can train a model without using any sound class labels. Using a conditioning encoder model which is learned jointly with the source separation network, the trained model can be "configured" to filter arbitrary sound sources, even ones that it has not seen during training. Evaluated on the FSD50k dataset, our model obtains an SI-SDR improvement of 9.6 dB for mixtures of two sounds. When trained on Librispeech, our model achieves an SI-SDR improvement of 14.0 dB when separating one voice from a mixture of two speakers. Moreover, we show that the representation learned by the conditioning encoder clusters acoustically similar sounds together in the embedding space, even though it is trained without using any labels.

show abstract

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Beat Gfeller

Pre-Training Audio Representations With Self-Supervision

A randomized distributed algorithm for the maximal independent set problem in growth-bounded graphs

SPICE: Self-Supervised Pitch Estimation

Towards optimal range medians

One-Shot Conditional Audio Filtering of Arbitrary Sounds

Contact Info

Product

Resources

About