Sajjad Abdoli scite author profile

In this paper, we present an end-to-end approach for environmental sound classification based on a 1D Convolution Neural Network (CNN) that learns a representation directly from the audio signal. Several convolutional layers are used to capture the signal's fine time structure and learn diverse filters that are relevant to the classification task. The proposed approach can deal with audio signals of any length as it splits the signal into overlapped frames using a sliding window. Different architectures considering several input sizes are evaluated, including the initialization of the first convolutional layer with a Gammatone filterbank that models the human auditory filter response in the cochlea. The performance of the proposed end-to-end approach in classifying environmental sounds was assessed on the UrbanSound8k dataset and the experimental results have shown that it achieves 89% of mean accuracy. Therefore, the propose approach outperforms most of the state-of-the-art approaches that use handcrafted features or 2D representations as input. Furthermore, the proposed approach has a small number of parameters compared to other architectures found in the literature, which reduces the amount of data required for training.

show abstract

End-to-End Environmental Sound Classification using a 1D Convolutional Neural Network

Abdoli¹,

Cardinal²,

Koerich³

2019

Preprint

View full text Add to dashboard Cite

Speaker Detection in the Wild: Lessons Learned from JSALT 2019

García¹,

Villalba²,

Bredin³

et al. 2020

View full text Add to dashboard Cite

Speaker detection in the wild: Lessons learned from JSALT 2019

Garcia¹,

Villalba²,

Bredin³

et al. 2019

Preprint

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Sajjad Abdoli

End-to-end environmental sound classification using a 1D convolutional neural network

End-to-End Environmental Sound Classification using a 1D Convolutional Neural Network

Speaker Detection in the Wild: Lessons Learned from JSALT 2019

Speaker detection in the wild: Lessons learned from JSALT 2019

Contact Info

Product

Resources

About