Receptive Field Regularization Techniques for Audio Classification and Tagging with Deep Convolutional Neural Networks

Koutini, Khaled; Eghbal-zadeh, Hamid; Widmer, Gerhard

doi:10.1109/TASLP.2021.3082307

Computer Science > Sound

arXiv:2105.12395 (cs)

[Submitted on 26 May 2021]

Title:Receptive Field Regularization Techniques for Audio Classification and Tagging with Deep Convolutional Neural Networks

Authors:Khaled Koutini, Hamid Eghbal-zadeh, Gerhard Widmer

View PDF

Abstract:In this paper, we study the performance of variants of well-known Convolutional Neural Network (CNN) architectures on different audio tasks. We show that tuning the Receptive Field (RF) of CNNs is crucial to their generalization. An insufficient RF limits the CNN's ability to fit the training data. In contrast, CNNs with an excessive RF tend to over-fit the training data and fail to generalize to unseen testing data. As state-of-the-art CNN architectures-in computer vision and other domains-tend to go deeper in terms of number of layers, their RF size increases and therefore they degrade in performance in several audio classification and tagging tasks. We study well-known CNN architectures and how their building blocks affect their receptive field. We propose several systematic approaches to control the RF of CNNs and systematically test the resulting architectures on different audio classification and tagging tasks and datasets. The experiments show that regularizing the RF of CNNs using our proposed approaches can drastically improve the generalization of models, out-performing complex architectures and pre-trained models on larger datasets. The proposed CNNs achieve state-of-the-art results in multiple tasks, from acoustic scene classification to emotion and theme detection in music to instrument recognition, as demonstrated by top ranks in several pertinent challenges (DCASE, MediaEval).

Comments:	Accepted in IEEE/ACM Transactions on Audio, Speech, and Language Processing. Code available: this https URL
Subjects:	Sound (cs.SD); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2105.12395 [cs.SD]
	(or arXiv:2105.12395v1 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2105.12395
Related DOI:	https://doi.org/10.1109/TASLP.2021.3082307

Submission history

From: Khaled Koutini [view email]
[v1] Wed, 26 May 2021 08:36:29 UTC (6,700 KB)

Computer Science > Sound

Title:Receptive Field Regularization Techniques for Audio Classification and Tagging with Deep Convolutional Neural Networks

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:Receptive Field Regularization Techniques for Audio Classification and Tagging with Deep Convolutional Neural Networks

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators