Adversarial Semi-Supervised Audio Source Separation Applied to Singing Voice Extraction

Stoller, Daniel; Ewert, Sebastian; Dixon, Simon

doi:10.1109/icassp.2018.8461722

Cited by 63 publications

(55 citation statements)

References 16 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…As a baseline system for SVS, we implemented a variant of the U-Net described in Section 3.2 and shown in Figure 1. The approach is similar to [19] and [11] and outputs a mask when given spectrogram magnitudes of a mixture excerpt. During training, audio excerpts are randomly selected from the multi-track dataset, and converted to a log-normalised spectrogram representation.…”

Section: Proposed Approachesmentioning

confidence: 99%

Jointly Detecting and Separating Singing Voice: A Multi-Task Approach

Stoller

Ewert²,

Dixon

2018

Latent Variable Analysis and Signal Separation

Self Cite

View full text Add to dashboard Cite

A main challenge in applying deep learning to music processing is the availability of training data. One potential solution is Multi-task Learning, in which the model also learns to solve related auxiliary tasks on additional datasets to exploit their correlation. While intuitive in principle, it can be challenging to identify related tasks and construct the model to optimally share information between tasks. In this paper, we explore vocal activity detection as an additional task to stabilise and improve the performance of vocal separation. Further, we identify problematic biases specific to each dataset that could limit the generalisation capability of separation and detection models, to which our proposed approach is robust. Experiments show improved performance in separation as well as vocal detection compared to single-task baselines. However, we find that the commonly used Signal-to-Distortion Ratio (SDR) metrics did not capture the improvement on non-vocal sections, indicating the need for improved evaluation methodologies.

show abstract

Section: Proposed Approachesmentioning

confidence: 99%

Jointly Detecting and Separating Singing Voice: A Multi-Task Approach

Stoller

Ewert²,

Dixon

2018

Latent Variable Analysis and Signal Separation

Self Cite

View full text Add to dashboard Cite

show abstract

“…As in the unconditional setting, the discriminator attempts to differentiate between real and generated images. Adversarial training was used for supervised source separation, where the distribution of each of the mixture components is known and modeled by a GAN, by Stoller et al [11] and Subkhan et al [12]. The adversarial training was motivated as being better able to deal with correlated sources.…”

Section: Related Workmentioning

confidence: 99%

Semi-supervised Monaural Singing Voice Separation with a Masking Network Trained on Synthetic Mixtures

Michelashvili

Benaim

Wolf

2019

ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

View full text Add to dashboard Cite

We study the problem of semi-supervised singing voice separation, in which the training data contains a set of samples of mixed music (singing and instrumental) and an unmatched set of instrumental music. Our solution employs a single mapping function g, which, applied to a mixed sample, recovers the underlying instrumental music, and, applied to an instrumental sample, returns the same sample. The network g is trained using purely instrumental samples, as well as on synthetic mixed samples that are created by mixing reconstructed singing voices with random instrumental samples. Our results indicate that we are on a par with or better than fully supervised methods, which are also provided with training samples of unmixed singing voices, and are better than other recent semi-supervised methods.

show abstract

“…The generation of sequential samples, however, heavily relies on the context information [24]. Albeit a handful of related studies reported in the audio processing domain, they either focus on speech enhancement [29], [30] or music creation [31].…”

Section: Introductionmentioning

confidence: 99%

Snore-GANs: Improving Automatic Snore Sound Classification With Synthesized Data

Zhang

Han

Qian

et al. 2020

IEEE J. Biomed. Health Inform.

View full text Add to dashboard Cite

One of the frontier issues that severely hamper the development of automatic snore sound classification (ASSC) associates to the lack of sufficient supervised training data. To cope with this problem, we propose a novel data augmentation approach based on semi-supervised conditional Generative Adversarial Networks (scGANs), which aims to automatically learn a mapping strategy from a random noise space to original data distribution. The proposed approach has the capability of well synthesizing 'realistic' high-dimensional data, while requiring no additional annotation process. To handle the mode collapse problem of GANs, we further introduce an ensemble strategy to enhance the diversity of the generated data. The systematic experiments conducted on a widely used Munich-Passau snore sound corpus demonstrate that the scGANs-based systems can remarkably outperform other classic data augmentation systems, and are also competitive to other recently reported systems for ASSC.

show abstract

Adversarial Semi-Supervised Audio Source Separation Applied to Singing Voice Extraction

Cited by 63 publications

References 16 publications

Jointly Detecting and Separating Singing Voice: A Multi-Task Approach

Jointly Detecting and Separating Singing Voice: A Multi-Task Approach

Semi-supervised Monaural Singing Voice Separation with a Masking Network Trained on Synthetic Mixtures

Snore-GANs: Improving Automatic Snore Sound Classification With Synthesized Data

Contact Info

Product

Resources

About