CrowdioSet and PaRIRset: Two Datasets Towards Live Music Source Separation

1Music Technology Group, Universitat Pompeu Fabra, Barcelona, Spain
2Eurecat, Centre Tecnològic de Catalunya, Barcelona, Spain

Abstract

Most Music Source Separation (MSS) models do not generalize well to live music recordings because they are trained on studio recordings alone, disregarding the venue acoustics, the speaker system's response and audience noise. We propose to bridge this gap by providing and training a model on two novel datasets.

  1. We present CrowdioSet: a noise dataset comprising 4800 real ambience tracks from Freesound and synthetic sing-alongs for the vocals in MUSDB18 and MOISESDB datasets, generated from zero-shot singing voice conversions. CrowdioSet enables effective audio denoising for live recordings, resulting in superior separation both in objective and subjective evaluations.
  2. We introduce PaRIRset: a stereo impulse response dataset captured across 40 professional concert venues using a microphone array. Our results show that adding PaRIRset RIRs increases the performance of a MSS model compared to using real RIRs from Speech Enhancement tasks alone. We make the examples, code, model weights, PaRIRset, and CrowdioSet freely available to the public.
proposed pipeline

(a) Baseline vs (b) Ours

CrowdioSet

We have synthesized preliminary sing-alongs for every one of the 384 vocals stems in MUSDB18HQ and MOISESDB by combining two techniques. First, we have applied the Antares AVOX Choir plugin (an effect that combines vibrato, detuning and delay) in its 32-voice configuration to each vocals stem.

Second, we have sourced a pool of 200 vocal samples from Freesound using the queries a cappella and vocals, and used each to convert every vocals stem using the pre-trained zero-shot voice conversion model HQ-SVC.

We have downloaded 11375 audio files from Freesound using the queries crowd, audience,cheering , applause, chatter, and protest.

crowdioset

CrowdioSet Audio Example

MOISESDB mixture

Original vocals

Five zero-shot voice conversions (HQ-SVC) used to synthesise sing-alongs:

HQ-SVC mix

AVOX Choir

Freesound Noise

Final live mixture (everything combined):

PaRIRset

We propose to simulate the acoustics of modern concerts. We have started from RIR collections commonly used in Speech Enhancement (SE): the ACE Challenge, the MIT IR Survey, and the SLR28 corpus.

However, these RIRs are limited for our purposes in two ways: first, the rooms captured (offices, classrooms, lecture halls) bear little resemblance to concert venues; second, modern live sound is delivered through a Public Address (PA) system with separate left and right loudspeaker arrays for stereo reproduction, a configuration absent from SE datasets.

We therefore measured our own set of impulse responses, which we call Public Address Room Impulse Response Set (PaRIRset) — the first dataset of real RIRs captured in professional concert venues with PA systems.

crowdioset

PaRIRset Audio Examples

MUSDB18HQ mixture (no RIR applied):

Mixture ⊛ Speech Enhancement RIRs :

Mixture ⊛ PaRIRset RIRs (ours):

Results

We have trained four models, each corresponding to a different training dataset configuration, which we refer to as clean, rev, noisy, and noisyrev respectively. The training datasets for each model are as follows:

  • clean: MUSDB18HQ + MOISESDB;
  • rev: (MUSDB18HQ + MOISESDB) ⊛ PaRIRset;
  • noisy: MUSDB18HQ + MOISESDB + CrowdioSet;
  • noisyrev: (MUSDB18HQ + MOISESDB + CrowdioSet) ⊛ PaRIRset.

For evaluation we have used the MUSDB18HQ test set, using our manually mixed audience stems test set and the PaRIRset test set. The same degradations that we have proposed for the training data were applied to the test data: we have evaluated each model under the four evaluation conditions clean, rev, noisy, noisyrev.

In addition, we have included SAM Audio to provide a baseline capable of audience isolation. We have included it in its textual prompting mode, taking the SAM Audio Large pre-trained variant, using the text queries “vocals”, “drums”, “bass”, “other”, and “crowd”

Note: take SAM Audio SDR results with caution, as the diffusion-based SAM Audio may exhibit sample misalignment.

results

We have complemented the objective evaluation with a subjective AB listening test, narrowing the tasks to vocals isolation and audience isolation. We have selected 10 real concert recordings from the audience, sourced from social media platforms, split into two groups of five. For the vocals isolation task, participants have been presented with a mixture and two separated vocals stems—one from our noisy model and one from the clean baseline—and asked which separation of the singer’s voice (as opposed to the crowd) they preferred, if any. For the audience isolation task, participants have compared the audience stem from our noisy model against SAM Audio. A total of 26 participants have taken part in the study, comprising media researchers and professional musicians or sound engineers.

subjective listening test results

Audio Samples

Sample 1

Mixture
Vocals Drums Bass Other Audience
clean Not supported
rev Not supported
noisy
noisyrev
SAM Audio

Sample 2

Mixture
Vocals Drums Bass Other Audience
clean Not supported
rev Not supported
noisy
noisyrev
SAM Audio

Sample 3

Mixture
Vocals Drums Bass Other Audience
clean Not supported
rev Not supported
noisy
noisyrev
SAM Audio

Sample 4

Mixture
Vocals Drums Bass Other Audience
clean Not supported
rev Not supported
noisy
noisyrev
SAM Audio

Sample 5

Mixture
Vocals Drums Bass Other Audience
clean Not supported
rev Not supported
noisy
noisyrev
SAM Audio

Sample 6

Mixture
Vocals Drums Bass Other Audience
clean Not supported
rev Not supported
noisy
noisyrev
SAM Audio

Sample 7

Mixture
Vocals Drums Bass Other Audience
clean Not supported
rev Not supported
noisy
noisyrev
SAM Audio

Sample 8

Mixture
Vocals Drums Bass Other Audience
clean Not supported
rev Not supported
noisy
noisyrev
SAM Audio

BibTeX

@inproceedings{guso2026crowdioset,
  title={CrowdioSet and PaRIRset: Two Datasets Towards Live Music Source Separation},
  author={Gus{\'o}, Enric and Serra, Xavier},
  booktitle={Proceedings of the 27th International Society for Music Information Retrieval Conference (ISMIR)},
  year={2026}
}

Acknowledgments

This work was financially supported by the Catalan Government through the funding grant ACCIO-Eurecat (Project TRAÇA: “IAGen” 2023-2026). We thank the band Cala Vento for allowing us to measure PaRIRset during their Brindis tour. We thank the Eurecat team for taking part in the listening test. In particular, we thank Umut Sayın for suggesting the use of the Zylia microphone array, and both him and Joanna Luberadzka for reviewing the manuscript. We also thank Pepe Ferrer from Global Audio Solutions for contributing to PaRIRset with an RIR.