Options
PolyBirdMix : A Large-Scale Synthetic Dataset and Benchmark for Bird Polyphony Estimation
Faculty
Contributor(s):
Publisher Information:
Zenodo
Year of publication:
2026
Language:
English
Abstract:
Automatic acoustic bird detection is an effective and noninvasive method for monitoring the health of ecosystems. The standard window-level presence/absence approach collapses any number of vocalizing individuals into a single detection. Estimating the number of vocalizing individuals per window would provide richer information, particularly valuable during high-activity periods like the dawn chorus or for monitoring rare species. Yet, no available dataset enables this task at scale due to the lack of individual-level bird annotations.
We present PolyBirdMix, a large-scale synthetic polyphonic bird call dataset. Sources are extracted from focal recordings via source separation and validated at the species level, supporting the assumption that each source corresponds to a single individual bird. Time-frequency masking is applied to isolate each source before mixing into polyphonic mixtures with controlled noise. The pipeline allows precise control over polyphony degree, SNR, and per-source signal level, and provides time-frequency bounds, noise-free audio, and per-source audio to support polyphony estimation, denoising, source separation, and curriculum learning. We describe the synthesis methodology and demonstrate its utility through baseline experiments.Changelog:- v1.0.0 (2026-07-12): Initial release — HSN, NES, PER, POW, SNE, SSW, UHH, XCM subsets.
We present PolyBirdMix, a large-scale synthetic polyphonic bird call dataset. Sources are extracted from focal recordings via source separation and validated at the species level, supporting the assumption that each source corresponds to a single individual bird. Time-frequency masking is applied to isolate each source before mixing into polyphonic mixtures with controlled noise. The pipeline allows precise control over polyphony degree, SNR, and per-source signal level, and provides time-frequency bounds, noise-free audio, and per-source audio to support polyphony estimation, denoising, source separation, and curriculum learning. We describe the synthesis methodology and demonstrate its utility through baseline experiments.Changelog:- v1.0.0 (2026-07-12): Initial release — HSN, NES, PER, POW, SNE, SSW, UHH, XCM subsets.
Type:
Collection
Keywords: ; ; ; ; ; ;
polyphony estimation
synthetic dataset
benchmark
bioacoustic monitoring
source-based annotation
passive acoustic monitoring
bird vocalization
Permalink
https://fis.uni-bamberg.de/handle/uniba/116531