Proceedings

Workshop on Detection and Classification of Acoustic Scenes and Events
27-29 October 2026, Boston, USA

Mark Cartwright, Frederic Font, Bashima Islam, Marko Stamenovic, and Shuo Zhang (eds.), Proceedings of the 11th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE 2026), October 2026.

ISBN (Electronic): This volume will be published at the time of DCASE Workshop 2026.

Link PDF

Abstract

The Detection and Classification of Acoustic Scenes and Events (DCASE) 2026 Challenge Task 1 Heterogeneous Audio Classification is a new task introduced in 2026 which focuses on the automatic classification of heterogeneous audio using the Broad Sound Taxonomy (BST), a general-purpose audio taxonomy that features 5 top-level and 23 second-level categories. The goal of this task is to evaluate sound classification models on diverse, real-world audio that varies widely in its nature, including duration and recording conditions. To that end, two complementary datasets were provided: an expert curated set, BSD10k-v1.2, and a larger crowd-sourced collection, BSD35k-CS, which presents real-world labeling variability. Participants were encouraged to explore audio-based, metadata-based, and multimodal approaches to sound classification, as well as to leverage hierarchical relationships between taxonomy categories. We received a total of 41 submissions from 12 teams. The highest scoring submissions include multimodal systems which take advantage of several pre-trained models for audio and text representation, using techniques for metadata processing and leveraging the two development sets. Overall, the best-scoring submitted system features an increase of ~13% above the provided baseline.

PDF
Abstract

Passive acoustic monitoring is increasingly used for biodiversity assessment, yet estimating the number of individuals of a species, rather than merely detecting its presence, remains an open problem. The BioDCASE~2026 Bird Population Counting task addresses this in multi-species zoo aviaries, where the number of individuals of each species is known from curator records, providing a controlled but realistic benchmark for acoustic abundance estimation. Participants estimate the abundance of three target species (Greater flamingo, Red-billed quelea, Hadada ibis) from continuous 48\,kHz recordings. The development set contains 140{,}899 three-second clips across six aviaries, and the evaluation set spans ten held-out aviaries ($\sim$380{,}000 clips), of which six are scored. An official baseline couples the ARIA detector with per-species detection-rate regression. Four teams submitted nine systems. The best reached a mean absolute error (MAE) of 9.17 individuals and the three top teams outperformed the baseline (MAE 24.50). The central finding is that detection-rate regression estimates the more separable quelea well but fails for the Greater flamingo, whose synchronous chorus saturates detection counts. The ibis is counted no better than a constant guess. The systems that improved most did so by modelling acoustic crowding directly, through synthetic-data density regression or a morphological measure of chorus energy, rather than by improving species detection.

PDF
Abstract

Self-supervised learning (SSL) for audio typically benefits from larger pre-training corpora, but this relationship assumes that training data carries meaningful learning signal. In ocean acoustics, this assumption is often violated: passive recordings are dominated by waves, wind, and silence, with biological vocalizations occupying a small fraction of the total duration. We show that this overrepresentation of noise causes SSL models to allocate representational capacity to acoustically trivial structure, degrading downstream performance. Using BEATs as a testbed, we demonstrate that pre-training on 30 hours of salient recordings (segments containing animal vocalizations) outperforms 1,240 hours of unfiltered data on a 31-class marine mammal classification benchmark. Under linear probing, the manually curated salient set achieves markedly better accuracy than the full corpus, despite using 41x less data. Automatic saliency-based filtering improves over unfiltered pre-training, but still falls well short of the manually curated salient set. Our findings suggest that, in domains where informative content is temporally sparse, data curation matters more than data scale for SSL pre-training.

PDF
Abstract

Audio moment retrieval consists of localizing the temporal segment, or segments, in a long audio recording that correspond to a natural-language query. Unlike conventional language-based audio retrieval, which typically operates at the clip level, this task requires open-vocabulary audio-text matching together with accurate temporal localization in untrimmed recordings. We propose a coarse-to-fine framework for audio moment retrieval that combines global proposal generation, local boundary refinement, and candidate-level reranking. The first stage applies a UVCOM-style temporal localization model to the full audio recording to obtain a ranked set of candidate moments. The second stage refines the most promising candidates on local crops using higher-resolution audio features, improving boundary accuracy. The final stage reranks the refined candidates using temporal descriptors, localization confidence, audio-text similarity, and boundary-context information. We additionally apply simple postprocessing steps, including timestamp rounding, minimum-duration enforcement, duplicate removal, and boundary clipping. Submitted to DCASE 2026 Task 6, our team achieved fifth place in the challenge, with an R1@0.7 of 40.01 for the best single system and 42.09 for the ensemble, however in the alternative metric R1@0.5 our team ranked 2nd in the evaluation with 64.41. These results indicate that decomposing audio moment retrieval into coarse localization, local refinement, and reranking is an effective strategy for language-based temporal localization in long audio recordings.

PDF
Abstract

This paper presents the Domain-Agnostic Incremental Learning for Audio Classification Task of the DCASE 2026 Challenge. Incremental learning refers to sequentially learning new tasks with the same system while maintaining its knowledge and performance on the previously learned task. Domain-incremental learning for sound classification refers to learning the same sound classes but in different acoustic domains, and was formalized as a data challenge for the first time in DCASE 2026. The task required participants to train a system to learn ten sound classes in three different domain, starting from a baseline system trained on the first domain. At each incremental task, learning had to be done without access to previous task data. The systems were ranked by the overall average accuracy calculated over the three domains. The task generated high interest and received submissions from 21 teams, with most submissions outperforming the provided baseline system. The highest performance of 79.6% was obtained using a combination of continual learning strategies that resulted in a balanced accuracy across domains coming, hence a good balance of retaining old and learning new information.

PDF
Abstract

Domain-agnostic incremental sound classification requires a model to sequentially learn new acoustic domains while preserving pre viously acquired knowledge and performing inference without ac cess to test-time domain labels. This setting is challenging because only data from the current domain are available during incremental learning, whereas evaluation samples may originate from any pre viously learned domain. Existing domain-specific classifiers often rely on known domain labels or unstable confidence-based routing, limiting their robustness in domain-agnostic inference. This paper proposes DIRNET-OOD (Domain-Agnostic Incremental Routing Network with Out-of-Distribution Residual Prior), a CNN14-based incremental learning framework with prototype residual routing for domain-agnostic inference. The proposed method performs pro totype residual routing by constructing class-wise prototypes for D2 and D3 in a learned routing space and estimating domain re sponsibility using probability-weighted prototype distances. Since D1 development data cannot be accessed during later incremental stages, the D1 branch is incorporated through an OOD residual prior instead of explicit prototypes. In addition, adaptive D2/D3 pairwise calibration refines the relative contributions of the D2 and D3 branches before soft probability fusion, improving routing sta bility when their prototype distances are similar. Experiments on the DCASE 2026 Task 7 benchmark show that DIRNET-OOD im proves the macro-average accuracy on D2/D3 evaluation samples from 52.6% to 74.89%, corresponding to a 22.29 percentage-point improvement over the official baseline. These results demonstrate that class-wise prototype routing, combined with an OOD resid ual prior and adaptive pairwise calibration, is effective for domain agnostic incremental sound classification.

PDF
Abstract

Audio moment retrieval (AMR) grounds free-form text queries as temporal spans in audio recordings. Target events range from short transients to tens-of-seconds scenes, so the tolerance of the tight-IoU metric varies substantially across queries. We address this with two design principles: modeling across temporal scales and aligning audio and text at multiple granularities. Our detector builds on QD-DETR over a single frozen M2D-CLAP encoder. For temporal scale, a coarse auxiliary branch supervises the shared encoder at halved resolution, denoising query training restores perturbed spans, a single-layer bidirectional Mamba encoder carries minute-long context, and a bounded cascade decoder re-attends to localized regions. For alignment, span re-ranking and boundary-contrastive objectives tie span and boundary regions to the query text, while a frame-level saliency objective aligns individual frames. At inference, zero-training saliency reuse converts this trained saliency head into a parallel proposal stream for proposal generation. Submitted to DCASE 2026 Task 6 as four cumulative variants, the system ranked joint sixth of 21 teams with 8.0 M trainable parameters per model. Paired per-query significance tests with false-discovery-rate control show that each stage improves at least one metric significantly, while several single-stage R1@0.7 differences remain within seed variability. Stratified analysis further shows that the largest gains occur for short, many-moment queries, where proposal coverage and boundary precision become tightly coupled at one-second temporal resolution.

PDF
Abstract

We present our system submitted to DCASE 2026 Task 6: Audio Moment Retrieval (AMR), which aims to retrieve a temporally grounded moment within a long audio recording given a natural language query. Following the AM-DETR baseline framework, we adopt frame-level audio feature extraction by segmenting long recordings into non-overlapping one-second clips, yielding a temporally ordered sequence of clip-level embeddings. Building on this, our primary contribution is a systematic investigation of pretrained audio and text encoders as replacements for the baseline MS-CLAP features, including M2D-CLAP and LAION-CLAP for both audio and text. The selected feature representations are fused into an ensemble and fed to an AM-DETR-based retrieval head for temporal boundary regression. We further incorporate frame masking during training and an IoU-based loss to improve localisation. Our system achieves a Recall1@0.7 score of 26.43 on the CASTELLA test split, surpassing the baseline score of 13.59.

PDF
Abstract

Domain-incremental learning (DIL) requires a single model to classify a fixed set of sound classes as the recording conditions change across a sequence of domains, without access to earlier-domain data and without forgetting. We study this problem in the setting of DCASE 2026 Task 7. Starting from the released baseline, which adds per-domain Batch Normalization (BN) and a per-domain classifier to a shared, frozen CNN14 backbone, we identify limited per-domain capacity as the main bottleneck on the harder domains. We insert, on top of the frozen backbone, a stack of zero-initialised residual bottleneck adapters after every convolutional block, and we set the stack depth per domain so that a domain receives capacity in proportion to how far it drifts from the frozen backbone. Combined with a warm-started per-domain head and a supervised contrastive loss, this raises the domain-agnostic average accuracy on the development set from 50.0% to 68.5%, while reproducing the first domain bit-for-bit so nothing is forgotten. To test generality we build a second DIL benchmark from TAU Urban Acoustic Scenes 2019, with four cities forming the domain sequence, and the same recipe again improves on the ADIL baseline across the sequence. Because these cities are balanced in difficulty, an equal (symmetric) depth is best here, whereas Task 7’s uneven domains call for an asymmetric one. Both follow a single rule: give each domain adapter capacity in proportion to its difficulty. Finally, we compare three domain-routing strategies and find the parameter-free entropy router the most robust: learned routers overfit the development split, and, on genuinely unseen domains (ESC-50), entropy routing misroutes with high confidence rather than abstaining, a failure mode we argue is more consequential for deployment than raw accuracy.

PDF
Abstract

Machine anomaly detection is widely used for predictive maintenance. When anomalous sound detection (ASD) systems are deployed in environments such as factories, we have noticed that a large number of false alarms are caused by human speech and nonverbal sounds such as laughing. In this paper, we investigate using a controlled evaluation framework where we contaminate existing ASD datasets with human vocalizations under various signal-to-noise ratios and durations. Experiments on different ASD systems show that human vocalizations significantly degrade anomaly detection performance. We further demonstrate that training with vocalization-contaminated data and applying source separation as a front-end both mitigate this degradation and partially restore anomaly-score distributions towards uncontaminated machine behavior.

PDF
Abstract

We present an empirical study of the factors that affect performance in first-shot unsupervised anomalous sound detection (ASD), based on our system developed for DCASE 2026 Task 2. We organize the study into two stages consisting of representation learning and anomaly scoring. For representation learning, we examine pretrained audio encoders, layer selection, encoder ensembling, and LoRA-based adaptation. For anomaly scoring, we investigate whitening-based normalization and several scoring back-ends, including local-density k-nearest neighbors (kNN), relative Mahalanobis distance (RMD), and their combination. Our results show that aggregating selected layers outperforms using only the final layer on the DCASE 2026 Task 2 development set. LoRA-based adaptation primarily improves cross-domain performance, while its in-domain effect varies across encoders. The effect of encoder ensembling depends on model composition, and adding more encoders does not necessarily improve performance. For anomaly scoring, whitening substantially improves performance, with local-density kNN providing the strongest individual back-end among those evaluated. Combining it with RMD further improves in-domain performance, whereas standalone RMD performs close to the chance level. The submitted system achieves 66.65% on the DCASE 2026 Task 2 development set and 65.45% on the DCASE 2026 Task 2 official evaluation set. These results are obtained using only near-field, single-channel audio. Index Terms—anomalous sound detection, first-shot unsupervised learning, encoder ensemble, LoRA, domain generalization

PDF
Abstract

In this paper, we introduce noise-aware self-supervised learning (NA-SSL) models for noise-aware anomalous sound detection (NA-ASD). NA-ASD is an ASD task with two-channel audio recordings, where one microphone is located close to the target machine and the other is located farther away to capture noise. For this task, we simulate two-channel recordings using diverse audio datasets and train NA-SSL models to extract clean SSL representations of the close-microphone signal by using the far-microphone recording dominated by background noise as auxiliary information. The NA-SSL models are then used as frontends in the standard ASD framework. Our experimental evaluation on the DCASE 2026 Challenge Task 2 development dataset demonstrates the effectiveness of the NA-SSL framework across three base SSL models (BEATs, EAT, and Dasheng), both with and without discriminative fine-tuning. Furthermore, the challenge results proved the effectiveness of the proposed approach, where the NA-BEATs system won the challenge by a large margin, achieving an official score of 70.24%, while the second-place system achieved 65.46%.

PDF
Abstract

Semi-supervised sound event detection models are routinely trained on synthetic strongly labeled, real weakly labeled, and real unlabeled audio, and selected on validation data that contains no strongly labeled real recordings. We show that this practice can fail silently. Under the classical DCASE protocol, Mean Teacher can degrade and even collapse to near-zero PSDS1 on real recordings while both quantities used for selection, the validation scores on the synthetic data and on a weakly labeled real data subset, remain unremarkable. Across nine architecture configurations we identify two geometrically distinct failure modes: sharp minima with elevated real-domain curvature, and flat plateaus of elevated real-domain loss. No quantity we examined that is computable without frame-level real labels separates failing from non-degraded runs. Among the diagnostics we test, the cheapest that separates non-degraded, degraded, and collapsed runs is the ratio of supervised loss on real versus synthetic audio, computed on held-out real evaluation data with frame-wise labels and therefore available only post hoc; its frame-level component reveals that what silently breaks is temporal localization. A seed-paired supervised-only control shows the failure is not specific to Mean Teacher; its objective widens a supervised instability in the sharp mode, while the flat mode did not occur without it. Replacing Mean Teacher with FlatMatch avoids both failure modes. On a 3M-parameter MobileNet where Mean Teacher reaches at most 0.33 PSDS1 and four of twelve seeds collapse to zero, FlatMatch reaches 0.46, within 0.04 of a 30-times larger CRNN+BEATs baseline. FlatMatch outperforms Mean Teacher on the recurrence-free MobileNet configurations and the transformer, while Mean Teacher stays ahead on every recurrent configuration; flatness-aware training acts as insurance against silent failure rather than as a universal performance booster.

PDF
Abstract

Unsupervised anomalous sound detection (ASD) for machine condition monitoring degrades when the background noise is non-stationary, leading to a mismatch between training and test conditions. DCASE 2026 Task 2, therefore, introduced noise-aware ASD with synchronized near/far microphone signals, while retaining the first-shot and domain-generalization requirements of DCASE 2025 Task 2. Because existing ASD datasets did not provide this observation, we recorded ToyCar, ToyDrone, and ToothBrush, and generated ToyCarEmu using measured impulse responses. The microphone arrangement is fixed within each machine type, and the two channels provide different target-to-noise dominance and arrival-time cues. We also adapted the DCASE 2025 baseline to stereo WAV input; it uses only the near channel as a reference for two-channel systems. We describe the four machine-type recodings, baseline, and evaluation results, providing a reproducible benchmark for noise-aware first-shot ASD.

PDF
Abstract

Contrastive language–audio pretraining (CLAP) provides a foundation for audio–language embeddings. Although several spatially-aware extensions have been developed, they are limited to static sources, leaving the modeling of moving sound sources as a critical challenge. We thus propose Moving-CLAP, a framework designed to capture time-varying sound source directions. We integrate an audio encoder recognizing the direction of a sound source in frame-level segments with a text encoder enhanced by spatially-biased attention that emphasizes directional tokens. We introduce two loss functions to improve spatial discriminability and directional consistency of temporal changes. Experiments show that Moving-CLAP learns embeddings that capture single-source motion, particularly its start and end directions, while effectively structuring the embedding space by spatial information. The embeddings also outperform the conventional framework in capturing two-source motions.

PDF
Abstract

DCASE~2026 Task~5 introduces Audio-Dependent Question Answering (ADQA), which tests whether large audio-language models answer from the audio rather than from textual priors. An Audio-Dependency Filtering (ADF) pipeline combines silent-audio probing, per-option perplexity, a large language model (LLM) commonsense check, and human review to remove items solvable from text alone. The 3000 items that pass form the ADQA-Bench evaluation set, spanning music, speech, and environmental audio. The inaugural edition draws 14 teams and 36 submissions across two tracks defined by total parameter count (up to 100B and under 10B). A Chung-Ang University ensemble of MOSS-Audio-8B-Thinking and Qwen3-Omni-30B reaches the top overall accuracy at \pct{58.33}, and a MOSS-only configuration from the same team leads the sub-10B track at \pct{57.30}. Across the 32 submissions with a published development score, evaluation accuracy falls by 11.93 percentage points (pp) on average (median 11.16\,pp) on the hidden evaluation split, consistent with its deliberately harder design and, for some teams, development-set overfitting. The most common building blocks are: the MOSS-Audio-8B-Thinking backbone (13 of 36 submissions), Low-Rank Adaptation (LoRA) fine-tuning on AudioMCQ-StrongAC, and preference or reinforcement-learning objectives---Group Relative Policy Optimization (GRPO) in five teams, Group reward-Decoupled Normalization Policy Optimization (GDPO) in two. At test time, prompt engineering is near-universal, and majority or choice-permutation voting is common. Every system misses the same 233 evaluation items.

PDF
Abstract

Audio-Dependent Question Answering (ADQA) requires Large Audio-Language Models (LALMs) to answer questions whose correct answers depend on the given audio content. Successful ADQA requires accurate audio perception, identification of question-relevant evidence, and cross-modal reasoning. Using the official ADQA dataset of DCASE 2026 Task 5, we investigate reasoning-oriented post-training with Low-Rank Adaptation (LoRA) and inference-time LoRA rescaling for both Qwen2.5-Omni and MOSS-Audio-8B-Thinking. We introduce a structured Chain-of-Thought (CoT) framework that decomposes the reasoning process into question analysis, question type, audio evidence, and reasoning. We then analyze how task-specific LoRA adaptation affects the two backbones and further explore inference-time rescaling of trained LoRA adapters. Experiments on the development set reveal markedly backbone-dependent behavior: post-training improves the Qwen-based systems but substantially degrades MOSS-Audio under our supervised fine-tuning configuration. Moderate LoRA rescaling further improves the best Qwen system's top-1 accuracy from 58.93% to 61.05% and partially restores the performance of the fine-tuned MOSS-Audio models, while the best MOSS-Audio system achieves 67.70% top-1 accuracy. Our submitted systems ranked third overall and second among lightweight systems under 10B parameters in the challenge.

PDF
Abstract

This report describes our first-place system for DCASE 2026 Task 5, Audio-Dependent Question Answering (ADQA), a multiple-choice benchmark designed to assess whether audio-language systems ground their answers in the input audio rather than relying on textual priors. Our system is built around MOSS-Audio-8B-Thinking and emphasizes training-free audio-aware inference using calibrated answer-letter prompting, rule-routed acoustic tagger assistance, and permutation-based ensemble voting. We additionally investigated GDPO-based LoRA adaptation with audio-contrastive hard examples selected by comparing original-audio and silent-control option scores. On the official ADQA-Bench evaluation set, our training-free tagger-assisted MOSS system achieved 57.30% Top-1 accuracy and ranked first among lightweight submissions. Although the GDPO-adapted MOSS system achieved a slightly lower accuracy of 57.07%, it contributed complementary predictions for ensemble construction. The final mixed ensemble, combining permutation-based MOSS predictions with Qwen-based predictions, achieved 58.33% Top-1 accuracy and ranked first overall. These results highlight the effectiveness of training-free acoustic evidence injection and prediction-level diversity for hard-filtered audio-dependent question answering.

PDF
Abstract

Abstract— Existing anomalous sound detection (ASD) datasets capture industrial noise, domain shifts, and clip-level operating conditions, but provide limited time-aligned evidence for relating detector responses to the evolving operation of a complex machine. We present ASD3DP, a process- aligned operating-sound dataset for fused filament fabrication (FFF) 3D printing. The dataset contains 41 recording sessions and 28.99 h of machine operation acquired through four synchronized microphones, together with 9,067 selected 10 s clips covering normal operation, belt- tension conditions, extruder failures, and toolhead collisions. Each audio interval is linked to a time-aligned command stream, telemetry, process states, print features, and fault-active intervals on a common release-audio timeline. The release supports conventional clip-based ASD and annotation-linked analysis of audio-only model responses. Baseline experiments establish compatibility with reconstruction-based and pretrained-encoder ASD systems. Process attributes are reserved for analysis after the systems and decision thresholds are fixed. The dataset, annotations, and accompanying documentation are publicly available at https://github.com/LUDO-Lab-Research/ASD3DP

PDF
Abstract

This paper presents QAM-DETR, our system for DCASE 2026 Challenge Task 6: Audio Moment Retrieval from Long Audio. The task requires localizing the temporal segment in a long audio recording that best matches a given natural-language query. Unlike clip-level audio-text retrieval, audio moment retrieval demands precise boundary prediction, long-range temporal modeling, and reliable ranking of candidate moments. QAM-DETR builds on a DETR-style moment localization framework using frozen pretrained audio and text representations. LAION-CLAP is used as the primary audio-language representation, while MS-CLAP, WavLM, and RoBERTa features are selectively integrated through lightweight learnable gates to exploit complementary acoustic and linguistic information. To improve query-conditioned temporal localization, the proposed system introduces 3-way cross-attention-based cross-modal fusion, a BiMamba-based bidirectional temporal encoder, lightweight multi-scale temporal fusion, and a quality prediction branch for candidate ranking. For selected final submissions, a frozen audio-language LLM verifier is further used as a post-processing re-ranker, which re-scores candidate segments without modifying their predicted temporal boundaries or updating model parameters. Experiments on the CASTELLA development splits show that the proposed components consistently improve the DETR-style baseline, with BiMamba temporal encoding and quality-aware ranking providing substantial gains. Final ensemble systems further improve retrieval performance by combining diverse feature configurations and, in selected variants, LLM-based candidate re-ranking. The final system achieved joint first place in DCASE 2026 Task 6, demonstrating the effectiveness of quality-aware temporal modeling and candidate re-ranking for long-audio moment retrieval.

PDF
Abstract

The Cross-Domain Mosquito Species Classification (CD-MSC) task focuses on mosquito species recognition under domain shift caused by variations across devices, environments, and collection protocols. Models trained on multiple source domains may perform poorly when evaluated on recordings collected in different settings. To address these challenges, we propose a robust classification framework that integrates diverse data augmentation, domain-invariant representation learning, and ensemble inference. Specifically, we employ multiple data augmentation strategies to enhance the diversity of training samples. To learn cross-domain features, we insert MixStyle modules into a DenseNet backbone to dynamically perturb feature statistics. Furthermore, the feature embedding is optimized using four complementary losses to mitigate cross-domain generalization issues. Finally, we adopt an ensemble inference strategy that leverages the strengths of different network architectures and training configurations for specific species. Experiments on the test set of the development dataset show that the proposed system achieves 0.4016 balanced accuracy on unseen domains, substantially outperforming the baseline of 0.1751.

PDF
Abstract

We introduce F1Audio, a novel dataset of features extracted from Formula-1 onboard engine audio, synchronized with car telemetry comprising approximately 300 hours of data from 10 Grand Prix events of the 2024 Formula-1 season. This is the first large-scale dataset combining engine audio features with vehicle speed, engine revolutions per minute (RPM), gear selection, and throttle percentage. We also provide benchmark results for the prediction of telemetry targets from the engine audio. Index Terms—Formula-1, engine acoustics, telemetry estimation, dataset, audio-derived features

PDF
Abstract

This paper presents an overview of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2026 Challenge Task 6, Audio Moment Retrieval (AMR) from Long Audio. Given a several-minute-long audio recording and a free-form text query, AMR aims to retrieve temporal moments in the recording that match the query, where each moment is represented by a pair of start and end timestamps. This task requires effective cross-modal alignment and long-range temporal modeling. We describe the task definition, the evaluation metrics, the development and evaluation datasets, and a baseline system that combines a pre-trained MS-CLAP feature extractor with a Detection Transformer (DETR)-based moment-detection network. On the development data, the baseline trained on a manually annotated dataset and a synthetic dataset achieved Recall1@0.7 of 13.56%, indicating that AMR in long audio remains a challenging problem. The challenge attracted 21 teams, which submitted 59 systems in total. The best system achieved Recall1@0.7 of 48.59%, roughly 3.5 times the baseline score. The results show that strengthening the audio-text feature extractor and the moment-detection network led to substantial performance improvements. Furthermore, the top three teams boosted performance by applying confidence score calibration or ensembling across different temporal resolutions of features.

PDF
Abstract

Antarctic whale vocalization detection requires temporally localizing call boundaries in long underwater recordings and classifying them into predefined call types. Training reliable machine learning models for this task is difficult due to the variability and sparsity of whale calls in passive acoustic monitoring datasets. These difficulties are further amplified by time-varying noise that can mask whale calls and by site- specific ambient sound characteristics to which a trained model may not generalize. In this work, we present NAVE, a Conformer-based detector that improves whale-call detection via three components. First, its input representation combines per-bin complex mean subtraction, which suppresses constant site-specific ambient sound characteristics, with an additional per-channel energy normalization (PCEN) channel that adapts to temporal variations in the noise floor. Second, a frequency-dynamic stem processes these noise-aware representations with convolutions that adapt across the characteristic frequency bands of the target call types. Third, we use a Conformer backbone with a convolution kernel matched to the duration of the target calls that adds an inductive bias and improves temporal localization. We evaluate NAVE on BioDCASE 2026 Task 2 using leave-one-site-out cross-validation across the validation sites, which approximates deployment on unseen recording sites. A single NAVE model with a tuned per-class threshold exceeds the stronger WhaleVAD- BPN baseline, using no boundary-proposal network and no search over post-processing parameters. A controlled ablation study shows that the normalized front-end and a duration-matched convolutional kernel each contribute substantially to recovering the D downsweep, the hardest class for previous detectors, with the wide kernel giving the largest aggregate gain.

PDF
Abstract

The performance of acoustic machine learning systems is commonly evaluated using fully annotated test sets. In real-world deployments, however, exhaustively labeling large volumes of continuously collected audio data is often infeasible. Consequently, performance assessment typically relies on a small labeled subset of the available data, introducing a sampling bias that can severely distort evaluation metrics. This paper studies methods for compensating the bias in evaluation-labeled subsets under strict annotation-budget constraints. We study whether importance weighting techniques can mitigate this discrepancy by compensating for the selection bias. Specifically, we implement and compare three density-ratio estimation methods: kernel density estimation (KDE), logistic regression, and k-nearest neighbors (kNN), utilizing feature-space representations of the deployed audio. To emulate realistic deployment scenarios, the labeled subsets are generated using five distinct sampling strategies based on active learning techniques. Experiments conducted on an audio scene classification (ASC) benchmark demonstrate that importance weighting consistently yields more realistic accuracy estimates, significantly reducing the gap between subset-based metrics and the true evaluation performance.

PDF
Abstract

This paper presents an overview of DCASE 2026 Challenge Task 2, titled “Noise-aware unsupervised anomalous sound detection (UASD) for machine condition monitoring.” Reliable detection under noisy conditions is crucial for practical deployment, but previous DCASE Task 2 settings provided limited information about environmental noise, potentially limiting UASD performance in highly noisy situations. To address this limitation, DCASE 2026 allows participants to exploit two-channel audio samples simultaneously captured at locations near and far from the target machine. Since the distant microphone is expected to contain relatively stronger environmental noise and weaker direct machine sounds, it may help distinguish environmental noise components from the target machine sounds. We received 173 submissions from 52 teams, and an analysis of these submissions has been made in this paper. The results revealed that using the second channel can improve detection performance under noisy conditions. Specifically, the first ranked team used a self-supervised learning approach that tries to directly compute an SSL representation of the clean target machine sound from the two-channel recordings.

PDF
Abstract

Environmental sound understanding, integrating sound event detection (SED), environmental sound classification (ESC), and signal-to-noise ratio (SNR) estimation in noisy environments, supports emergency response and disaster monitoring. However, three challenges remain. First, background noise degrades detection and classification performance. Second, although source separation improves noise robustness, language-driven source separation models may generate target-like signals even when the target sound is absent, causing false detections and misclassifications. Third, conventional SNR estimation directly regresses SNR from noisy mixtures without an explicit reference representation, limiting its use as a confidence measure. To address these challenges, we propose an integrated framework combining the language-driven source separation model AudioSep with the audio-language foundation model CLAP. Target and noise text embeddings are used as anchors. Source separation is combined with parameter-efficient adaptation to align acoustic embeddings of noisy mixtures with target anchors, improving noise robustness. Contrastive learning is further introduced to separate embeddings generated by mismatched source separation from the target anchors, suppressing false detections and misclassifications. A continuous, sigmoid-interpolated anchor is constructed between the target and noise anchors, and SNR is estimated from embedding similarity to this anchor. Experiments show that the proposed method achieves a recall of 51.2% and a specificity of 96.4% for SED of sirens at -20 dB SNR. For ESC over seven sound categories, it achieves an accuracy of 77.3% at -20 dB, substantially outperforming the baseline CLAP model (16.2%). The proposed anchors also support an SNR estimate that reflects the continuous ordering of SNR conditions. These results demonstrate the effectiveness of the proposed framework for robust environmental sound understanding under severe noise conditions.

PDF
Abstract

Context is central to human auditory perception, as listeners interpret sounds within their acoustic environment to resolve ambiguity. However, large-scale datasets that jointly capture sound events and acoustic scenes remain limited. We address this by augmenting AudioSet Strong with clip-level scene annotations under two taxonomies for joint scene-event modeling. A corpus-scale analysis shows that scene identity is better captured by events distinctive to a scene than by those that merely occur frequently. Reweighting event co-occurrences toward scene-specific associations produces representations that better align with human-derived taxonomies, while also exposing scenes where the acoustic and semantic views diverge. This geometric divergence predicts classification errors, while supervised validation confirms that the proposed scene annotations provide meaningful complementary information for developing scalable context-aware machine listening systems.

PDF
Abstract

Large Audio Language Models (LALMs) enable generative reasoning over complex acoustic scenes but often exhibit prediction instability. We show that this variability can serve as a reliability metric via inference-time self-consistency and confidence-aware aggregation. Using a log-probability–based scoring framework, candidate answers are weighted by model likelihoods, leveraging reasoning diversity and self-certainty without additional training. Experiments on four LALMs, including Qwen3-Omni and Audio Flamingo 3, demonstrate that confidence-aware aggregation improves accuracy and calibration, with additive methods like Borda voting outperforming classical Majority Vote, especially under high disagreement. This approach transforms LALM variability into a source of reliability for scalable and robust multimodal audio understanding. Project code is available at https://github.com/michelolzam/lalm-self-consistency.

PDF
PDF
Abstract

To successfully deploy a model in time-varying environments such as streaming data prediction and sensing control, domain-incremental learning (DIL) has attracted attention since it aims to adapt a previously trained model to newly arriving domains, while preserving knowledge from earlier domains without accessing their data. Incremental learning across domains can be regarded as a recurrent update, in which the current model is obtained by updating the model carried over from previous domains. Conventional DIL approaches that rely on domain-invariant feature learning and weight regularization gradually overwrite or constrain parameters learned in previous domains, leading to catastrophic forgetting. Instead, this paper proposes a new domain-specific parameter-isolation architecture that retains all past domains. The proposed architecture mitigates catastrophic forgetting through a full-order recurrent update, constructing a new expert using domain-specific data conditioned on all previously frozen models. To achieve this, we incorporate data-free generative replay to reconstruct previous-domain data and cross-domain feature generation to recover later expert features missing from earlier domain samples. Finally, we apply the proposed model architecture to domain-agnostic incremental learning for audio classification, as defined in the DCASE 2026 Challenge Task 7. Consequently, we achieve micro and macro accuracies of 78.4% and 78.9%, respectively, representing increases of 33 and 25 percentage points over the Challenge baseline. Ablation studies are conducted to examine the effectiveness of each processing component in terms of classification accuracy.

PDF
Abstract

The Broad Sound Taxonomy (BST) provides a compact hierarchy for organizing arbitrary sound recordings into 5 top-level and 23 second-level categories. The DCASE 2026 Challenge Task 1, Heterogeneous Audio Classification, evaluates systems for automatically categorizing recordings from the Freesound platform into the correct second-level BST category based on audio and textual metadata. In this work, we investigate a metadata-aware classification approach that uses a large language model to predict BST class probabilities from clip metadata and fuses these predictions with audio-text embeddings. In addition, we study strategies for exploiting a larger crowdsourced and noisily annotated training set. To this end, we generate soft pseudo-labels from models trained on clean data and use them as training targets for the noisy subset. We evaluate our approach on the BSD10k dataset in the context of DCASE 2026 Task 1 using M2D- and CLAP-based classifiers. Our results suggest that using LLM-derived class predictions is complementary to processing metadata only with a pretrained CLAP text encoder. For using a larger dataset with crowdsourced annotations we find that training on soft pseudo-labels is more effective than using the noisy labels, but that our simple pseudo-labeling strategy does not improve over training on the clean data alone; performance appears to saturate even as the pseudo-labeled set grows. Finally, error analysis shows that our model's class confusions are concentrated in classes that also exhibit human annotation ambiguity, suggesting that further progress may require sharper BST class definitions in addition to improved audio-text modeling.

PDF
Abstract

Sperm whales (Physeter macrocephalus) communicate through short click sequences called codas, whose structure includes the rhythmic pattern of inter-click intervals (ICI) and the spectral composition of the clicks. In this paper, we propose the Dual-Channel Contrastive Encoder (DCCE), which explicitly encodes rhythm and spectral information with separate encoders before fusing them. We evaluate DCCE on three coda-understanding tasks: social unit, coda type, and individual identification. We compare DCCE against strong baselines from literature under a common linear-probe protocol and analyze the role of each channel and training objective through ablations. On 1,383 codas, shared partitions and validation-only layer selection support three random splits and three recording-date-disjoint folds. DCCE with unit pairing and type supervision achieves mean random-split macro-F1 across three runs of 0.901/0.865/0.597 for unit/type/identity. Raw ICI remains strongest for type (0.987), and ICI plus temporal mel statistics is a strong identity reference (0.724). An auxiliary identity objective improves spectral-branch identity prediction but does not improve the mean fused-identity score in this configuration. Date-grouped evaluation and four year holdouts expose generalization limits. The results support task-dependent combinations of timing and spectral information, without establishing independent biological channels or general architectural superiority.

PDF
Abstract

In this study, we address the automated detection of whale vocalizations in low-sampling-rate passive acoustic monitoring (PAM) recordings from the BioDCASE Challenge 2026 Task 2. We investigate two spectrogram-based object detectors from the computer vision domain, YOLOv11 and RT-DETR, alongside an audio-native transformer model, BEATs, fine-tuned on upsampled waveforms. Using an event-level evaluation based on 1D intersection-over-union (IoU) along the temporal axis, we observe that the BEATs model delivers the highest performance among the base models on the test set (F1 = 0.655), indicating that large-scale audio pre-training transfers effectively even to narrowband, low-frequency whale vocalizations. RT-DETR clearly surpasses YOLOv11 in its base configuration in the test set of the development set (F1 = 0.517 vs. 0.430), and a detailed analysis of model-agnostic post-processing reveals that adding a simple non-maximum suppression (NMS) step is essential to improve instance-level performance, increasing RT-DETR’s test-set F1 score to 0.679. On the evaluation set, the same overall ranking is observed, with BEATs achieving an F1 score of 0.376, RT-DETR with NMS reaching 0.276, and YOLOv11 with augmentation obtaining 0.166. These results demonstrate that (i) self-supervised transformer models pre-trained on large-scale audio datasets already provide strong baselines for marine bioacoustic event detection, and (ii) transformer-based detectors with carefully tuned NMS remain competitive for precise temporal localization, offering a complementary approach when frame-level spectrograms and high-resolution instance-level annotations are available.

PDF
Abstract

This paper describes a multi-stage separation and classification system for DCASE 2026 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes (S5). Inspired by the top-ranked system in DCASE 2025 Task 4, we adopt a cascaded framework consisting of universal sound separation (USS) with source counting, source classification, and class-aware refinement. In the first stage, a TF-Locoformer-based USS model separates multi-channel mixtures into single-channel foreground and interference signals. Each separated signal is then classified as one of 18 foreground classes or as interference. The separated foreground signals are further refined by another TF-Locoformer-based model conditioned on the predicted class labels and the observed mixture. Our best system achieves a CA-PI-SDRi of 12.94 dB on the final evaluation set, ranking third overall. Notably, the system achieves a label prediction accuracy of 76.92%, the best classification performance among all submissions.

PDF
Abstract

This paper presents DOOM-SED, a downstream-oriented open-vocabulary model for sound event detection. Open-vocabulary SED aims to detect sound events described by arbitrary textual labels. Existing methods have been formulated as query-based detection, in which each target class is specified and detected independently. This formulation, however, ignores relationships among semantically related target labels, which often increases crosstalk errors. In this paper, we propose a downstream-oriented training framework and network architecture that explicitly handle a given set of target labels. Specifically, we train a detection model to solve dynamically generated detection episodes that mimic downstream SED tasks with a small set of target classes. The model is designed and trained to capture inter-label relationships through self-attention over label embeddings. Each episode is constructed on the fly from temporally sound event labels in AudioSet Strong and carefully designed based on the hierarchical structure of the AudioSet ontology. Experimental results on the publicly available DESED, URBAN-SED, and MAESTRO Real benchmarks demonstrate that DOOM-SED significantly improves downstream performance over existing open-vocabulary models.

PDF
Abstract

This paper presents RealDESED, a real-world domestic sound event detection (SED) benchmark comprising 5,710 audio recordings collected by 652 participants in their homes. Each recording is between 15 and 35 seconds long and contains temporally precise annotations for 15 common domestic sound classes. In contrast to existing SED datasets, which typically rely on simulated soundscapes or broad web-crawled audio, RealDESED consists exclusively of recordings captured in natural domestic environments, reflecting realistic variability in recording devices, device placement, acoustic conditions, background sounds, and naturally occurring event co-occurrences. A distinguishing characteristic of the dataset is its multi-annotator labeling scheme, where each recording is independently annotated by multiple annotators, while the validation and test sets undergo an additional review process to ensure high annotation quality and reliable benchmarking. Furthermore, the dataset provides rich metadata, including recording device, device placement, environment labels, and textual scene descriptions. We establish a strong transformer-based baseline and investigate annotation aggregation strategies, post-processing methods, long-form inference, and the impact of recording metadata on model performance. Our baseline achieves a macro-averaged PSDS1 score of 0.731 on the test set. We believe RealDESED provides a valuable benchmark for developing and evaluating robust SED systems under realistic domestic conditions, helping to bridge the gap between current research benchmarks and real-world deployment.

PDF
Abstract

Diffusion-based text-to-audio generative models such as AudioLDM achieve high perceptual quality and strong semantic consistency; however, their practical deployment is hindered by the substantial computational cost of the U-Net denoising backbone. In this work, we apply model pruning to improve the computational efficiency of AudioLDM, a U-Net based text-conditioned audio latent diffusion model. We analyse parameter redundancy across U-Net convolutional blocks and evaluate a filter-pruning strategy. Pruning is guided by norm-based criteria and followed by lightweight finetuning to recover performance losses. Experimental results demonstrate that up to 83% of the parameters and 39% of the multiply–accumulate operations of U-Net have been reduced while maintaining, and in some cases improving, generation quality compared to the baseline unpruned network. We find that pruning affects AudioLDM’s ability to generate certain sound events including safety-critical sounds such as gunshots, sirens, and explosions, as well as mechanical sounds such as drills and sewing machines, and other sounds such as sprays and tick-tocks, which are mostly recovered by lightweight finetuning of the pruned model.

PDF
Abstract

Spatial semantic segmentation of sound scenes (S5) involves the detection and extraction of target sound events in audio files. However, in typical detection-extraction pipelines, errors in the detection module propagate to the extraction module, degrading overall performance. In this work, we develop a multichannel detection-extraction model and evaluate inference-time algorithms to improve the class detection and target sound extraction in a multi-stage setup. We evaluate two fitness functions: binary cross entropy and mixture consistency; two relabeling strategies: naive and conditional confusion matrix relabeling; and three source estimate evaluation models: a multi-label classifier, a fine-tuned single-source classifier, and an off-the-shelf audio judge. We thoroughly evaluate the impact of these design choices on the DCASE 2025 Task 4 dataset, and validate entropy-based gating of inference-time updates on the held-out evaluation set, laying the groundwork for inference-time updates of S5 systems in real-world scenarios.

PDF
Abstract

We describe our submission to DCASE 2026 Task 5, Audio-Dependent Question Answering (ADQA), an entry to the sub-10B lightweight champion track. The official training set pairs each multiple-choice question with a chain-of-thought (CoT) trace generated by Gemini 3.1 Pro, but our audit finds these traces systematically flawed: 68.5% never engage the answer options they should discriminate, and 10.1% are contaminated, reasoning from a written ``transcript'' that does not exist at inference time and teaching exactly the textual-hallucination behaviour the benchmark penalizes. We therefore repair the data rather than the model, regenerating the 14,129 flagged traces with a blind, audio-only Gemma-4-E4B-it→Qwen3-Omni-30B-A3B teacher cascade under rejection sampling: a trace is kept only when the teacher independently selects the gold option from the audio alone. We fine-tune Qwen2.5-Omni-7B on the repaired data and further optimize the best variant with GRPO; a reference-guided reasoning judge adds no accuracy, so our strongest model uses answer-accuracy and format rewards alone. Using only the 19,480 officially provided items and no external data or models, our best system scores 58.4% on the development set (+5.4 points over the base model) and 50.6% on the withheld ADQA-Bench evaluation set.

PDF
Abstract

This paper presents an overview of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2026 Challenge Task 4, Spatial Semantic Segmentation of Sound Scenes (S5). The S5 task focuses on the joint detection and separation of sound events in complex spatial audio mixtures, supporting many applications, such as immersive communication. First introduced in DCASE 2025, the S5 task continues in DCASE 2026 with key changes to better reflect real-world conditions, including allowing mixtures to contain multiple sources of the same class and to contain no target sources. In this paper, we describe the task setting, including updates to the evaluation metrics and dataset, and report the results of 31 submissions from 10 teams. The official access point for data and code is https://github.com/nttcslab/dcase2026_task4_baseline.

PDF
Abstract

Conventional general sound recognition systems typically output deterministic sound event labels, implicitly assuming that the target sound class can be correctly identified from the input audio. However, in real listening situations, the sound event class is not always clearly identifiable. Human listeners may nevertheless understand their surroundings from an ambiguous sound without identifying its exact sound event class. This motivates a discussion of how the outputs of sound recognition systems should be redesigned under such uncertainty. As a basis for this discussion, this paper proposes an output representation for sound event recognition that combines a sound event class, its confidence score, and an onomatopoeic description of the sound. The proposed representation preserves conventional class-based recognition while providing an additional onomatopoeic description of acoustic characteristics that can remain informative even when the class prediction is uncertain. Experiments using ESC-50 and ESC-50-Onomatopoeia show that the proposed method achieves sound recognition performance comparable to that of a conventional recognition-only system. In addition, an LLM-as-a-judge evaluation and subjective listening experiments indicate that the proposed output is preferred over conventional deterministic outputs based on the sound event label, particularly when used to support understanding of the surrounding environment. These results suggest that such output representations can make sound event recognition more informative and communicative under uncertainty.

PDF
Abstract

Active learning mitigates annotation costs in bioacoustic classification; however, effective batch selection remains challenging for multi-label audio data. Conventional uncertainty-based strategies frequently sample redundant instances, whereas diversity-based approaches tend to overlook informative samples crucial for optimizing the current classifier. To address these limitations, we propose Pareto-Balanced Multi-Funnel Selection (PB-MFS), a novel batch active learning strategy that decouples candidate filtering from final batch selection. Specifically, PB-MFS initially constructs a survivor set via multiple complementary acquisition funnels. Subsequently, it determines the optimal annotation batch by harmonizing uncertainty and diversity through a Pareto optimization objective. The proposed method integrates diversity-driven cold-start selection, multi-funnel survivor voting, and an uncertainty-diversity trade-off leveraging pretrained audio embeddings. Comprehensive evaluations on the four BioDCASE 2026 Task 4 subsets (HSN, POW, UHH, and ATBFL) demonstrate that PB-MFS achieves an aggregated mean AULC of $0.510$. Furthermore, it outperforms the strongest baseline, CoreSet, by an absolute margin of $0.050$, delivering a $10.9\%$ relative performance gain.

PDF
Abstract

Tiny bioacoustic recognition on microcontrollers requires balancing classification accuracy with stringent resource constraints. Existing approaches mainly focus on model compression while paying limited attention to acoustic front-end representations and the practical execution characteristics of target hardware. To address this challenge, we propose a hardware-aware, representation-asymmetric teacher-student framework with a diverse teacher pool. Instead of requiring the deployable student model to learn from the same low-cost representation as the teachers, our framework constructs multiple teachers using diverse network architectures and feature-rich front-end representations. Their soft predictions are aggregated and distilled into a compact student model operating solely on computationally efficient front-end inputs. To further improve deployment performance, the student combines global average pooling and global max pooling. On the BioDCASE-Tiny 2026 Task 3 benchmark, the proposed framework ranked first among all submissions on the ESP32-S3 evaluation track. These results highlight the effectiveness of jointly optimizing acoustic representation, knowledge distillation, and hardware-aware deployment for tiny bioacoustic recognition.

PDF
Abstract

In daily life, people hear speech, footsteps, and music around them. We can often recognize these sounds and judge where they come from. Each sound source can be shown on a separate acoustic map, a rectangular image covering 360° horizontally and 180° vertically. The map shows the directions occupied by the source as a region and the sound energy within that region. A class label identifies the sound. Predicting these labeled acoustic maps from audio is called semantic acoustic imaging. Such maps could help robots perceive their surroundings and allow augmented reality displays to show sound regions and classes over the real world. Existing models can recognize sound classes and estimate a direction for each source. However, a direction alone does not describe the source region or its energy. Acoustic imaging must also distinguish sound sources in nearby directions, while the number of active sources and the regions they occupy can change over time. We therefore propose the Semantic Acoustic Imaging Detector (SAID), which predicts a separate labeled acoustic map for each active source from audio. First, we pretrain Audio2Sph, SAID's audio encoder, through sound energy estimation across directions without class labels. Then, we train the complete SAID model to predict source regions, energy, and classes together. We also develop a pipeline that generates simulated recordings for pretraining and supports fine-tuning on real recordings. On the official DCASE2026 Task 3 Track A evaluation set, our submitted system ranks first with 0.1080 macro-averaged mean average precision (Macro mAP) and 0.3962 Macro Pearson r.

PDF
Abstract

Pretrained audio spectrogram transformers are widely used as backbones of machine listening systems for sound event classification. However, they are typically trained using spectrograms with a fixed Short-Time Fourier Transform (STFT) configuration, implicitly binding the model to the time-frequency resolution seen during pretraining. In practice, adapting these systems to new domains or computational constraints often requires different STFT frontend parameters. In this work, we conduct a controlled study of the sensitivity of the foundational Audio Spectrogram Transformer to variations in STFT window and hop size, showing that audio classification performance degrades substantially as frontend parameters deviate from pretraining settings. As a mitigation strategy, we introduce a lightweight conditioning mechanism using FiLM that enables a pretrained model to adapt to variable input configurations using limited additional training data. Our method generalizes well to unseen STFT configurations, improving model robustness without sacrificing performance in the original setting.

PDF
Abstract

Passive acoustic monitoring enables population size estimation from audio recordings, but existing methods show higher estimation error in large populations, because larger groups tend to produce more frequent and severe call overlap that distorts the acoustic evidence such methods rely on. To address this problem, we propose a population estimation system based on overlapping call density regression. The proposed system employs a convolutional recurrent regressor that predicts a call density for each audio segment from its log-mel spectrogram, trained with a square-root-domain mean squared error loss to reduce the dominance of high-density samples during training. The segment-level density predictions are then aggregated into the mean and standard deviation, which are mapped to the final population estimate through linear regression. We conduct experiments on the real-world dataset from the BioDCASE 2026 Bird Counting Challenge, using the Greater flamingo as a representative flock-calling species. Experimental results show that the proposed system substantially outperforms the challenge baseline. Under leave-one-out cross-validation on the four development-set aviaries, the mean absolute error (MAE) is reduced from 23.0 to 1.5. In the official challenge evaluation, the proposed system ranked first among all submissions.

PDF
Abstract

Audiovisual sound source localization and detection aims to identify sound-emitting objects and estimate their spatial acoustic distributions by jointly exploiting visual and multichannel audio information. In this paper, we present an acoustic map-guided cross-modal fusion framework for DCASE 2026 Task 3 Track B. The proposed system represents spatial acoustic cues as equirectangular acoustic maps derived from multichannel microphone signals, while visual features are extracted using a pretrained ResNet-50-FPN backbone. To better exploit the complementarity between the two modalities, acoustic spatial cues are injected into visual feature maps through gated residual attention, and a ResNet-Conformer adaptation module is used to enhance contextual modeling. At the region level, visual ROI and acoustic ROI features are further fused with cross-modal attention to improve localization and classification. In addition, the system incorporates priors from an AV-SELD teacher model, together with score calibration and region-wise acoustic support estimation, to improve class ranking, source-count control, and acoustic energy reconstruction. Our best system achieved a Macro mAP of 0.013 and a Macro Pearson correlation of 0.400 on the official evaluation set, ranking first in DCASE 2026 Task 3 Track B.

PDF
Abstract

Temporal misalignment is common in multichannel bioacoustic analysis, especially when recordings are acquired using independent devices or unsynchronized channels. We formulate alignment as the estimation of a time-varying inter-channel delay trajectory and adapt cost-volume matching to this setting. The peak-aware cost-volume model constructs explicit matching evidence over candidate temporal lags from encoded two-channel short-time Fourier transform magnitude features. Peak-aware pooling then summarizes the lag-domain evidence into dominant responses and reliability cues, while a bidirectional long short-term memory models their temporal evolution. The estimated frame-level delay trajectory is then interpolated to the target keypoints and converted into matching timestamps in the other channel. On the zebra finch validation subset, the neural model improves coarse alignment over the official deep-learning baseline. In the final submission, the neural estimator is combined with generalized cross-correlation with phase transform (GCC-PHAT) through confidence-based routing, with the GCC-PHAT branch providing the fine-alignment gains on aru.

PDF
Abstract

Heterogeneous audio classification under the Broad Sound Taxonomy (BST) requires accurate second-level predictions while maintaining consistency with the top-level taxonomy structure. CLAP, a pre-trained model based on contrastive audio-text learning, has been adopted for heterogeneous audio classification due to its strong capability in learning high-level semantic representations. However, its emphasis on semantic alignment limits its ability to capture fine-grained acoustic variations, thereby constraining its performance in hierarchical sound classification. To address this issue, we propose an enhanced CLAP-based representation learning framework that integrates an auxiliary acoustic encoder branch and a gated text adaptation module, aiming to enrich acoustic details and improve task-specific text alignment. In addition, we introduce a K-nearest-neighbor (KNN)-based neighborhood-constrained inference strategy and construct an expanded training set to further improve prediction performance. Experimental results on the official DCASE 2026 evaluation set show that the proposed model achieves a hierarchical F1-score of 80.0%. Furthermore, an ensemble of complementary models within the proposed framework yields a final hierarchical F1-score of 81.1%, securing second place in the DCASE 2026 Challenge Task 1.

PDF
Abstract

Bioacoustic call-type classification relies on costly expert annotation. Active learning can reduce this burden by selecting a small batch of segments for expert annotation and using the labeled segments for training the classifier. The setting is hard: the target calls are extremely sparse and the call-type distribution is long-tailed, so a tight budget must be spent on the few rare, informative segments. We propose BADGE-Greedy-DPP, a deterministic batch selector that greedily adds the segment whose BADGE gradient embedding most enlarges the volume spanned by the batch; because this log-volume objective is submodular, the greedy rule guarantees a batch value at least a (1-1/e) fraction of the optimum of this objective, a guarantee not provided by BADGE's existing k-means++ and MCMC DPP sampling heuristics. There is also a temporal granularity mismatch in the task. The acquisition function scores whole segments, yet the informative frames inside them are few. Uniform averaging therefore washes them out. The BADGE construction naturally addresses this mismatch when applied frame-wise, as prediction residuals weight the aggregated pseudo-gradient, so confidently predicted no-call frames contribute little while a single uncertain rare-call frame can still set the segment's direction. Across 10 runs on a sparse, imbalanced hyena call-type dataset, BADGE-Greedy-DPP achieves the best overall and rare-call-type performance among all compared query strategies, including MFFT, the strongest non-BADGE baseline, and the two vanilla BADGE traversals.

PDF
Abstract

Sound event detection relies on frame-level strong labels whose annotation is expensive. Active learning addresses this problem by selecting the audio segments whose labels help the classifier most. One of the prevailing acquisition strategies for this task, mismatch-first farthest-traversal (MFFT), combines the disagreement between two classifiers and the diversity of the selected segments through hard sequential decisions. It selects whole groups of high-disagreement segments first and spreads only the remaining budget by farthest traversal. On two multi-label datasets we show that this design is blind to the similarity among the selected segments and fails under low budgets, with every mismatch-first variant ending below the plain geometric strategy it builds on. We propose mismatch-weighted facility location (MW-FL), which spends the entire budget through a disagreement-weighted coverage objective that penalizes similarity among the selected segments. The disagreement signal from MFFT is used to obtain the nonnegative weights of this facility-location objective, using fixed smoothing without dataset-specific tuning. Experiments across two geometric mechanisms with three ways of using disagreement show that coverage of the selected segments is the dominant factor, hard disagreement gating of selection is harmful on both mechanisms, and soft disagreement weighting helps on top of coverage. MW-FL attains the best area under the learning curve on both datasets.

PDF
Abstract

Sound event detection (SED) has mainly focused on the temporal extent of acoustic events, but certain applications require the detection of their time-frequency bounding boxes (BBs), which is referred to as time-frequency SED (TF-SED). In industrial audio, detecting BBs of whistle-like sounds caused by defects on a wind turbine blade is valuable because their shape in the spectrogram carries information on the damage severity. However, TF-SED remains underexplored partly because annotated datasets are scarce owing to the expertise required for annotation in bioacoustics and data confidentiality in industry. Prior work addressed TF-SED from a model-design perspective by combining an audio spectrogram transformer backbone with a detection transformer detector, while the training strategy was left unexplored. This study investigates a multi-stage transfer learning strategy that bridges modality and domain gaps and trains the model in three steps: training on image object detection data (COCO), bird audio data (BirdSet), and wind turbine audio data (WindTurbine). In the final stage, we apply component-wise progressive unfreezing, gradually expanding the set of learnable components. Our experiments revealed three findings. First, transfer learning from COCO is beneficial despite the modality gap, as the benefit of large-scale pretraining outweighs the cost of switching modality. Second, when transferring through the intermediate audio domain, keeping the backbone frozen, thereby preserving the general-purpose representation from large-scale pretraining and not specializing in the intermediate domain, yields better generalization to the target domain. Third, in the final adaptation stage, unfreezing the object query before the rest of the detector is effective because it encodes the characteristics of the target sound events. These findings provide practical guidance for efficiently training TF-SED models in other target domains using limited annotated data.

PDF