AI

Finetuning Strategies for Querying Sounds by Vocal Imitation

Researchers from a team of authors have published a paper describing their winning submission to the AES AIMLA 2025 Challenge. They investigated two fine-tuning strategies for querying sound effects by vocal imitation: contrastive learning with a frozen CED encoder and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. The report includes details released after the challenge.
Researchers from a team of authors have published a paper describing their winning submission to the AES AIMLA 2025 Challenge. They investigated two fine-tuning strategies for querying sound effects by vocal imitation: contrastive learning with a frozen CED encoder and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. The report includes details released after the challenge. --- Why it matters: This research is important to engineers working on audio processing tasks, as it presents two new fine-tuning strategies that can improve the performance of sound effect querying by vocal imitation. These methods can be applied to various applications such as music information retrieval and audio analysis. Source: https://arxiv.org/abs/2608.19174

This article was originally published at: https://arxiv.org/abs/2608.19174