publications
publications by categories in reversed chronological order. generated by jekyll-scholar.
2026
- SpidR-Adapt: A Universal Speech Representation Model for Few-Shot AdaptationMahi Luthra, Jiayi Shen, Maxime Poli, Angelo Ortiz Tandazo, and 13 more authorsIn Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Jul 2026
Human infants, with only a few hundred hours of speech exposure, acquire basic units of new languages, highlighting a striking efficiency gap compared to the data-hungry self-supervised speech models. To address this gap, this paper introduces SpidR-Adapt for rapid adaptation of speech units to new languages using minimal unlabeled data. We cast such low-resource speech representation learning as a meta-learning problem and construct a multi-task adaptive pre-training (MAdaPT) protocol which formulates the adaptation process as a bi-level optimization framework. To enable scalable meta-training under this framework, we propose a novel heuristic solution, first-order bi-level optimization (FOBLO), avoiding heavy computation costs. Finally, we stabilize meta-training by using a robust initialization through interleaved supervision which alternates self-supervised and supervised objectives. Empirically, SpidR-Adapt achieves rapid gains in phonemic discriminability (ABX) and downstream spoken language modeling scores (sWUGGY, sBLIMP, tSC), surpassing in-domain toplines after training on less than 1h of target-language audio and delivering 100\times greater data efficiency than standard multi-task training. These findings highlight a practical, architecture-agnostic path toward biologically inspired, data-efficient representations. We open-source the training code and model checkpoints at https://github.com/facebookresearch/spidr-adapt.
@inproceedings{luthra-etal-2026-spidr, title = {{S}pid{R}-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation}, author = {Luthra, Mahi and Shen, Jiayi and Poli, Maxime and Ortiz Tandazo, Angelo and Higuchi, Yosuke and Benchekroun, Youssef and Gleize, Martin and Saint-James, Charles-Eric and Lin, Dongyan and Rust, Phillip and Villar, Angel and Parimi, Surya and Stark, Vanessa and Moritz, Rashel and Pino, Juan and LeCun, Yann and Dupoux, Emmanuel}, editor = {Liakata, Maria and Moreira, Viviane P. and Zhang, Jiajun and Jurgens, David}, booktitle = {Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)}, month = jul, year = {2026}, address = {San Diego, California, United States}, publisher = {Association for Computational Linguistics}, url = {https://aclanthology.org/2026.acl-long.1325/}, doi = {10.18653/v1/2026.acl-long.1325}, pages = {28705--28728}, isbn = {979-8-89176-390-6}, } - MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units DiscoveryAngelo Ortiz Tandazo, Manel Khentout, Youssef Benchekroun, Thomas Hueber, and 1 more authorIn Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Jul 2026
This paper introduces MauBERT, a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representation learning. We continue HuBERT pre-training with supervision based on a phonetic-to-articulatory feature mapping in 55 languages. Our models learn from multilingual data to predict articulatory features or phones, resulting in language-independent representations that capture multilingual phonetic properties. Through comprehensive ABX discriminability testing, we show MauBERT models produce more context-invariant representations than state-of-the-art multilingual self-supervised learning models. Additionally, the models effectively adapt to unseen languages and casual speech with minimal self-supervised fine-tuning (10 hours of speech). This establishes an effective approach for instilling linguistic inductive biases in self-supervised speech models.
@inproceedings{ortiz-tandazo-etal-2026-maubert, title = {{M}au{BERT}: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery}, author = {Ortiz Tandazo, Angelo and Khentout, Manel and Benchekroun, Youssef and Hueber, Thomas and Dupoux, Emmanuel}, editor = {Liakata, Maria and Moreira, Viviane P. and Zhang, Jiajun and Jurgens, David}, booktitle = {Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)}, month = jul, year = {2026}, address = {San Diego, California, United States}, publisher = {Association for Computational Linguistics}, url = {https://aclanthology.org/2026.acl-long.24/}, doi = {10.18653/v1/2026.acl-long.24}, pages = {568--585}, isbn = {979-8-89176-390-6}, } - DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech UnitsMaxime Poli, Manel Khentout, Angelo Ortiz Tandazo, Ewan Dunbar, and 2 more authors2026
@misc{poli2026discophonbenchmarkingunsuperviseddiscovery, title = {DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units}, author = {Poli, Maxime and Khentout, Manel and {Ortiz Tandazo}, Angelo and Dunbar, Ewan and Chemla, Emmanuel and Dupoux, Emmanuel}, year = {2026}, archiveprefix = {arXiv}, primaryclass = {cs.CL}, url = {https://arxiv.org/abs/2603.18612}, }
2024
- Simulating articulatory trajectories with phonological feature interpolationAngelo Ortiz Tandazo, Thomas Schatz, Thomas Hueber, and Emmanuel DupouxIn Interspeech 2024, 2024
As a first step towards a complete computational model of speech learning involving perception-production loops, we investigate the forward mapping between pseudo-motor commands and articulatory trajectories. Two phonological feature sets, based respectively on generative and articulatory phonology, are used to encode a phonetic target sequence. Different interpolation techniques are compared to generate smooth trajectories in these feature spaces, with a potential optimisation of the target value and timing to capture co-articulation effects. We report the Pearson correlation between a linear projection of the generated trajectories and articulatory data derived from a multi-speaker dataset of electromagnetic articulography (EMA) recordings. A correlation of 0.67 is obtained with an extended feature set based on generative phonology and a linear interpolation technique. We discuss the implications of our results for our understanding of the dynamics of biological motion.
@inproceedings{ortiztandazo24_interspeech, title = {Simulating articulatory trajectories with phonological feature interpolation}, author = {{Ortiz Tandazo}, Angelo and Schatz, Thomas and Hueber, Thomas and Dupoux, Emmanuel}, year = {2024}, booktitle = {Interspeech 2024}, pages = {3595--3599}, doi = {10.21437/Interspeech.2024-2192}, issn = {2958-1796}, }