We are a reading group meeting regularly to discuss papers on speech-based deep learning models and their use in modelling human speech processing and acquisition. We currently meet every third Thursday of the month at 3pm Amsterdam time, for informal journal club discussions and occasional invited talks. Our meetings are hybrid, with our physical meeting location at LAB42 (University of Amsterdam), but many participants joining virtually.

Anyone with an interest in the topic is very welcome to join our meetings! Contact Marianne or request to join our mailing list if you’d like to participate (please include a message describing who you are, if it might not be obvious from your e-mail address).

Next meeting(s)

Our first session of the new academic year takes place on September 17th! Greta Tuckute and Klemen Kotar will present recent work with their AuriStream model. Relevant readings:

Tuckute, G., Kotar, K., Fedorenko, E., Yamins, D. (2025). Representing Speech Through Autoregressive Prediction of Cochlear Tokens. Proc. Interspeech.

Tuckute, G., Kotar, K., Yamins, D. L., & Konkle, T. (2026). Learning Language by Listening: A Computational Learnability Account. 9th Annual Conference on Cognitive Computational Neuroscience.

The next meeting after that is planned on October 15th, with Michele Gubian presenting about compensation for tonal context.

See an archive of our past meetings below.

Archive

2026
Aug 13th
Invited talk
Mohammad Javad Ranjbar (EPFL NLP lab)
Audio and Text Understanding for Low-Resource Languages
Audio and Text Understanding for Low-Resource Languages
Mohammad Javad Ranjbar, EPFL NLP lab

Modern language and speech models still struggle with low resource languages such as Persian, where progress is limited not only by data scarcity but also by language specific, cultural, and multimodal challenges. This talk follows my work on building the resources and evaluations needed to understand and reduce these gaps.

The starting point is a large scale Persian speech corpus built from long form audiobook recordings, where making the source usable for TTS meant handling noisy segmentation, imperfect ASR, missing punctuation, speaker variation, and audio text quality filtering. Those resources then made evaluation possible, and on our Persian audio benchmark current audio language models show large gaps between text only and audio based performance, especially on culturally grounded tasks such as poetry meter, which unvowelled Persian script cannot convey at all. That gap is what my current work at EPFL is aimed at, pretraining audio understanding into an open multilingual model using mixtures of real and synthetic speech, which brings the data scarcity problem back at training scale.
June 18th
Invited talk
Maya Nachesa (University of Amsterdam)
Your Multimodal Speech Model Says I Have a Face for Radio
[arXiv preprint]
May 21st
Invited talk
Stephen McIntosh (UTokyo)
Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling [arXiv preprint]
Apr 16th
Journal club
Poli, M., Luthra, M., Benchekroun, Y., Higuchi, Y., Gleize, M., Shen, J., Algayres, R., Chung, Y., Assran, M., Pino, J., & Dupoux, E. (2025). SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision. Transactions on Machine Learning Research.
Mar 19th
Invited talk
Kwanghee Choi (UT Austin)
Self-supervised Speech Models are Phonological Vector Machines
[ACL + Interspeech preprints]
Feb 19th
Invited talk
Michaela Watkins (University of Amsterdam)
Mapping phonological features to phonetic cues: A (symbolic) Neural Network proposal for laryngeal stops in Seoul Korean using the BiPhon model
Jan 15th
Journal club
Dubiel, M., Sergeeva, A. & Leiva, L. (2024). Impact of Voice Fidelity on Decision Making: A Potential Dark Pattern? International Conference on Intelligent User Interfaces.
2025
Nov 20th
Journal club
Khorrami, K. & Räsänen, O. (2025). A model of early word acquisition based on realistic-scale audiovisual naming events. Speech Communication.
Oct 16th
Journal club
Zhang, Y., Leonard, M. K., Gwilliams, L., Bhaya-Grossman, I., & Chang, E. F. (2025). Dynamics of auditory word form encoding in human speech cortex. bioRxiv.
Sept 26th
Invited talk
Bart de Boer (VUB Brussels)
Artificial Neural Networks in the 1930s
Sept 16th
Journal club
Roll, N., Graham, C., Tatsumi, Y., Nguyen, K. T., Sumner, M., & Jurafsky, D. (2025). In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties. EMNLP.
July 2nd
Journal club
Défossez, A., Mazaré, L., Orsini, M., Royer, A., Pérez, P., Jégou, H., Grave, E. & Zeghidour, N. (2024). Moshi: a speech-text foundation model for real-time dialogue. arXiv.
Nguyen, T. A., Muller, B., Yu, B., et al. (2025). SpiRit-LM: Interleaved spoken and written language model. TACL.
June 18th
Journal club
Cruz Blandón, M.A., Gonzalez-Gomez, N., Lavechin, M., & Räsänen, O. (2025). Simulating prenatal language exposure in computational models: An exploration study. Cognition.
May 12th
Invited talk
Oli Liu (University of Edinburgh)
A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
Apr 29th
Journal club
Cho, C. J., Wu, P., Prabhune, T. S., Agarwal, D., & Anumanchipalli, G. K. (2024). Coding Speech through Vocal Tract Kinematics. IEEE Journal of Selected Topics in Signal Processing.
Apr 14th
Journal club
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., & Tagliasacchi, M. (2021). SoundStream: An End-to-End Neural Audio Codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing.
Apr 3rd
Journal club
Joint meeting with the SignLab!
Gueuwou, S., Du, X., Shakhnarovich, G., Livescu, K., & Liu, A. H. (2025). SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster Prediction. ACL.
Mar 18th
Journal club
Joint meeting with the Music Cognition Group!
Kim, G., Kim, D. K., & Jeong, H. (2024). Spontaneous emergence of rudimentary music detectors in deep neural networks. Nature Communications.
Mar 10th
Journal club
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Roblek, D., Teboul, O., Grangier, D., Tagliasacchi, M. & Zeghidour, N. (2023). AudioLM: A Language Modeling Approach to Audio Generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing.
Jan 27th
Journal club
Lavechin, M., de Seyssel, M., Métais, M., Metze, F., Mohamed, A., Bredin, H., Dupoux, E. & Cristia, A. (2024). Modeling early phonetic acquisition from child-centered audio data. Cognition.
Jan 16th
Journal club
Poli, M., Schatz, T., Dupoux, E., & Lavechin, M. (2024). Modeling the initial state of early phonetic learning in infants. Language Development Research.
2024
Dec 19th
Journal club
Fucci, D., Gaido, M., Savoldi, B., Negri, M., Cettolo, M., & Bentivogli, L. (2024). SPES: Spectrogram perturbation for explainable speech-to-text generation. arXiv.
Nov 28th
Journal club
Khorrami, K., Cruz Blandón, M. A., & Räsänen, O. (2023). Computational Insights to Acquisition of Phonemes, Words, and Word Meanings in Early Language: Sequential or Parallel Acquisition? CogSci Proceedings.
Oct 31st
Journal club
Taguchi, C., & Chiang, D. (2024). Language complexity and speech recognition accuracy: Orthographic complexity hurts, phonological complexity doesn’t. ACL.
Sept 26th
Journal club
Orhan, P., Boubenec, Y., & King, J. R. (2024). Algebraic structures emerge from the self-supervised learning of natural sounds. bioRxiv.
June 20th
Journal club
Hofer, M., Le, T. A., Levy, R., & Tenenbaum, J. (2021). Learning evolved combinatorial symbols with a neuro-symbolic generative model. arXiv. [see e-mail for an updated manuscript]
Apr 25th
Journal club
Nortje, L., Oneaţă, D., Matusevych, Y., & Kamper, H. (2024). Visually Grounded Speech Models have a Mutual Exclusivity Bias. TACL.
Mar 28th
Journal club
Pasad, A., Chien, C. M., Settle, S., & Livescu, K. (2024). What do self-supervised speech models know about words? TACL.
Feb 15th
Journal club
Algayres, R., Adi, Y., Nguyen, T. A., Copet, J., Synnaeve, G., Sagot, B., & Dupoux, E. (2023). Generative Spoken Language Model based on continuous word-sized audio tokens. EMNLP.
2023
Dec 21st
Journal club
Bartelds, M., San, N., McDonnell, B., Jurafsky, D., & Wieling, M. (2023). Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation. ACL.
Nov 16th
Journal club
Gaskell, M. G., & Marslen-Wilson, W. D. (2001). Lexical ambiguity resolution and spoken word recognition: Bridging the gap. Journal of Memory and Language.
Oct 19th
Journal club
Anderson, A. J., Davis, C., & Lalor, E. C. (2023). Context and Attention Shape Electrophysiological Correlates of Speech-to-Language Transformation. bioRxiv.
Oct 19th
Journal club
Anderson, A. J., Davis, C., & Lalor, E. C. (2023). Context and Attention Shape Electrophysiological Correlates of Speech-to-Language Transformation. bioRxiv.
Sept 21st
Journal club
Ashihara, T., Moriya, T., Matsuura, K., Tanaka, T., Ijima, Y., Asami, T., Delcroix, M., Honma, Y. (2023). SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge? Proc. Interspeech.
Martin, K., Gauthier, J., Breiss, C., Levy, R. (2023). Probing Self-supervised Speech Models for Phonetic and Phonemic Information: A Case Study in Aspiration. Proc. Interspeech.
June 14th
Journal club
Beguš, G., Zhou, A. & Zhao, T.C. Encoding of speech in convolutional layers and the brain stem based on language experience. Scientific Reports.
May 17th
Journal club
Beguš, G., Leban, A., & Gero, S. (2023). Approaching an unknown communication system by latent space exploration and causal inference. arXiv.
Apr 5th
Journal club
Adolfi, F., Bowers, J. S., & Poeppel, D. (2023). Successes and critical failures of neural networks in capturing human-like speech recognition. Neural Networks.
Mar 22nd
Journal club
Lakhotia, K., Kharitonov, E., Hsu, W. N., Adi, Y., Polyak, A., Bolte, B., Nguyen, T.-A., Copet, J., Baevski, A., Mohamed, A. & Dupoux, E. (2021). On Generative Spoken Language Modeling from Raw Audio. TACL.
Mar 9th
Journal club
Scharenborg, O., Tiesmeyer, S., Hasegawa-Johnson, M., & Dehak, N. (2018). Visualizing Phoneme Category Adaptation in Deep Neural Networks. Proc. Interspeech.
Scharenborg, O., van der Gouw, N., Larson, M., & Marchiori, E. (2019). The representation of speech in deep neural networks. MultiMedia Modeling: 25th International Conference.
Feb 23rd
Journal club
Millet, J., Chitoran, I., & Dunbar, E. (2021). Predicting non-native speech perception using the Perceptual Assimilation Model and state-of-the-art acoustic models. CoNLL.
Millet, J., & Dunbar, E. (2022). Do self-supervised speech models develop human-like perception biases? ACL.
Jan 26th
Journal club
Boersma, P., Benders, T., & Seinhorst, K. (2020). Neural network models for phonology and phonetics. Journal of Language Modelling.
Beguš, G. (2020). Generative Adversarial Phonology: Modeling Unsupervised Phonetic and Phonological Learning With Neural Networks. Frontiers in Artificial Intelligence.
2022
Dec 22nd
Journal club
Jiang, B., Dunbar, E., Sonderegger, M., Clayards, M., & Dupoux, E. (2020). Modelling Perceptual Effects of Phonology with ASR Systems. CogSci Proceedings.