<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>(D)NN-Speech Reading Group</title><link>https://dnn-speech.github.io/</link><description>Recent content on (D)NN-Speech Reading Group</description><generator>Hugo -- gohugo.io</generator><language>en-US</language><managingEditor>mariannedhk@gmail.com (Marianne de Heer Kloots)</managingEditor><webMaster>mariannedhk@gmail.com (Marianne de Heer Kloots)</webMaster><copyright>Marianne de Heer Kloots</copyright><atom:link href="https://dnn-speech.github.io/index.xml" rel="self" type="application/rss+xml"/><item><title>(D)NN-speech reading group</title><link>https://dnn-speech.github.io/reading-group/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><author>mariannedhk@gmail.com (Marianne de Heer Kloots)</author><guid>https://dnn-speech.github.io/reading-group/</guid><description>&lt;p&gt;We are a reading group meeting regularly to discuss papers on speech-based deep learning models and their use in modelling human speech processing and acquisition. We currently meet &lt;strong&gt;every third Thursday of the month&lt;/strong&gt; at 3pm Amsterdam time, for informal journal club discussions and occasional invited talks. Our meetings are hybrid, with our physical meeting location at &lt;a href="https://lab42.uva.nl/"&gt;LAB42&lt;/a&gt; (University of Amsterdam), but many participants joining virtually.&lt;/p&gt;
&lt;p&gt;Anyone with an interest in the topic is very welcome to join our meetings! Contact &lt;a href="https://mdhk.net/"&gt;Marianne&lt;/a&gt; or &lt;a href="https://groups.google.com/g/dnn-speech"&gt;request to join our mailing list&lt;/a&gt; if you&amp;rsquo;d like to participate (please include a message describing who you are, if it might not be obvious from your e-mail address).&lt;/p&gt;</description><content:encoded><![CDATA[<p>We are a reading group meeting regularly to discuss papers on speech-based deep learning models and their use in modelling human speech processing and acquisition. We currently meet <strong>every third Thursday of the month</strong> at 3pm Amsterdam time, for informal journal club discussions and occasional invited talks. Our meetings are hybrid, with our physical meeting location at <a href="https://lab42.uva.nl/">LAB42</a> (University of Amsterdam), but many participants joining virtually.</p>
<p>Anyone with an interest in the topic is very welcome to join our meetings! Contact <a href="https://mdhk.net/">Marianne</a> or <a href="https://groups.google.com/g/dnn-speech">request to join our mailing list</a> if you&rsquo;d like to participate (please include a message describing who you are, if it might not be obvious from your e-mail address).</p>
<h4 id="next-meetings">Next meeting(s)</h4>
<p>Our first session of the new academic year takes place on <strong>September 17th</strong>! <a href="http://www.tuckute.com/">Greta Tuckute</a> and <a href="https://klemenkotar.github.io/">Klemen Kotar</a> will present recent work with their <em>AuriStream</em> model. Relevant readings:</p>
<blockquote>
<p><span style="font-size:12pt">Tuckute, G., Kotar, K., Fedorenko, E., Yamins, D. (2025). <a href="http://doi.org/10.21437/Interspeech.2025-2044">Representing Speech Through Autoregressive Prediction of Cochlear Tokens</a>. <em>Proc. Interspeech</em>.</span></p>
</blockquote>
<blockquote>
<p><span style="font-size:12pt">Tuckute, G., Kotar, K., Yamins, D. L., &amp; Konkle, T. (2026). <a href="https://openreview.net/forum?id=Z4VJEMX8pS">Learning Language by Listening: A Computational Learnability Account</a>. <em>9th Annual Conference on Cognitive Computational Neuroscience</em>.</span></p>
</blockquote>
<p>The next meeting after that is planned on <strong>October 15th</strong>, with <a href="https://github.com/uasolo">Michele Gubian</a> presenting about <a href="https://arxiv.org/abs/2606.17835">compensation for tonal context</a>.</p>
<p>See an archive of our past meetings below.</p>
<h4 id="archive">Archive</h4>


<script src="https://kit.fontawesome.com/2502d9ad4b.js" crossorigin="anonymous"></script>
<div class="wrapper">
<table border="0">
  <tr  class="header">
      <th colspan="2"><span>▾</span> 2026</th>
  </tr>
  <tr>
    <td><b>Aug 13th</b><br>Invited talk</td>
    <td>
    <a href="https://mohammadjranjbar.github.io/">Mohammad Javad Ranjbar</a> (EPFL NLP lab) <br>
    <i>Audio and Text Understanding for Low-Resource Languages</i><br> <span class="inline-details"><input type="checkbox" id="inline-details-0" class="inline-details-toggle"><label for="inline-details-0"
        class="inline-details-summary">See abstract</label><span class="inline-details-content"><b>Audio and Text Understanding for Low-Resource Languages</b><br>
Mohammad Javad Ranjbar, EPFL NLP lab<br><br>
Modern language and speech models still struggle with low resource languages such as Persian, where progress is limited not only by data scarcity but also by language specific, cultural, and multimodal challenges. This talk follows my work on building the resources and evaluations needed to understand and reduce these gaps. <br><br>
The starting point is a large scale Persian speech corpus built from long form audiobook recordings, where making the source usable for TTS meant handling noisy segmentation, imperfect ASR, missing punctuation, speaker variation, and audio text quality filtering. Those resources then made evaluation possible, and on our Persian audio benchmark current audio language models show large gaps between text only and audio based performance, especially on culturally grounded tasks such as poetry meter, which unvowelled Persian script cannot convey at all. That gap is what my current work at EPFL is aimed at, pretraining audio understanding into an open multilingual model using mixtures of real and synthetic speech, which brings the data scarcity problem back at training scale.</span></span>

    </td>
  </tr>
  <tr>
    <td><b>June 18th</b><br>Invited talk</td>
    <td>
    <a href="https://www.linkedin.com/in/mknachesa/">Maya Nachesa</a> (University of Amsterdam) <br>
    <i>Your Multimodal Speech Model Says I Have a Face for Radio</i><br> [<a href="https://arxiv.org/abs/2605.30472">arXiv preprint</a>]
    </td>
  </tr>
  <tr>
    <td><b>May 21st</b><br>Invited talk <a href="https://DNN-speech.github.io/files/DNNspeech_260521_StephenMcIntosh_Kanade.pdf"><i class="fa-solid fa-file"></i></a></td>
    <td>
    <a href="https://www.linkedin.com/in/stephen-mcintosh-4566a6136/">Stephen McIntosh</a> (UTokyo) <br>
    <i>Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling</i> [<a href="https://arxiv.org/abs/2602.00594">arXiv preprint</a>]
    </td>
  </tr>
  <tr>
    <td><b>Apr 16th</b><br>Journal club</td>
    <td>
    Poli, M., Luthra, M., Benchekroun, Y., Higuchi, Y., Gleize, M., Shen, J., Algayres, R., Chung, Y., Assran, M., Pino, J., & Dupoux, E. (2025). <a href="https://openreview.net/forum?id=E7XAFBpfZs">SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision</a>. <i>Transactions on Machine Learning Research</i>.
    </td>
  </tr>
  <tr>
    <td><b>Mar 19th</b><br>Invited talk <a href="https://www.youtube.com/watch?v=DtFYKvNo9IQ"><i class="fa-solid fa-video fa-sm"></i></a></td>
    <td>
    <a href="https://kwangheechoi.notion.site/Kwanghee-Choi-9250bb79f1664e5b8e8c5553adf20068">Kwanghee Choi</a> (UT Austin) <br>
    <i>Self-supervised Speech Models are Phonological Vector Machines</i> <br>
    [<a href="https://arxiv.org/abs/2602.18899">ACL</a> + <a href="https://arxiv.org/abs/2603.12642">Interspeech</a> preprints]
    </td>
  </tr>
  <tr>
    <td><b>Feb 19th</b><br>Invited talk</td>
    <td>
    <a href="https://sites.google.com/view/michaelawatkins">Michaela Watkins</a> (University of Amsterdam)<br>
    <i>Mapping phonological features to phonetic cues: A (symbolic) Neural Network proposal for laryngeal stops in Seoul Korean using the BiPhon model</i>
    </td>
  </tr>
  <tr>
    <td><b>Jan 15th</b><br>Journal club</td>
    <td> Dubiel, M., Sergeeva, A. & Leiva, L. (2024). <a href="https://doi.org/10.1145/3640543.3645202">Impact of Voice Fidelity on Decision Making: A Potential Dark Pattern? </a><i>International Conference on Intelligent User Interfaces</i>.
    </td>
  </tr>
  <tr class="header">
    <th colspan="2"><span>▾</span> 2025</th>
  </tr>
  <tr>
    <td><b>Nov 20th</b><br>Journal club</td>
    <td>
    Khorrami, K. & Räsänen, O. (2025). <a href="https://doi.org/10.1016/j.specom.2024.103169">A model of early word acquisition based on realistic-scale audiovisual naming events</a>. <i>Speech Communication</i>.
    </td>
  </tr>
  <tr>
    <td><b>Oct 16th</b><br>Journal club</td>
    <td> Zhang, Y., Leonard, M. K., Gwilliams, L., Bhaya-Grossman, I., & Chang, E. F. (2025). <a href="https://www.biorxiv.org/content/10.1101/2025.05.05.651964v1.full">Dynamics of auditory word form encoding in human speech cortex</a>. <i>bioRxiv</i>.
    </td>
  </tr>
  <tr>
    <td><b>Sept 26th</b><br>Invited talk</td>
    <td> <a href="https://ai.vub.ac.be/team/bart-de-boer/">Bart de Boer</a> (VUB Brussels) <br>
    <i>Artificial Neural Networks in the 1930s</i>
    </td>
  </tr>
  <tr>
    <td><b>Sept 16th</b><br>Journal club</td>
    <td> Roll, N., Graham, C., Tatsumi, Y., Nguyen, K. T., Sumner, M., & Jurafsky, D. (2025). <a href="https://arxiv.org/abs/2505.14887">In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties</a>. <i>EMNLP</i>.
    </td>
  </tr>
  <tr>
    <td><b>July 2nd</b><br>Journal club</td>
    <td>
    Défossez, A., Mazaré, L., Orsini, M., Royer, A., Pérez, P., Jégou, H., Grave, E. & Zeghidour, N. (2024). <a href="https://arxiv.org/abs/2410.00037">Moshi: a speech-text foundation model for real-time dialogue</a>. <i>arXiv</i>. <br>
    Nguyen, T. A., Muller, B., Yu, B., et al. (2025). <a href="https://aclanthology.org/2025.tacl-1.2/">SpiRit-LM: Interleaved spoken and written language model</a>. <i>TACL</i>.
    </td>
  </tr>
  <tr>
    <td><b>June 18th</b><br>Journal club</td>
    <td>
    Cruz Blandón, M.A., Gonzalez-Gomez, N., Lavechin, M., & Räsänen, O. (2025). <a href="https://doi.org/10.1016/j.cognition.2024.106044">Simulating prenatal language exposure in computational models: An exploration study</a>. <i>Cognition</i>.
    </td>
  </tr>
  <tr>
    <td><b>May 12th</b><br>Invited talk</td>
    <td>
    <a href="https://uililo.github.io/">Oli Liu</a> (University of Edinburgh)<br>
    <i>A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech</i>
    </td>
  </tr>
  <tr>
    <td><b>Apr 29th</b><br>Journal club</td>
    <td>
    Cho, C. J., Wu, P., Prabhune, T. S., Agarwal, D., & Anumanchipalli, G. K. (2024). <a href="https://doi.org/10.1109/JSTSP.2024.3497655">Coding Speech through Vocal Tract Kinematics</a>. <i>IEEE Journal of Selected Topics in Signal Processing</i>.
    </td>
  </tr>
  <tr>
    <td><b>Apr 14th</b><br>Journal club</td>
    <td>
    Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., & Tagliasacchi, M. (2021). <a href="https://arxiv.org/abs/2107.03312">SoundStream: An End-to-End Neural Audio Codec</a>. <i>IEEE/ACM Transactions on Audio, Speech, and Language Processing</i>.
    </td>
  </tr>
  <tr>
    <td><b>Apr 3rd</b><br>Journal club</td>
    <td>
    <b>Joint meeting with the <a href="https://www.signlab-amsterdam.nl/">SignLab</a>!</b> <br>
    Gueuwou, S., Du, X., Shakhnarovich, G., Livescu, K., & Liu, A. H. (2025). <a href="https://arxiv.org/abs/2411.16765">SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster Prediction</a>. <i>ACL</i>.
    </td>
  </tr>
  <tr>
    <td><b>Mar 18th</b><br>Journal club</td>
    <td>
    <b>Joint meeting with the <a href="https://musicreadinggroup.wordpress.com/">Music Cognition Group</a>!</b> <br>
    Kim, G., Kim, D. K., & Jeong, H. (2024). <a href="http://doi.org/10.1038/s41467-023-44516-0">Spontaneous emergence of rudimentary music detectors in deep neural networks</a>. <i>Nature Communications</i>.
    </td>
  </tr>
  <tr>
    <td><b>Mar 10th</b><br>Journal club</td>
    <td>
  Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Roblek, D., Teboul, O., Grangier, D., Tagliasacchi, M. & Zeghidour, N. (2023). <a href="http://doi.org/10.1109/TASLP.2023.3288409">AudioLM: A Language Modeling Approach to Audio Generation</a>. <i>IEEE/ACM Transactions on Audio, Speech, and Language Processing</i>.
    </td>
  </tr>
  <tr>
    <td><b>Jan 27th</b><br>Journal club</td>
    <td>
    Lavechin, M., de Seyssel, M., Métais, M., Metze, F., Mohamed, A., Bredin, H., Dupoux, E. & Cristia, A. (2024). <a href="https://doi.org/10.1016/j.cognition.2024.105734">Modeling early phonetic acquisition from child-centered audio data</a>. <i>Cognition</i>.
    </td>
  </tr>
  <tr>
    <td><b>Jan 16th</b><br>Journal club</td>
    <td>
    Poli, M., Schatz, T., Dupoux, E., & Lavechin, M. (2024). <a href="https://ldr.lps.library.cmu.edu/article/717/galley/581/view/">Modeling the initial state of early phonetic learning in infants</a>. <i>Language Development Research</i>.
    </td>
  </tr>
<tr class="header">
    <th colspan="2"><span>▾</span> 2024</th>
</tr>
  <tr>
    <td><b>Dec 19th</b><br>Journal club</td>
    <td>
    Fucci, D., Gaido, M., Savoldi, B., Negri, M., Cettolo, M., & Bentivogli, L. (2024). <a href="https://arxiv.org/abs/2411.01710">SPES: Spectrogram perturbation for explainable speech-to-text generation</a>. <i>arXiv</i>.
    </td>
  </tr>
  <tr>
    <td><b>Nov 28th</b><br>Journal club</td>
    <td>
    Khorrami, K., Cruz Blandón, M. A., & Räsänen, O. (2023). <a href="https://escholarship.org/uc/item/79t028n8">Computational Insights to Acquisition of Phonemes, Words, and Word Meanings in Early Language: Sequential or Parallel Acquisition?</a> <i>CogSci Proceedings</i>.
    </td>
  </tr>
  <tr>
    <td><b>Oct 31st</b><br>Journal club</td>
    <td>
    Taguchi, C., & Chiang, D. (2024). <a href="https://aclanthology.org/2024.acl-long.827/">Language complexity and speech recognition accuracy: Orthographic complexity hurts, phonological complexity doesn’t</a>. <i>ACL</i>.
    </td>
  </tr>
  <tr>
    <td><b>Sept 26th</b><br>Journal club</td>
    <td>
    Orhan, P., Boubenec, Y., & King, J. R. (2024). <a href="https://doi.org/10.1101/2024.03.13.584776">Algebraic structures emerge from the self-supervised learning of natural sounds</a>. <i>bioRxiv</i>.
    </td>
  </tr>
  <tr>
    <td><b>June 20th</b><br>Journal club</td>
    <td>
    Hofer, M., Le, T. A., Levy, R., & Tenenbaum, J. (2021). <a href="https://arxiv.org/abs/2104.08274">Learning evolved combinatorial symbols with a neuro-symbolic generative model</a>. <i>arXiv</i>. [see e-mail for an updated manuscript]
    </td>
  </tr>
  <tr>
    <td><b>Apr 25th</b><br>Journal club</td>
    <td>
    Nortje, L., Oneaţă, D., Matusevych, Y., & Kamper, H. (2024). <a href="https://arxiv.org/abs/2403.13922">Visually Grounded Speech Models have a Mutual Exclusivity Bias</a>. <i>TACL</i>.
    </td>
  </tr>
  <tr>
    <td><b>Mar 28th</b><br>Journal club</td>
    <td>
    Pasad, A., Chien, C. M., Settle, S., & Livescu, K. (2024). <a href="https://arxiv.org/abs/2307.00162">What do self-supervised speech models know about words?</a> <i>TACL</i>.
    </td>
  </tr>
  <tr>
    <td><b>Feb 15th</b><br>Journal club</td>
    <td>
    Algayres, R., Adi, Y., Nguyen, T. A., Copet, J., Synnaeve, G., Sagot, B., & Dupoux, E. (2023). <a href="https://aclanthology.org/2023.emnlp-main.182/">Generative Spoken Language Model based on continuous word-sized audio tokens</a>. <i>EMNLP</i>.
    </td>
  </tr>
<tr class="header">
    <th colspan="2"><span>▾</span> 2023</th>
</tr>
  <tr>
    <td><b>Dec 21st</b><br>Journal club</td>
    <td>
    Bartelds, M., San, N., McDonnell, B., Jurafsky, D., & Wieling, M. (2023). <a href="https://aclanthology.org/2023.acl-long.42/">Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation</a>. <i>ACL</i>.
    </td>
  </tr>
  <tr>
    <td><b>Nov 16th</b><br>Journal club</td>
    <td>
    Gaskell, M. G., & Marslen-Wilson, W. D. (2001). <a href="https://doi.org/10.1006/jmla.2000.2741">Lexical ambiguity resolution and spoken word recognition: Bridging the gap</a>. <i>Journal of Memory and Language</i>.
    </td>
  </tr>
  <tr>
    <td><b>Oct 19th</b><br>Journal club</td>
    <td>
    Anderson, A. J., Davis, C., & Lalor, E. C. (2023). <a href="https://doi.org/10.1101/2023.09.24.559177">Context and Attention Shape Electrophysiological Correlates of Speech-to-Language Transformation</a>. <i>bioRxiv</i>.
    </td>
  </tr>
  <tr>
    <td><b>Oct 19th</b><br>Journal club</td>
    <td>
    Anderson, A. J., Davis, C., & Lalor, E. C. (2023). <a href="https://doi.org/10.1101/2023.09.24.559177">Context and Attention Shape Electrophysiological Correlates of Speech-to-Language Transformation</a>. <i>bioRxiv</i>.
    </td>
  </tr>
  <tr>
    <td><b>Sept 21st</b><br>Journal club</td>
    <td>
    Ashihara, T., Moriya, T., Matsuura, K., Tanaka, T., Ijima, Y., Asami, T., Delcroix, M., Honma, Y. (2023). <a href="https://doi.org/10.21437/Interspeech.2023-1823">SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?</a> <i>Proc. Interspeech</i>. <br>
    Martin, K., Gauthier, J., Breiss, C., Levy, R. (2023). <a href="https://doi.org/10.21437/Interspeech.2023-235">Probing Self-supervised Speech Models for Phonetic and Phonemic Information: A Case Study in Aspiration</a>. <i>Proc. Interspeech</i>.
    </td>
  </tr>
  <tr>
    <td><b>June 14th</b><br>Journal club</td>
    <td>
    Beguš, G., Zhou, A. & Zhao, T.C. <a href="https://doi.org/10.1038/s41598-023-33384-9">Encoding of speech in convolutional layers and the brain stem based on language experience</a>. <i>Scientific Reports</i>.
    </td>
  </tr>
  <tr>
    <td><b>May 17th</b><br>Journal club</td>
    <td>
    Beguš, G., Leban, A., & Gero, S. (2023). <a href="https://arxiv.org/abs/2303.10931">Approaching an unknown communication system by latent space exploration and causal inference</a>. <i>arXiv</i>.
    </td>
  </tr>
  <tr>
    <td><b>Apr 5th</b><br>Journal club</td>
    <td>
    Adolfi, F., Bowers, J. S., & Poeppel, D. (2023). <a href="https://doi.org/10.1016/j.neunet.2023.02.032">Successes and critical failures of neural networks in capturing human-like speech recognition</a>. <i>Neural Networks</i>.
    </td>
  </tr>
  <tr>
    <td><b>Mar 22nd</b><br>Journal club</td>
    <td>
    Lakhotia, K., Kharitonov, E., Hsu, W. N., Adi, Y., Polyak, A., Bolte, B., Nguyen, T.-A., Copet, J., Baevski, A., Mohamed, A. & Dupoux, E. (2021). <a href="https://doi.org/10.1162/tacl_a_00430">On Generative Spoken Language Modeling from Raw Audio</a>. <i>TACL</i>.
    </td>
  </tr>
  <tr>
    <td><b>Mar 9th</b><br>Journal club</td>
    <td>
    Scharenborg, O., Tiesmeyer, S., Hasegawa-Johnson, M., & Dehak, N. (2018). <a href="https://doi.org/10.21437/Interspeech.2018-1707">Visualizing Phoneme Category Adaptation in Deep Neural Networks</a>. <i>Proc. Interspeech</i>. <br>
    Scharenborg, O., van der Gouw, N., Larson, M., & Marchiori, E. (2019). <a href="https://doi.org/10.1007/978-3-030-05716-9_16">The representation of speech in deep neural networks</a>. <i>MultiMedia Modeling: 25th International Conference</i>.
    </td>
  </tr>
  <tr>
    <td><b>Feb 23rd</b><br>Journal club</td>
    <td>
    Millet, J., Chitoran, I., & Dunbar, E. (2021). <a href="https://doi.org/10.18653/v1/2021.conll-1.51">Predicting non-native speech perception using the Perceptual Assimilation Model and state-of-the-art acoustic models</a>. <i>CoNLL</i>. <br>
    Millet, J., & Dunbar, E. (2022). <a href="https://aclanthology.org/2022.acl-long.523/">Do self-supervised speech models develop human-like perception biases?</a> <i>ACL</i>.
    </td>
  </tr>
  <tr>
    <td><b>Jan 26th</b><br>Journal club</td>
    <td>
    Boersma, P., Benders, T., & Seinhorst, K. (2020). <a href="https://doi.org/10.15398/jlm.v8i1.224">Neural network models for phonology and phonetics</a>. <i>Journal of Language Modelling</i>. <br>
    Beguš, G. (2020). <a href="https://doi.org/10.3389/frai.2020.00044">Generative Adversarial Phonology: Modeling Unsupervised Phonetic and Phonological Learning With Neural Networks</a>. <i>Frontiers in Artificial Intelligence</i>.
    </td>
  </tr>
<tr class="header">
    <th colspan="2"><span>▾</span> 2022</th>
</tr>
  <tr>
    <td><b>Dec 22nd</b><br>Journal club</td>
    <td>
    Jiang, B., Dunbar, E., Sonderegger, M., Clayards, M., & Dupoux, E. (2020). <a href="https://cognitivesciencesociety.org/cogsci20/papers/0669/0669.pdf">Modelling Perceptual Effects of Phonology with ASR Systems</a>. <i>CogSci Proceedings</i>.
    </td>
  </tr>
</table>
</div>

]]></content:encoded></item></channel></rss>