Abstract
Speech perception, the process of understanding a fluctuating sound signal as a meaningful sequence of linguistic units, is a fundamental human task that enables verbal communication. However, the neural basis of speech perception is not well understood. In particular, how top-down contextual information influences speech perception is unclear, particularly at a neural level. Furthermore, while artificial neural networks (NNs) are able to recognize speech with very high accuracy, whether they are able to serve as accurate models of human speech perception is an open question. This dissertation sought to better understand (i) the neural basis of human speech perception and in particular top-down contextual effects on speech perception and (ii) whether artificial NNs process speech in a way that is analogous to biological NNs (i.e. humans). Two main studies were conducted in this thesis. Study 1: The EEG study showed a specific brain response when listeners consciously perceived speech compared when they perceived identical stimuli as noise. These results, highlight the importance of mid-latency, negative-going responses in sensory cortex in conscious speech perception and perception, more generally. Study 2: this study created a dataset (Wordsworth) specifically designed to challenge NNs and facilitate comparisons between speech processing in humans and artificial NNs. Our results show that although artificial neural networks can predict sets of words that strictly control for semantically irrelevant features and can confuse phonetically similar words in the same way as humans do, these models cannot fully use the acoustic components of speech as a basis for discriminating semantics, nor can they transfer to sine-wave speech in the same way as humans do, suggesting they may not fully capture the complexities of human speech learning.