
AI-Generated versus Human-Recorded Voice Input for L2 English Vowel Perception: A Classroom-Based Perception-Production Training with Korean EFL Middle School Learners
© 2026 KASELL All rights reserved
This is an open-access article distributed under the terms of the Creative Commons License, which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.
Abstract
This study examined whether AI-generated speech can effectively support classroom-based, high-variability phonetic training (HVPT)-informed pronunciation training that combines perception and production practice in L2 vowel learning, and whether its effects differ from those of human-recorded speech in terms of learning magnitude, generalization, and contrast-specific outcomes. Seventy-four Korean middle school learners of English completed 8 weeks of pronunciation training in which identification and production practice were delivered with either AI-generated or human-recorded stimuli targeting six English vowel contrasts. Phoneme identification accuracy was assessed using a two-alternative forced-choice (2AFC) task before and after training, with generalization examined across voice types (new vs. familiar), word types (trained vs. untrained), and vowel contrasts. Data were analyzed using binomial generalized linear mixed-effects models. Results showed significant pre-post improvement in both groups, with the AI-voice group demonstrating greater overall gains. Learners in both groups generalized to new voices and untrained lexical items. However, no significant three-way interactions were found for voice or word-type generalization, indicating that the pattern of transfer did not differ between training conditions. In addition, although improvement varied across vowel contrasts, the relative hierarchy of phonological difficulty was preserved across groups. These findings suggest that AI-generated speech can enhance the magnitude of phonetic learning without altering the underlying mechanisms of generalization or the organization of L2 vowel perception. Moreover, AI-generated speech appears to support category-level restructuring in ways comparable to human speech input. The greater gains observed in the AI condition may be related to differences in the variability and consistency of AI-generated input, a possibility that warrants further investigation. Overall, the results highlight the pedagogical potential of AI-generated speech as a scalable and theoretically grounded tool for L2 pronunciation training.
Keywords:
AI-generated speech, neural text-to-speech (TTS), pronunciation training, high variability phonetic training (HVPT), L2 vowel perception, generalization, Korean EFL learnersReferences
-
Al-Shami, F. and W. Cardoso. 2025. Text-to-speech in high-variability phonetic training: Focus on L2 phonological awareness. Computer-Assisted Language Learning Electronic Journal 26(6), 21-42.
[https://doi.org/10.54855/callej.123123]
- Audacity Team. 2025. Audacity (version 3.7.5) [Computer software]. Available online at https://www.audacityteam.org
-
Baayen, R. H., D. J. Davidson and D. M. Bates. 2008. Mixed-effects modelling with crossed random effects for subjects and items. Journal of Memory and Language 59(4), 390-412.
[https://doi.org/10.1016/j.jml.2007.12.005]
-
Barrington, S., E. A. Cooper and H. Farid. 2025. People are poorly equipped to detect AI-powered voice clones. Scientific Reports 15, 11004.
[https://doi.org/10.1038/s41598-025-94170-3]
-
Bates, D., M. Mächler, B. Bolker and S. Walker. 2015. Fitting linear mixed-effects models using lme4. Journal of Statistical Software 67(1), 1-48.
[https://doi.org/10.18637/jss.v067.i01]
-
Best, C. T. and M. D. Tyler. 2007. Nonnative and second-language speech perception: Commonalities and complementarities. In O.-S. Bohn and M. J. Munro, eds., Language Experience in Second Language Speech Learning: In Honor of James Emil Flege, 13-34. Amsterdam: John Benjamins.
[https://doi.org/10.1075/lllt.17.07bes]
-
Bione, T. and W. Cardoso. 2020. Synthetic voices in the foreign language context. Language Learning and Technology 24(1), 169-186.
[https://doi.org/10.64152/10125/44715]
- Boersma, P. and D. Weenink. 2022. Praat: Doing phonetics by computer (version 6.2.09) [Computer software]. Available online at http://www.praat.org
-
Bohn, O.-S. and M. J. Munro. 2007. Language Experience in Second Language Speech Learning: In Honor of James Emil Flege. Amsterdam: John Benjamins.
[https://doi.org/10.1075/lllt.17]
-
Bradlow, A. R., R. Akahane-Yamada, D. B. Pisoni and Y. Tohkura. 1999. Training Japanese listeners to identify English /r/ and /l/: Long-term retention of learning in perception and production. Perception and Psychophysics 61(5), 977-985.
[https://doi.org/10.3758/BF03206911]
-
Bradlow, A. R., D. B. Pisoni, R. Akahane-Yamada and Y. Tohkura. 1997. Training Japanese listeners to identify English /r/ and /l/: IV. Some effects of perceptual learning on speech production. Journal of the Acoustical Society of America 101(4), 2299-2310.
[https://doi.org/10.1121/1.418276]
-
Brekelmans, G., B. G. Evans and E. Wonnacott. 2025. Training child learners on nonnative vowel contrasts with phonetic training: The role of task and variability. Language Learning 75(3), 666-701.
[https://doi.org/10.1111/lang.12677]
- Carlet, A. and J. Cebrian. 2014. Training Catalan speakers to identify L2 consonants and vowels: A short-term high variability training study. Concordia Working Papers in Applied Linguistics 5, 85-98.
-
Cheng, B., X. Zhang, S. Fan and Y. Zhang. 2019. The role of temporal acoustic exaggeration in high variability phonetic training: A behavioral and ERP study. Frontiers in Psychology 10(1178), 1-28.
[https://doi.org/10.3389/fpsyg.2019.01178]
-
Choe, S., H. Park and H. Ahn. 2025. The efficacy of high variability phonetic training for L2 speech perception in EFL contexts: A meta-analytic approach. Korean Journal of English Language and Linguistics 25, 1416-1443.
[https://doi.org/10.15738/kjell.25..202510.1416]
- Davies, M. 2008-. The Corpus of Contemporary American English (COCA). Available online at https://www.english-corpora.org/coca
- Flege, J. E. 1995. Second language speech learning: Theory, findings, and problems. In W. Strange, ed., Speech Perception and Linguistic Experience: Issues in Cross-Language Research, 233-277. Timonium, MD: York Press.
-
Flege, J. E. and O.-S. Bohn. 2021. The revised Speech Learning Model (SLM-r). In R. Wayland, ed., Second Language Speech Learning: Theoretical and Empirical Progress, 3-83. Cambridge: Cambridge University Press.
[https://doi.org/10.1017/9781108886901.002]
-
Flege, J. E., O.-S. Bohn and S. Jang. 1997. Effects of experience on non-native speakers’ production and perception of English vowels. Journal of Phonetics 25(4), 437-470.
[https://doi.org/10.1006/jpho.1997.0052]
- Fox, J. and S. Weisberg. 2019. An R Companion to Applied Regression. 3rd ed. Thousand Oaks, CA: Sage.
- Hong, S. 2007. The characteristics of vowel identification errors of university-level Korean students of American English: HCA. Language and Linguistics 39, 257-277.
-
Hong, S. 2012. The relative perceptual easiness between perceptually assimilated vowels for university-level Korean learners of American English and measurement bias in an identification test. Studies in Phonetics, Phonology and Morphology 18(3), 491-511.
[https://doi.org/10.17959/sppm.2012.18.3.491]
- Hwang, H. and H. Y. Lee. 2015. The effect of high variability phonetic training on the production of English vowels and consonants. In Proceedings of 18th International Congress of Phonetic Sciences, 1041-1045.
-
Ingram, J. C. L. and S.-G. Park. 1997. Cross-language vowel perception and production by Japanese and Korean learners of English. Journal of Phonetics 25(3), 343-370.
[https://doi.org/10.1006/jpho.1997.0048]
- John, P. and W. Cardoso. 2016. A comparative study of text-to-speech and native speaker output. In Proceedings of the 4th Annual Meeting on Language Teaching, 78-96.
- Kahoot! 2025. Kahoot! [Game-based learning platform]. Available online at https://kahoot.com
- Kim, J.-E. 2010. Perception and production of English front vowels by Korean speakers. Phonetics and Speech Sciences 2(1), 51-58.
-
Kim, Y. and C. Seong. 2025. Detection of AI-generated speech using acoustic parameters. Phonetics and Speech Sciences 17(3), 15-22.
[https://doi.org/10.13064/KSSS.2025.17.3.015]
-
Lee, A. H. and R. Lyster. 2017. Can corrective feedback on perception affect production? Applied Psycholinguistics 38(2), 371-393.
[https://doi.org/10.1017/S0142716416000254]
- Lee, B., S. Lee, C. Kim, M. Ko and S. Kim. 2015. Middle School English 3. Seoul: Dong-A Publishing.
-
Lee, B., L. Plonsky and K. Saito. 2020. The effects of perception-vs. production-based pronunciation instruction. System 88, 102185.
[https://doi.org/10.1016/j.system.2019.102185]
-
Lee, H. Y. and H. Hwang. 2016. Gradient of learnability in teaching English pronunciation to Korean learners. Journal of the Acoustical Society of America 139, 1859-1872.
[https://doi.org/10.1121/1.4945716]
-
Lee, K. and M. Cho. 2015. Perception of English vowels by Korean learners: Comparisons between new and similar L2 vowel categories. The Journal of the Korea Contents Association 15(8), 579-587.
[https://doi.org/10.5392/JKCA.2015.15.08.579]
-
Lee, S. and H. Baek. 2025. The role of perceptual and acoustic similarity in learners’ perception of L2 vowels. Korean Journal of English Language and Linguistics 25, 1084-1101.
[https://doi.org/10.15738/kjell.25..202508.1084]
-
Lee, S. and M. Cho. 2018. Predicting L2 vowel identification accuracy from cross-language mappings between L2 English and L1 Korean. Language Sciences 66, 183-198.
[https://doi.org/10.1016/j.langsci.2017.09.006]
-
Lee, S. A. S. and G. K. Iverson. 2012. Vowel category formation in Korean-English bilingual children. Journal of Speech, Language, and Hearing Research 55(5), 1449-1462.
[https://doi.org/10.1044/1092-4388(2012/11-0150)]
- Lenth, R. and J. Piaskowski. 2026. emmeans: Estimated Marginal Means, aka Least-Squares Means (version 2.0.2) [R package]. Available online at https://rvlenth.github.io/emmeans
-
Leong, C. X. R., J. M. Price, N. J. Pitchford and W. J. B. van Heuven. 2018. High variability phonetic training in adaptive adverse conditions is rapid, effective, and sustained. PLoS ONE 13(10), e0204888.
[https://doi.org/10.1371/journal.pone.0204888]
-
Li, Y. 2024. A comparison of perception-based and production-based training approaches to adults’ learning of L2 sounds. Language Learning and Development 20(3), 232-248.
[https://doi.org/10.1080/15475441.2023.2285776]
-
Liakin, D., W. Cardoso and N. Liakina. 2017. The pedagogical use of mobile speech synthesis (TTS): Focus on French liaison. Computer Assisted Language Learning 30(3-4), 325-342.
[https://doi.org/10.1080/09588221.2017.1312463]
-
Lively, S. E., J. S. Logan and D. B. Pisoni. 1993. Training Japanese listeners to identify English /r/ and /l/: The role of phonetic environment and talker variability. Journal of the Acoustical Society of America 94(3), 1242-1255.
[https://doi.org/10.1121/1.408177]
-
Logan, J. S., S. E. Lively and D. B. Pisoni. 1991. Training Japanese listeners to identify English /r/ and /l/: A first report. Journal of the Acoustical Society of America 89(2), 874-886.
[https://doi.org/10.1121/1.1894649]
- LOVO. 2023. Genny [AI voice generator and text-to-speech platform]. Available online at https://lovo.ai
-
Mohammadkarimi, E. 2024. Exploring the use of artificial intelligence in promoting English language pronunciation skills. LLT Journal: A Journal on Language and Language Teaching 27(1), 98-115.
[https://doi.org/10.24071/llt.v27i1.8151]
- Moodle. 2025. Moodle [Learning management system]. Available online at https://moodle.org
-
Nagle, C. L. and M. M. Baese-Berk. 2022. Advancing the state of the art in L2 speech perception-production research: Revisiting theoretical assumptions and methodological practices. Studies in Second Language Acquisition 44(2), 580-605.
[https://doi.org/10.1017/S0272263121000371]
-
Park, M. and J. Lee. 2024. Korean ESL learners’ production of English vowel contrasts: Developmental variations in L2 sound learning. Korean Journal of English Language and Linguistics 24, 1318-1332.
[https://doi.org/10.15738/kjell.24..202412.1318]
-
Piske, T. 2007. Implications of James E. Flege’s research for the foreign language classroom. In O.-S. Bohn and M. J. Munro, eds., Language Experience in Second Language Speech Learning: In Honor of James Emil Flege, 301-314. Amsterdam: John Benjamins.
[https://doi.org/10.1075/lllt.17.26pis]
- R Core Team. 2025. R: A language and environment for statistical computing. Vienna, Austria: R Foundation for Statistical Computing [Computer software]. Available online at https://www.R-project.org
-
Ren, Y., X. Tan, T. Qin, Z. Zhao and T.-Y. Liu. 2022. Revisiting over-smoothness in text to speech. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 8197-8213.
[https://doi.org/10.18653/v1/2022.acl-long.564]
-
Sakai, M. and C. Moorman. 2018. Can perception training improve the production of second language phonemes? A meta-analytic review of 25 years of perception training research. Applied Psycholinguistics 39(1), 187-224.
[https://doi.org/10.1017/S0142716417000418]
-
Shen, J., R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang and Q. V. Le. 2018. Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions. In Proceedings of 2018 IEEE International Conference on Acoustics, Speech, and Signal Processing, 4779-4783.
[https://doi.org/10.1109/ICASSP.2018.8461368]
- Smith, G., W. Cardoso and C. García Fuentes. 2016. Text-to-speech synthesizers: Are they ready for the second language classroom? In Proceedings of the 4th Annual Meeting on Language Teaching, 97-111.
- Tan, X., T. Qin, F. Soong and T.-Y. Liu. 2021. A survey on neural speech synthesis. arXiv preprint arXiv:2106.15561, .
-
Thomson, R. I. 2012. Improving L2 listeners’ perception of English vowels: A computer-mediated approach. Language Learning 62(4), 1231-1258.
[https://doi.org/10.1111/j.1467-9922.2012.00724.x]
-
Thomson, R. I. 2018. High variability [pronunciation] training (HVPT). A proven technique about which every language teacher and learner ought to know. Journal of Second Language Pronunciation 4(2), 208-231.
[https://doi.org/10.1075/jslp.17038.tho]
-
Tsukada, K., D. Birdsong, E. Bialystok, M. Mack, H. Sung and J. E. Flege. 2005. A developmental study of English vowel production and perception by native Korean adults and children. Journal of Phonetics 33(3), 263-290.
[https://doi.org/10.1016/j.wocn.2004.10.002]
- Tyler, M. D. 2019. PAM-L2 and phonological category acquisition in the foreign language classroom. In A. M. Nyvad, M. Hejná, A. Højen, A. Jespersen, M. Hjortshøj and M. Sørensen, eds., A Sound Approach to Language Matters - In Honor of Ocke-Schwen Bohn, 607-630. Aarhus, Denmark: Aarhus University Press.
-
Uchihara, T., M. Karas and R. I. Thomson. 2025. High variability phonetic training (HVPT): A meta-analysis of L2 perceptual training studies. Studies in Second Language Acquisition 47, 794-827.
[https://doi.org/10.1017/S0272263125100879]
-
Vancova, H. 2023. AI and AI-powered tools for pronunciation training. Journal of Language and Cultural Education 11(3), 12-24.
[https://doi.org/10.2478/jolace-2023-0022]
-
Wong, J. W. S. 2026. What do the stimuli in high variability phonetic training tell us about second language perception? Current findings, implications, and the way forward. In M. Reed and J. M. Levis, eds., The Handbook of Second Language Listening, 155-168. Hoboken, NJ: John Wiley & Sons.
[https://doi.org/10.1002/9781394312375.ch12]
-
Yang, B. 1996. A comparative study of American English and Korean vowels produced by male and female speakers. Journal of Phonetics 24(2), 245-261.
[https://doi.org/10.1006/jpho.1996.0013]
-
Yang, B. 2010. College students’ production and perception of English vowels. English Language Teaching 22(4), 165-184.
[https://doi.org/10.17936/pkelt.2010.22.4.007]
-
Zhang, X., B. Cheng, D. Qin and Y. Zhang. 2021. Is talker variability a critical component of effective phonetic training for nonnative speech? Journal of Phonetics 87, 101071.
[https://doi.org/10.1016/j.wocn.2021.101071]