Artificial intelligence-driven Sinhala auditory–verbal tutoring system for hearing-impaired children
Speech therapy for hearing-impaired children in Sri Lanka is significantly limited by factors such as cost constraints and the lack of automated, language-specific tools to aid in therapy sessions. This study presents a Sinhala auditory training platform for hearing-impaired children that combines adaptive learning with automatic pronunciation evaluation. The platform uses Bayesian knowledge tracing and multi-armed bandit task sequencing to estimate learner competence and personalize training. Its task system is based on four language comprehension task types that serve as blueprints for automatically generating activities of varying difficulty, supported by 353 audio assets. The platform covers phoneme discrimination, syllable processing, and word recognition, while providing analytics for therapists. A pronunciation evaluation module built on a fine-tuned Wav2Vec2 model delivers phoneme-level feedback using forced alignment and goodness of pronunciation scoring. To support this module, 1,534 Sinhala speech recordings from hearing-impaired speakers were collected for training. Experimental results using the model achieved a phoneme error rate of 0.289, demonstrating stable convergence during training. The system highlights the feasibility of combining adaptive tutoring and speech-based feedback for scalable, personalized auditory rehabilitation in Sinhala.
Andrijašević, A., & Vukelić, B. (2024). Generating Speech Material for Auditory Training Exercises Using ChatGPT Chatbot. In: Proceedings of the 2024 47th MIPRO ICT and Electronics Convention (MIPRO), May 20-24, 2024, Opatija, Croatia. 126-131. https://doi.org/10.1109/MIPRO60963.2024.10569423
Babu, A., Wang, C., Tjandra, A., Lakhotia, K., Xu, Q., Goyal, N., Singh, K., von Platen, P., Saraf, Y., Pino, J., Baevski, A., Conneau, A., & Auli, M. (2022). XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale. arXiv. arXiv:2111.09296. https://doi.org/10.21437/interspeech.2022-143
Baevski, A., Zhou, H., Mohamed, A., & Auli, M. (2020). wav2vec 2.0: A framework for self-supervised learning of speech representations (Version 3). Advances in neural information processing systems, 33, 12449-12460. https://doi.org/10.48550/arXiv.2006.11477
Colibaba, S., Gheorghiu, I., Colibaba, A., Ursa, O., Antoniţă, C., & Cîrşmari, R. (2022, November 17–19). The Voice Project: Habilitating hearing-impaired children to recover hearing and lead a normal life. In: Proceedings of the 2022 E-Health and Bioengineering Conference (EHB), Online. 1-4. https://doi.org/10.1109/EHB55594.2022.9991663
Corbett, A. T., & Anderson, J. R. (1994). Knowledge tracing: Modeling the acquisition of procedural knowledge. User Modelling and User-Adapted Interaction, 4(4), 253-278. https://doi.org/10.1007/BF01099821
Fernando, S., Leo, C., Tharushika, C., Mawaththa, R., Perera, J., Vidhanaarachchi, S., & Kasthurirathna, D. (2024). Assisting Hearing Impaired Children by Personalized Gamification of Auditory Verbal Therapy. In: Proceedings of the 2024 6th International Conference on Advancements in Computing (ICAC), December 12-13, 2024, Colombo, Sri Lanka. 414-419. https://doi.org/10.1109/ICAC64487.2024.10850970
Gnadlinger, F., Werminghaus, M., Selmanagić, A., Filla, T., Richter, J. G., Kriglstein, S., & Klenzner, T. (2024). Incorporating an intelligent tutoring system into a game-based auditory rehabilitation training for adult cochlear implant recipients: Algorithm development and validation. JMIR Serious Games, 12(1), e55231. https://doi.org/10.2196/55231
Graves, A., Fernández, S., Gomez, F., & Schmidhuber, J. (2006). Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In: Proceedings of the 23rd International Conference on Machine Learning, June 25-29, 2006, Pittsburgh, PA. 369-376 https://doi.org/10.1145/1143844.1143891
Ishihara, M., & Tsuda, M. (2021). Hearing support system for the hearing impaired. In: Proceedings of the 2021 IEEE 3rd Global Conference on Life Sciences and Technologies (LifeTech), March 9-11, 2021, Nara, Japan. 473-474. https://doi.org/10.1109/LifeTech52111.2021.9391875
Lee, W. S., & Tsoi, I. C. Y. (2021, January 24-27). Acoustical Characteristics of the Cantonese Vowels and Tones Produced by Hearing Impaired Speakers. In: Proceedings of the 2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP), Online. 1-5. https://doi.org/10.1109/ISCSLP49672.2021.9362118
Liu, C., & Fu, Q. J. (2007). Estimation of vowel recognition with cochlear implant simulations. IEEE Transactions on Biomedical Engineering, 54(1), 74-81. https://doi.org/10.1109/TBME.2006.883800
Morris, A. C., Maier, V., & Green, P. D. (2004). From WER and RIL to MER and WIL: improved evaluation measures for connected speech recognition. In: Proceedings of InterSpeech 2004, October 4-8, 2004, Jeju Island, Korea. 2765–2768. https://doi.org/10.21437/Interspeech.2004-668
Nonis, P. D. M., & Hettiarachchi, S. (2014). A study of phonetic and phonological development of Sinhala speaking children in the Puttalam District aged 3; 0-3; 11 years. Journal of the Faculty of Graduate Studies, 3, 8-25. https://doi.org/10.13140/RG.2.1.1152.0486
Segal, A., Ben David, Y., Williams, J. J., Gal, K., & Shalom, Y. (2018). Combining difficulty ranking with multi-armed bandits to sequence educational content. In: Proceedings of the International Conference on Artificial Intelligence in Education, June 27-30, 2018, London, United Kingdom. 317-321). https://doi.org/10.1007/978-3-319-93846-2_59
Seo, D. G. (2017). Overview and current management of computerized adaptive testing in licensing/certification examinations. Journal of Educational Evaluation for Health Professions, 14, 17. https://doi.org/10.3352/jeehp.2017.14.17
van De Sande, B. (2013). Properties Of The Bayesian Knowledge Tracing Model. Journal of Educational Data Mining, 5(2), 1-10. https://doi.org/10.5281/zenodo.3554629
Weerasinghe, R., Wasala, A., & Gamage, K. (2005). A rule based syllabification algorithm for Sinhala. In: Proceedings of the International Conference on Natural Language Processing, October 11-13, 2005, Jeju Island, Korea. 438-449. https://doi.org/10.1007/11562214_39
Witt, S. M., & Young, S. J. (2000). Phone-level pronunciation scoring and assessment for interactive language learning. Speech Communication, 30(2-3), 95-108. https://doi.org/10.1016/S0167-6393(99)00044-8
Yang, M., Hirschi, K., Looney, S. D., Kang, O., & Hansen, J. H. L. (2022). Improving Mispronunciation Detection with Wav2vec2-based Momentum Pseudo-Labeling for Accentedness and Intelligibility Assessment. In: Proceedings of Interspeech 2022, September 18-22, 2022, Incheon, Korea. 4481–4485. https://doi.org/10.21437/interspeech.2022-11039
Yong-Xin, L., Shuang, L., & De-Min, H. (2007). Research advances in post-operative rehabilitation following cochlear implant. Journal of Otology, 2(2), 92-96. https://doi.org/10.1016/S1672-2930(07)50019-X
