Research Article

A Yoruba Language Automatic Speech Recognition System using Deep Learning Approach

by  Ogunjobi Olumide Jospeh, Agbonifo Oluwatoyin Catherine, Akinwonmi Akintoba E.
journal cover
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Issue 124
Published: July 2026
Authors: Ogunjobi Olumide Jospeh, Agbonifo Oluwatoyin Catherine, Akinwonmi Akintoba E.
10.5120/ijca0dcb3a1b147c
PDF

Ogunjobi Olumide Jospeh, Agbonifo Oluwatoyin Catherine, Akinwonmi Akintoba E. . A Yoruba Language Automatic Speech Recognition System using Deep Learning Approach. International Journal of Computer Applications. 187, 124 (July 2026), 60-65. DOI=10.5120/ijca0dcb3a1b147c

                        @article{ 10.5120/ijca0dcb3a1b147c,
                        author  = { Ogunjobi Olumide Jospeh,Agbonifo Oluwatoyin Catherine,Akinwonmi Akintoba E. },
                        title   = { A Yoruba Language Automatic Speech Recognition System using Deep Learning Approach },
                        journal = { International Journal of Computer Applications },
                        year    = { 2026 },
                        volume  = { 187 },
                        number  = { 124 },
                        pages   = { 60-65 },
                        doi     = { 10.5120/ijca0dcb3a1b147c },
                        publisher = { Foundation of Computer Science (FCS), NY, USA }
                        }
                        %0 Journal Article
                        %D 2026
                        %A Ogunjobi Olumide Jospeh
                        %A Agbonifo Oluwatoyin Catherine
                        %A Akinwonmi Akintoba E.
                        %T A Yoruba Language Automatic Speech Recognition System using Deep Learning Approach%T 
                        %J International Journal of Computer Applications
                        %V 187
                        %N 124
                        %P 60-65
                        %R 10.5120/ijca0dcb3a1b147c
                        %I Foundation of Computer Science (FCS), NY, USA
Abstract

In this paper, a Yorùbá Language Automatic Speech Recognition System capable of recognizing words spoken by the users based on preprocessed stored data was designed and implemented. Dataset from Data Science and Computational Research Laboratory, Federal University of Technology, Akure, Nigeria, were made used of. Speech feature sequences were extracted using Mel-frequency cepstral coefficients (MFCC) technique. Kaggle environment was deployed, where python programming language was used to implement the Transformer and LSTM (Long Short-Term Memory) models. The preprocessed data were fed into both models for training, both models were tested and evaluated across the following standard evaluation metrics WER, MER and CER. The results obtained from both models shows a promising approach that could be adopted for Yorùbá Language speech recognition system. The research work describes ASR for standard Yorúbà language. The source language and target language is Yorùbá language, with focus on speaking in Yorùbá language (voice data) and the output in Yorùbá language text form (text data).

References
  • IBM Cloud Education (2020). Natural Language Processing (NLP). Retrieved from https://www.ibm.com/cloud/learn/natural-language-processing
  • Brownlee, J. (2019). Deep learning for Natural Language Processing (NLP). Retrieved from https://machinelearningmastery.com/natural-langage-processing/
  • IBM Cloud Education (2020). What is Speech Recognition?. Retrieved from https://www.ibm.com/topics/speech-recognition
  • IGI Global (2021). What is ASR. Retrieved from https://www.igiglo.com/dictionary/asr/1566
  • Zajechowski, M. (2021). Automatic Speech Recognition (ASR) Software: An Introduction. Retrieved from. https://usabilitygeek.com/automatic-speech-recognition-asr-software-an-introduction/
  • Raut and Deoghare (2016). Automatic Speech Recognition and its Applications. IRJET, 03(05), 2368-2371
  • Adetunmbi, O. A., Obe, O. O., and Iyanda, J. N. (2016). Development of Standard Yorùbá speech-to-text system using HTK. International Journal of Speech Technology, 19(4), 92-94.
  • Abiola, O.B., Adeyemo O.A., Saka-Balogun O.Y., and Okesola F. (2020). A Web-Based Yorùbá to English Bilingual Lexicon for Building Technicians. International Journal of Advanced Trends in Computer Science and Engineering, 9(1), 793- 800
  • Eberhard, D. M., Gary F. S., and Fennig, C. D. (2021). Ethnologue: Languages of the World. Twenty-fourth. Retrieved from. https://www.ethnologue.com/language/yor
  • Okanlawon, J. (2016). An Analysis of the Yorùbá Language with English Phonetics, Phonology, Morphology and Syntax. Retrieved from. https://cos.northeastern.edu/wp-content/uploads/2018/09/Jolaade-Okanlawon-An-Analysis-of-Yorùbá-with-English.pdf
  • Ibiyemi, T.S., and Akintola, A.G. (2012). Automatic Speech Recognition for Telephone Voice Dialing in Yorùbá. International Journal of Engineering Research & Technology (IJERT), 1(4), 1-6
  • Zhang, Y., and Yzhang5, S. I. (2013). Speech Recognition Using Deep Learning Algorithms, pp. 1-5
  • Modupe, I. A, Sefara, T. J, and Ojo, S.O. (2019). Yorùbá Gender Recognition from Speech using Attention-based BiLSTM. In Proceedings of the First International Workshop on NLP Solutions for Under Resourced Languages (NSURL) co-located with ICNLSP. Short papers, pp. 16-22
  • Dong, L., Xu, S. and Xu, B. (2018). Speech-Transformer: A No-Recurrence Sequence-To-Sequence Model for Speech Recognition. International Conference on Acoustics, Speech and signal processing, pp. 5884-5888
Index Terms
Computer Science
Information Sciences
No index terms available.
Keywords

Automatic Speech Recognition (ASR) Mel-Frequency Cepstral Coefficients (MFCC) Speech-to-Text Low-Resource Language Deep Learning

Powered by PhDFocusTM