Spoken Language Identification Using Prosody, Phonotactics, and Acoustics: A Review

Thukroo, Irshad Ahmad; Bashir, Rumaan; Giri, Kaiser J.

Spoken Language Identification Using Prosody, Phonotactics, and Acoustics: A Review

Irshad Ahmad Thukroo (), Rumaan Bashir () and Kaiser J. Giri ()
Additional contact information
Irshad Ahmad Thukroo: Department of Computer Science, Islamic University of Science & Technology Kashmir, India
Rumaan Bashir: Department of Computer Science, Islamic University of Science & Technology Kashmir, India
Kaiser J. Giri: Department of Computer Science, Islamic University of Science & Technology Kashmir, India

Journal of Information & Knowledge Management (JIKM), 2022, vol. 21, issue 04, 1-45

Abstract: Spoken language identification (LID) is the identification of language present in a speech segment despite its size (duration and speed), ambiance (topic and emotion), and moderator (gender, age, demographic region). Information Technology has touched new vistas for a couple of decades mostly to simplify the day-to-day life of humans. One of the key contributions of Information Technology is the application of Artificial Intelligence to achieve better results. The advent of artificial intelligence has given rise to a new branch of Natural Language Processing (NLP) called Computational Linguistics, which generates frameworks for intelligently manipulating spoken language knowledge and has brought humanâ€“machine into a new stage. In this context, speech has arisen to be one of the imperative forms of interfaces, which is the basic mode of communication for us, and generally the most preferred one. Recognition of the spoken language is a frontend for several technologies, like multiple languages conversation systems, expressed translation software, multilingual speech recognition, spoken word extraction, speech production systems. This paper reviews and summarises the different levels of information that can be used for language identification. A broad study of acoustic, phonetic, and prosody features has been provided and various classifiers have been used for spoken language identification specifically for Indian languages. This paper has investigated various existing spoken language identification models implemented using prosodic, phonotactic, acoustic, and deep learning approaches, the datasets used, and performance measures utilized for their analysis. It also highlights the main features and challenges faced by these models. Moreover, this review analyses the efficiency of the spoken language models that can help the researchers to propose new language identification models for speech signals.

Keywords: Acoustics; deep learning; spoken language identification; prosody; phonotactics; spectral features (search for similar items in EconPapers)
Date: 2022
References: Add references at CitEc
Citations:

Downloads: (external link)
http://www.worldscientific.com/doi/abs/10.1142/S0219649222500575
Access to full text is restricted to subscribers

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:wsi:jikmxx:v:21:y:2022:i:04:n:s0219649222500575

Ordering information: This journal article can be ordered from

DOI: 10.1142/S0219649222500575

Access Statistics for this article

Journal of Information & Knowledge Management (JIKM) is currently edited by Professor Suliman Hawamdeh

More articles in Journal of Information & Knowledge Management (JIKM) from World Scientific Publishing Co. Pte. Ltd.
Bibliographic data for series maintained by Tai Tone Lim ().