An Enhanced Human Speech Based Emotion Recognition
M. Narendra and
Lankala Suvarchala
International Journal of Scientific Research in Science and Technology, 2024, vol. 11, issue 3, 518-528
Abstract:
Speech Emotion Recognition (SER) is a Machine Learning (ML) topic that has attracted substantial attention from researchers, particularly in the field of emotional computing. This is because of its growing potential, improvements in algorithms, and real-world applications. Pitch, intensity, and Mel-Frequency Cepstral Coefficients (MFCC) are examples of quantitative variables that can be used to represent the paralinguistic information found in human speech. The three main processes of data processing, feature selection/extraction, and classification based on the underlying emotional traits are typically followed to achieve SER. The use of ML techniques for SER implementation is supported by the nature of these processes as well as the unique characteristics of human speech. Several ML techniques were used in recent affective computing research projects for SER tasks; Only a few number of them, nevertheless, adequately convey the fundamental strategies and tactics that can be applied to support the three essential phases of SER implementation. Additionally, these works either overlook or just briefly explain the difficulties involved in completing these tasks and the cutting-edge methods employed to overcome them. With a focus on the three SER implementation processes, we give a comprehensive assessment of research conducted over the past ten years that tackled SER challenges from machine learning perspectives in this study. A number of difficulties are covered in detail, including the problem of Speaker-Independent experiments' low classification accuracy and related solutions. The review offers principles for SER evaluation as well, emphasizing indicators that can be experimented with and common baselines. The purpose of this paper is to serve as a a thorough manual that SER researchers may use to build SER solutions using ML techniques, inspire potential upgrades to current SER models, or spark the development of new methods to improve SER performance.
Keywords: Speech Emotion Recognition; Machine Learning; Emotional Computing,; MFCC; Actors audio dataset (search for similar items in EconPapers)
Date: 2024
References: Add references at CitEc
Citations:
Downloads: (external link)
https://ijsrst.com/home/article/view/IJSRST24113128 Abstract page (text/html)
https://ijsrst.com/home/article/download/IJSRST24113128/IJSRST24113128 Full text (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:etm:ijsrst:v11:y2024:i3:id:215
DOI: 10.32628/IJSRST24113128
Access Statistics for this article
More articles in International Journal of Scientific Research in Science and Technology from Technoscience Academy
Bibliographic data for series maintained by Pankaj Sharma ().