aalto1 untyped-item.component.html
Automatic classification of vocal intensity category from speech
Loading...
URL
Journal Title
Journal ISSN
Volume Title
Sähkötekniikan korkeakoulu |
Master's thesis
Electronic archive copy is available via Aalto Thesis Database.
Authors
Date
Department
Major/Subject
Mcode
ELEC3029
Language
en
Pages
39 + 7
Series
Abstract
Vocal intensity regulation is a fundamental phenomenon in speech communication. In speech science, the term vocal intensity is referred to as the acoustic energy of speech, and it is quantified by sound pressure level (SPL). Unlike, for example, loudspeaker amplifies, which adjust the sound intensity by affecting only the gain, the regulation of intensity in speech is much more complex and challenging because it is based on the physiological speech production mechanism. The speech signal carries acoustical cues about the vocal intensity category/ SPL that the speaker used when the corresponding speech signal was produced. Due to the lack of proper calibration information in existing speech databases, it is not possible to estimate the true vocal intensity category/SPL used in recordings. In addition, there is only one previous study on the automatic classification of vocal intensity category. In this current study, a large speech database representing four vocal intensity categories (soft, normal, loud, and very loud) was recorded from 50 speakers by including calibration information. Two automatic machine learning-based classification systems were developed using Support Vector Machines (SVMs) and Convolutional Neural Networks (CNNs) and using Mel-Frequency Cepstral Coefficients (MFCCs) as features. The results show that the best classification accuracy (of about 65%) was obtained using the SVM classifier.