aalto1 untyped-item.component.html

Automatic classification of vocal intensity category from speech

Loading...
Thumbnail Image

URL

Journal Title

Journal ISSN

Volume Title

Sähkötekniikan korkeakoulu | Master's thesis
Electronic archive copy is available via Aalto Thesis Database.

Department

Mcode

ELEC3029

Language

en

Pages

39 + 7

Series

Abstract

Vocal intensity regulation is a fundamental phenomenon in speech communication. In speech science, the term vocal intensity is referred to as the acoustic energy of speech, and it is quantified by sound pressure level (SPL). Unlike, for example, loudspeaker amplifies, which adjust the sound intensity by affecting only the gain, the regulation of intensity in speech is much more complex and challenging because it is based on the physiological speech production mechanism. The speech signal carries acoustical cues about the vocal intensity category/ SPL that the speaker used when the corresponding speech signal was produced. Due to the lack of proper calibration information in existing speech databases, it is not possible to estimate the true vocal intensity category/SPL used in recordings. In addition, there is only one previous study on the automatic classification of vocal intensity category. In this current study, a large speech database representing four vocal intensity categories (soft, normal, loud, and very loud) was recorded from 50 speakers by including calibration information. Two automatic machine learning-based classification systems were developed using Support Vector Machines (SVMs) and Convolutional Neural Networks (CNNs) and using Mel-Frequency Cepstral Coefficients (MFCCs) as features. The results show that the best classification accuracy (of about 65%) was obtained using the SVM classifier.

Description

Supervisor

Alku, Paavo

Thesis advisor

Kadiri, Sudarsana

Other note

Citation

Endorsement

Review

Supplemented By

Referenced By