OPENGLOT – An open environment for the evaluation of glottal inverse filtering

dc.contributorAalto-yliopistofi
dc.contributorAalto Universityen
dc.contributor.authorAlku, Paavoen_US
dc.contributor.authorMurtola, Tiinaen_US
dc.contributor.authorMalinen, Jarmoen_US
dc.contributor.authorKuortti, Juhaen_US
dc.contributor.authorStory, Braden_US
dc.contributor.authorAiraksinen, Manuen_US
dc.contributor.authorSalmi, Mikaen_US
dc.contributor.authorVilkman, Erkkien_US
dc.contributor.authorGeneid, Ahmeden_US
dc.contributor.departmentDepartment of Signal Processing and Acousticsen
dc.contributor.departmentDepartment of Mathematics and Systems Analysisen
dc.contributor.departmentDepartment of Energy and Mechanical Engineeringen
dc.contributor.groupauthorSpeech Communication Technologyen
dc.contributor.groupauthorNumerical Analysisen
dc.contributor.groupauthorJorma Skyttä's Groupen
dc.contributor.groupauthorAdvanced Manufacturing and Materialsen
dc.contributor.organizationUniversity of Arizonaen_US
dc.contributor.organizationUniversity of Helsinkien_US
dc.date.accessioned2019-02-25T08:55:49Z
dc.date.available2019-02-25T08:55:49Z
dc.date.embargoinfo:eu-repo/date/embargoEnd/2021-02-06en_US
dc.date.issued2019-02-01en_US
dc.description.abstractGlottal inverse filtering (GIF) refers to technology to estimate the source of voiced speech, the glottal flow, from speech signals. When a new GIF algorithm is proposed, its accuracy needs to be evaluated. However, the evaluation of GIF is problematic because the ground truth, the real glottal volume velocity signal generated by the vocal folds, cannot be recorded non-invasively from natural speech. This absence of the ground truth has been circumvented in most previous GIF studies by using simple linear source-filter synthesis techniques with known artificial glottal flow models and all-pole vocal tract filters. Moreover, in a few previous studies, physical modeling of speech production has been utilized in synthesis of the test data for GIF evaluation. The evaluation strategy in previous GIF studies is, however, scattered between individual investigations and there is currently a lack of a coherent, common platform to be used in GIF evaluation. In order to address this shortcoming, the current study introduces a new environment, called OPENGLOT, for GIF evaluation. The key ideas of OPENGLOT are twofold: the environment is versatile (i.e., it provides different types of test signals for GIF evaluation) and open (i.e., the system can be used by anyone who wants to evaluate her or his new GIF method and compare it objectively to previously developed benchmark techniques). OPENGLOT consists of four main parts, Repositories I–IV, that contain data and sound synthesis software. Repository I contains a large set of synthetic glottal flow waveforms, and speech signals generated by using the Liljencrants–Fant (LF) waveform as an artificial excitation, and a digital all-pole filter to model the vocal tract. Repository II contains glottal flow and speech pressure signals generated using physical modeling of human speech production. Repository III contains pairs of glottal excitation and speech pressure signal generated by exciting 3D printed plastic vocal tract replica with LF excitations via a loudspeaker. Finally, Repository IV contains multichannel recordings (speech pressure signal, electroglottogram, high-speed video of the vocal folds) from natural production of speech. After presenting these four core parts of OPENGLOT, the article demonstrates the platform by presenting a typical use case.en
dc.description.versionPeer revieweden
dc.format.extent10
dc.format.mimetypeapplication/pdfen_US
dc.identifier.citationAlku, P, Murtola, T, Malinen, J, Kuortti, J, Story, B, Airaksinen, M, Salmi, M, Vilkman, E & Geneid, A 2019, 'OPENGLOT – An open environment for the evaluation of glottal inverse filtering', Speech Communication, vol. 107, pp. 38-47. https://doi.org/10.1016/j.specom.2019.01.005en
dc.identifier.doi10.1016/j.specom.2019.01.005en_US
dc.identifier.issn0167-6393
dc.identifier.issn1872-7182
dc.identifier.otherPURE UUID: f45fbdaf-1d99-4023-a5b4-bfd9cc4a67e7en_US
dc.identifier.otherPURE ITEMURL: https://research.aalto.fi/en/publications/f45fbdaf-1d99-4023-a5b4-bfd9cc4a67e7en_US
dc.identifier.otherPURE LINK: http://www.sciencedirect.com/science/article/pii/S0167639318303509
dc.identifier.otherPURE FILEURL: https://research.aalto.fi/files/31671883/ELEC_Alku_OPENGLOT_Speech_Communications.pdf
dc.identifier.urihttps://aaltodoc.aalto.fi/handle/123456789/36939
dc.identifier.urnURN:NBN:fi:aalto-201902252096
dc.language.isoenen
dc.publisherElsevier
dc.relation.ispartofseriesSpeech Communicationen
dc.relation.ispartofseriesVolume 107, pp. 38-47en
dc.rightsopenAccessen
dc.subject.keywordSpeech productionen_US
dc.subject.keywordGlottal flowen_US
dc.subject.keywordGlottal inverse filteringen_US
dc.subject.keywordEvaluation toolen_US
dc.titleOPENGLOT – An open environment for the evaluation of glottal inverse filteringen
dc.typeA1 Alkuperäisartikkeli tieteellisessä aikakauslehdessäfi
dc.type.versionacceptedVersion

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
ELEC_Alku_OPENGLOT_Speech_Communications.pdf
Size:
2.85 MB
Format:
Adobe Portable Document Format