Fundamental Frequency Model for Postfiltering at Low Bitrates in a Transform-Domain Speech and Audio Codec
| dc.contributor | Aalto-yliopisto | fi |
| dc.contributor | Aalto University | en |
| dc.contributor.author | Das, Sneha | en_US |
| dc.contributor.author | Bäckström, Tom | en_US |
| dc.contributor.author | Fuchs, Guillaume | en_US |
| dc.contributor.department | Department of Signal Processing and Acoustics | en |
| dc.contributor.groupauthor | Speech Communication Technology | en |
| dc.contributor.groupauthor | Speech Interaction Technology | en |
| dc.contributor.organization | Fraunhofer Institute for Integrated Circuits | en_US |
| dc.date.accessioned | 2021-01-25T10:08:42Z | |
| dc.date.available | 2021-01-25T10:08:42Z | |
| dc.date.issued | 2020 | en_US |
| dc.description.abstract | Speech codecs can use postfilters to improve the quality of the decoded signal. While postfiltering is effective in reducing coding artifacts, such methods often involve processing in both the encoder and the decoder, rely on additional transmitted side information, or are highly dependent on other codec functions for optimal performance. We propose a low-complexity postfiltering method to improve the harmonic structure of the decoded signal, which models the fundamental frequency of the signal. In contrast to past approaches, the postfilter operates at the decoder as a standalone function and does not need the transmission of additional side information. It can thus be used to enhance the output of any codec. We tested the approach on a modified version of the EVS codec in TCX mode only, which is subject to more pronounced coding artefacts when used at its lowest bitrate. Listening test results show an average improvement of 7 MUSHRA points for decoded signals with the proposed harmonic postfilter. | en |
| dc.description.version | Peer reviewed | en |
| dc.format.extent | 5 | |
| dc.format.mimetype | application/pdf | en_US |
| dc.identifier.citation | Das, S, Bäckström, T & Fuchs, G 2020, Fundamental Frequency Model for Postfiltering at Low Bitrates in a Transform-Domain Speech and Audio Codec. in Proceedings of Interspeech. vol. 2020-October, Interspeech, International Speech Communication Association (ISCA), pp. 2837-2841, Interspeech, Shanghai, China, 25/10/2020. https://doi.org/10.21437/Interspeech.2020-1067 | en |
| dc.identifier.doi | 10.21437/Interspeech.2020-1067 | en_US |
| dc.identifier.issn | 1990-9772 | |
| dc.identifier.other | PURE UUID: 20eab96e-675b-450b-a81f-2a06fbdfcd13 | en_US |
| dc.identifier.other | PURE ITEMURL: https://research.aalto.fi/en/publications/20eab96e-675b-450b-a81f-2a06fbdfcd13 | en_US |
| dc.identifier.other | PURE FILEURL: https://research.aalto.fi/files/55066711/Fundamental_Frequency_Model_for_Postfiltering_at_Low_Bitrates_in_a_Transform_Domain_Speech_and_Audio_Codec.pdf | |
| dc.identifier.uri | https://aaltodoc.aalto.fi/handle/123456789/102093 | |
| dc.identifier.urn | URN:NBN:fi:aalto-202101251402 | |
| dc.language.iso | en | en |
| dc.relation.ispartof | Interspeech | en |
| dc.relation.ispartofseries | Proceedings of Interspeech | en |
| dc.relation.ispartofseries | Volume 2020-October, pp. 2837-2841 | en |
| dc.relation.ispartofseries | Interspeech | en |
| dc.rights | openAccess | en |
| dc.subject.keyword | fundamental frequency | en_US |
| dc.subject.keyword | postfiltering | en_US |
| dc.subject.keyword | speech coding | en_US |
| dc.title | Fundamental Frequency Model for Postfiltering at Low Bitrates in a Transform-Domain Speech and Audio Codec | en |
| dc.type | A4 Artikkeli konferenssijulkaisussa | fi |
| dc.type.version | publishedVersion |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- Fundamental_Frequency_Model_for_Postfiltering_at_Low_Bitrates_in_a_Transform_Domain_Speech_and_Audio_Codec.pdf
- Size:
- 944.81 KB
- Format:
- Adobe Portable Document Format