The Aalto system based on fine-tuned AudioSet features for DCASE2018 task2 - general purpose audio tagging

dc.contributorAalto-yliopistofi
dc.contributorAalto Universityen
dc.contributor.authorXu, Zhicunen_US
dc.contributor.authorSmit, Peteren_US
dc.contributor.authorKurimo, Mikkoen_US
dc.contributor.departmentDepartment of Signal Processing and Acousticsen
dc.contributor.groupauthorCentre of Excellence in Computational Inference, COINen
dc.contributor.groupauthorSpeech Recognitionen
dc.contributor.organizationDepartment of Signal Processing and Acousticsen_US
dc.date.accessioned2018-12-21T10:30:48Z
dc.date.available2018-12-21T10:30:48Z
dc.date.issued2018-11en_US
dc.description| openaire: EC/H2020/780069/EU//MeMAD
dc.description.abstractIn this paper, we presented a neural network system for DCASE 2018 task 2, general purpose audio tagging. We fine-tuned the Google AudioSet feature generation model with different settings for the given 41 classes on top of a fully connected layer with 100 units. Then we used the fine-tuned models to generate 128 dimensional features for each 0.960s audio. We tried different neural network structures including LSTM and multi-level attention models. In our experiments, the multi-level attention model has shown its superiority over others. Truncating the silence parts, repeating and splitting the audio into the fixed length, pitch shifting augmentation, and mixup techniques are all used in our experiments. The proposed system achieved a result with MAP@3 score at 0.936, which outperforms the baseline result of 0.704 and achieves top 8% in the public leaderboard.en
dc.description.versionPeer revieweden
dc.format.extent5
dc.format.mimetypeapplication/pdfen_US
dc.identifier.citationXu, Z, Smit, P & Kurimo, M 2018, The Aalto system based on fine-tuned AudioSet features for DCASE2018 task2 - general purpose audio tagging. in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2018 Workshop (DCASE2018)., 29, Tampere University of Technology, pp. 24-28, Detection and Classification of Acoustic Scenes and Events, Surrey, United Kingdom, 19/11/2018. < http://dcase.community/documents/workshop2018/proceedings/DCASE2018Workshop_Xu_29.pdf >en
dc.identifier.isbn978-952-15-4262-6
dc.identifier.otherPURE UUID: 5ce018ed-dfaa-4d8d-9966-9fc15f635869en_US
dc.identifier.otherPURE ITEMURL: https://research.aalto.fi/en/publications/5ce018ed-dfaa-4d8d-9966-9fc15f635869en_US
dc.identifier.otherPURE LINK: http://dcase.community/documents/workshop2018/proceedings/DCASE2018Workshop_Xu_29.pdfen_US
dc.identifier.otherPURE FILEURL: https://research.aalto.fi/files/30233157/DCASE2018Workshop_Xu_29.pdfen_US
dc.identifier.urihttps://aaltodoc.aalto.fi/handle/123456789/35661
dc.identifier.urnURN:NBN:fi:aalto-201812216670
dc.language.isoenen
dc.relationinfo:eu-repo/grantAgreement/EC/H2020/780069/EU//MeMADen_US
dc.relation.ispartofDetection and Classification of Acoustic Scenes and Eventsen
dc.relation.ispartofseriesProceedings of the Detection and Classification of Acoustic Scenes and Events 2018 Workshop (DCASE2018)en
dc.relation.ispartofseriespp. 24-28en
dc.rightsopenAccessen
dc.titleThe Aalto system based on fine-tuned AudioSet features for DCASE2018 task2 - general purpose audio taggingen
dc.typeA4 Artikkeli konferenssijulkaisussafi
dc.type.versionpublishedVersion

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
DCASE2018Workshop_Xu_29.pdf
Size:
347.3 KB
Format:
Adobe Portable Document Format