Towards Memory-Efficient Training for Extremely Large Output Spaces : Learning with 670k Labels on a Single Commodity GPU
| dc.contributor | Aalto-yliopisto | fi |
| dc.contributor | Aalto University | en |
| dc.contributor.author | Schultheis, Erik | en_US |
| dc.contributor.author | Babbar, Rohit | en_US |
| dc.contributor.department | Department of Computer Science | en |
| dc.contributor.editor | Koutra, Danai | en_US |
| dc.contributor.editor | Plant, Claudia | en_US |
| dc.contributor.editor | Gomez Rodriguez, Manuel | en_US |
| dc.contributor.editor | Baralis, Elena | en_US |
| dc.contributor.editor | Bonchi, Francesco | en_US |
| dc.contributor.groupauthor | Babbar Rohit group | en |
| dc.contributor.groupauthor | Computer Science Professors | en |
| dc.contributor.groupauthor | Computer Science - Artificial Intelligence and Machine Learning (AIML) - Research area | en |
| dc.date.accessioned | 2024-08-06T07:41:52Z | |
| dc.date.available | 2024-08-06T07:41:52Z | |
| dc.date.issued | 2023 | en_US |
| dc.description | Publisher Copyright: © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG. | |
| dc.description.abstract | In classification problems with large output spaces (up to millions of labels), the last layer can require an enormous amount of memory. Using sparse connectivity would drastically reduce the memory requirements, but as we show below, applied naïvely it can result in much diminished predictive performance. Fortunately, we found that this can be mitigated by introducing an intermediate layer of intermediate size. We further demonstrate that one can constrain the connectivity of the sparse layer to be of constant fan-in, in the sense that each output neuron will have the exact same number of incoming connections, which allows for more efficient implementations, especially on GPU hardware. The CUDA implementation of our approach is provided at https://github.com/xmc-aalto/ecml23-sparse. | en |
| dc.description.version | Peer reviewed | en |
| dc.format.extent | 16 | |
| dc.format.mimetype | application/pdf | en_US |
| dc.identifier.citation | Schultheis, E & Babbar, R 2023, Towards Memory-Efficient Training for Extremely Large Output Spaces : Learning with 670k Labels on a Single Commodity GPU. in D Koutra, C Plant, M Gomez Rodriguez, E Baralis & F Bonchi (eds), Machine Learning and Knowledge Discovery in Databases : Research Track - European Conference, ECML PKDD 2023, Proceedings. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), vol. 14171 LNAI, Springer, pp. 689-704, European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, Turin, Italy, 18/09/2023. https://doi.org/10.1007/978-3-031-43418-1_41 | en |
| dc.identifier.doi | 10.1007/978-3-031-43418-1_41 | en_US |
| dc.identifier.isbn | 978-3-031-43417-4 | |
| dc.identifier.issn | 0302-9743 | |
| dc.identifier.issn | 1611-3349 | |
| dc.identifier.other | PURE UUID: 543dff70-4944-4cdd-be8a-6df3365b6557 | en_US |
| dc.identifier.other | PURE ITEMURL: https://research.aalto.fi/en/publications/543dff70-4944-4cdd-be8a-6df3365b6557 | en_US |
| dc.identifier.other | PURE FILEURL: https://research.aalto.fi/files/150711131/SCI_Schultheis_etal_ECML_PKDD_2023.pdf | |
| dc.identifier.uri | https://aaltodoc.aalto.fi/handle/123456789/129661 | |
| dc.identifier.urn | URN:NBN:fi:aalto-202408065234 | |
| dc.language.iso | en | en |
| dc.relation.ispartof | European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases | en |
| dc.relation.ispartofseries | Machine Learning and Knowledge Discovery in Databases: Research Track - European Conference, ECML PKDD 2023, Proceedings | en |
| dc.relation.ispartofseries | pp. 689-704 | en |
| dc.relation.ispartofseries | Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) ; Volume 14171 LNAI | en |
| dc.rights | openAccess | en |
| dc.title | Towards Memory-Efficient Training for Extremely Large Output Spaces : Learning with 670k Labels on a Single Commodity GPU | en |
| dc.type | A4 Artikkeli konferenssijulkaisussa | fi |
| dc.type.version | publishedVersion |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- SCI_Schultheis_etal_ECML_PKDD_2023.pdf
- Size:
- 547.26 KB
- Format:
- Adobe Portable Document Format