Mapping Faroese in the Multilingual Representation Space: Insights for ASR Model Optimization
dc.contributor.author | Lág, Dávid í | |
dc.contributor.author | Scalvini, Barbara | |
dc.contributor.author | Gudnason, Jon | |
dc.contributor.editor | Johansson, Richard | |
dc.contributor.editor | Stymne, Sara | |
dc.coverage.spatial | Tallinn, Estonia | |
dc.date.accessioned | 2025-02-18T09:37:59Z | |
dc.date.available | 2025-02-18T09:37:59Z | |
dc.date.issued | 2025-03 | |
dc.description.abstract | ASR development for low-resource languages like Faroese faces significant challenges due to the scarcity of large, diverse datasets. While fine-tuning multilingual models using related languages is a common practice, there is no standardized method for selecting these auxiliary languages, leading to a computationally expensive trial-and-error process. By analyzing Faroese’s positioning among other languages in wav2vec2’s multilingual representation space, we find that Faroese's closest neighbors are influenced not only by linguistic similarity but also by historical, phonetic, and cultural factors. These findings open new avenues for auxiliary language selection to improve Faroese ASR and underscore the potential value of data-driven factors in ASR fine-tuning. | |
dc.identifier.uri | https://hdl.handle.net/10062/107229 | |
dc.language.iso | en | |
dc.publisher | University of Tartu Library | |
dc.relation.ispartofseries | NEALT Proceedings Series, No. 57 | |
dc.rights | Attribution-NonCommercial-NoDerivatives 4.0 International | |
dc.rights.uri | https://creativecommons.org/licenses/by-nc-nd/4.0/ | |
dc.title | Mapping Faroese in the Multilingual Representation Space: Insights for ASR Model Optimization | |
dc.type | Article |
Failid
Originaal pakett
1 - 1 1