Automatic Validation of the Non-Validated Spanish Speech Data of Common Voice 17.0
dc.contributor.author | Hernández Mena, Carlos Daniel | |
dc.contributor.author | Scalvini, Barbara | |
dc.contributor.author | Lág, Dávid í | |
dc.contributor.editor | Tudor, Crina Madalina | |
dc.contributor.editor | Debess, Iben Nyholm | |
dc.contributor.editor | Bruton, Micaella | |
dc.contributor.editor | Scalvini, Barbara | |
dc.contributor.editor | Ilinykh, Nikolai | |
dc.contributor.editor | Holdt, Špela Arhar | |
dc.coverage.spatial | Tallinn, Estonia | |
dc.date.accessioned | 2025-02-14T10:09:19Z | |
dc.date.available | 2025-02-14T10:09:19Z | |
dc.date.issued | 2025-03 | |
dc.description.abstract | Mozilla Common Voice is a crowdsourced project that aims to create a public, multilingual dataset of voice recordings for training speech recognition models. In Common Voice, anyone can contribute by donating or validating recordings in various languages. However, despite the availability of many recordings in certain languages, a significant percentage remains unvalidated by users. This is the case for Spanish, where in version 17.0 of Common Voice, 75\% of the 2,220 hours of recordings are unvalidated. In this work, we used the Whisper recognizer to automatically validate approximately 784 hours of recordings which are more than the 562 hours validated by users. To verify the accuracy of the validation, we developed a speech recognition model based on a version of NVIDIA-NeMo’s Parakeet, which does not have an official Spanish version. Our final model achieved a WER of less than 4\% on the test and validation splits of Common Voice 17.0. Both the model and the speech corpus are publicly available on Hugging Face. | |
dc.identifier.uri | https://aclanthology.org/2025.resourceful-1.0/ | |
dc.identifier.uri | https://hdl.handle.net/10062/107116 | |
dc.language.iso | en | |
dc.publisher | University of Tartu Library | |
dc.rights | Attribution-NonCommercial-NoDerivatives 4.0 International | |
dc.rights.uri | https://creativecommons.org/licenses/by-nc-nd/4.0/ | |
dc.title | Automatic Validation of the Non-Validated Spanish Speech Data of Common Voice 17.0 | |
dc.type | Article |
Failid
Originaal pakett
1 - 1 1
Laen...
- Nimi:
- 2025_resourceful_1_12.pdf
- Suurus:
- 116.57 KB
- Formaat:
- Adobe Portable Document Format