Voices of Luxembourg: Tackling Dialect Diversity in a Low-Resource Setting

Hosseini-Kivanani, Nina; Schommer, Christoph; Gilles, Peter

Voices of Luxembourg: Tackling Dialect Diversity in a Low-Resource Setting

Failid

2025_resourceful_1_29.pdf (159.93 KB)

Kuupäev

2025-03

Autorid

Hosseini-Kivanani, Nina

Schommer, Christoph

Gilles, Peter

Kirjastaja

University of Tartu Library

Abstrakt

Dialect classification is essential for preserving linguistic diversity, particularly in low-resource languages such as Luxembourgish. This study introduces one of the first systematic approaches to classifying Luxembourgish dialects, addressing phonetic, prosodic, and lexical variations across four major regions. We benchmarked multiple models, including state-of-the-art pre-trained speech models like Wav2Vec2, XLSR-Wav2Vec2, and Whisper, alongside traditional approaches such as Random Forest and CNN-LSTM. To overcome data limitations, we applied targeted data augmentation strategies and analyzed their impact on model performance. Our findings highlight the superior performance of CNN-Spectrogram and CNN-LSTM models while identifying the strengths and limitations of data augmentation. This work establishes foundational benchmarks and provides actionable insights for advancing dialectal NLP in Luxembourgish and other low-resource languages.

URI

https://hdl.handle.net/10062/107127

Kollektsioonid

Proceedings of the Third Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2025)

Kirje täielik lehekülg

Voices of Luxembourg: Tackling Dialect Diversity in a Low-Resource Setting

Failid

Kuupäev

Autorid

Ajakirja pealkiri

Ajakirja ISSN

Köite pealkiri

Kirjastaja

Abstrakt

Kirjeldus

Märksõnad

Viide

URI

Kollektsioonid