The Devil’s in the Details: the Detailedness of Classes Influences Personal Information Detection and Labeling

Szawerna, Maria Irena; Dobnik, Simon; Muñoz Sánchez, Ricardo; Volodina, Elena

The Devil’s in the Details: the Detailedness of Classes Influences Personal Information Detection and Labeling

Failid

2025_nodalida_1_70.pdf (308.21 KB)

Kuupäev

2025-03

Autorid

Szawerna, Maria Irena

Dobnik, Simon

Muñoz Sánchez, Ricardo

Volodina, Elena

Kirjastaja

University of Tartu Library

Abstrakt

In this paper, we experiment with the effect of different levels of detailedness or granularity—understood as i) the number of classes, and ii) the classes’ semantic depth in the sense of hypernym and hyponym relations — of the annotation of Personally Identifiable Information (PII) on automatic detection and labeling of such information. We fine-tune a Swedish BERT model on a corpus of Swedish learner essays annotated with a total of six PII tagsets at varying levels of granularity. We also investigate whether the presence of grammatical and lexical correction annotation in the tokens and class prevalence have an effect on predictions. We observe that the fewer total categories there are, the better the overall results are, but having a more diverse annotation facilitates fewer misclassifications for tokens containing correction annotation. We also note that the classes’ internal diversity has an effect on labeling. We conclude from the results that while labeling based on the detailed annotation is difficult because of the number of classes, it is likely that models trained on such annotation rely more on the semantic content captured by contextual word embeddings rather than just the form of the tokens, making them more robust against nonstandard language.

URI

https://hdl.handle.net/10062/107263

Kollektsioonid

Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025)

Kirje täielik lehekülg

The Devil’s in the Details: the Detailedness of Classes Influences Personal Information Detection and Labeling

Failid

Kuupäev

Autorid

Ajakirja pealkiri

Ajakirja ISSN

Köite pealkiri

Kirjastaja

Abstrakt

Kirjeldus

Märksõnad

Viide

URI

Kollektsioonid