| Naslov: | Weakly-supervised multilingual medical NER for symptom extraction for low-resource languages |
|---|
| Avtorji: | ID Sallauka, Rigon (Avtor) ID Arioz, Umut (Avtor) ID Rojc, Matej (Avtor) ID Mlakar, Izidor (Avtor) |
| Datoteke: | applsci-15-05585-v2.pdf (338,94 KB) MD5: 9E3606C205F09FCCA4B26DDF5C379DCF
|
|---|
| Jezik: | Angleški jezik |
|---|
| Vrsta gradiva: | Članek v reviji |
|---|
| Tipologija: | 1.01 - Izvirni znanstveni članek |
|---|
| Organizacija: | FERI - Fakulteta za elektrotehniko, računalništvo in informatiko
|
|---|
| Opis: | Patient-reported health data, especially patient-reported outcomes measures, are vital for improving clinical care but are often limited by memory bias, cognitive load, and inflexible questionnaires. Patients prefer conversational symptom reporting, highlighting the need for robust methods in symptom extraction and conversational intelligence. This study presents a weakly-supervised pipeline for training and evaluating medical Named Entity Recognition (NER) models across eight languages, with a focus on low-resource settings. A merged English medical corpus, annotated using the Stanza i2b2 model, was translated into German, Greek, Spanish, Italian, Portuguese, Polish, and Slovenian, preserving the entity annotations medical problems, diagnostic tests, and treatments. Data augmentation addressed the class imbalance, and the fine-tuned BERT-based models outperformed baselines consistently. The English model achieved the highest F1 score (80.07%), followed by German (78.70%), Spanish (77.61%), Portuguese (77.21%), Slovenian (75.72%), Italian (75.60%), Polish (75.56%), and Greek (69.10%). Compared to the existing baselines, our models demonstrated notable performance gains, particularly in English, Spanish, and Italian. This research underscores the feasibility and effectiveness of weakly-supervised multilingual approaches for medical entity extraction, contributing to improved information access in clinical narratives—especially in under-resourced languages. |
|---|
| Ključne besede: | low-resource languages, machine translation, medical entity extraction, NER, NLP, patient-reported outcomes, weakly-supervised learning |
|---|
| Status publikacije: | Objavljeno |
|---|
| Verzija publikacije: | Objavljena publikacija |
|---|
| Poslano v recenzijo: | 01.05.2025 |
|---|
| Datum sprejetja članka: | 13.05.2025 |
|---|
| Datum objave: | 16.05.2025 |
|---|
| Založnik: | MDPI |
|---|
| Leto izida: | 2025 |
|---|
| Št. strani: | 18 str. |
|---|
| Številčenje: | Vol. 15, iss. 10, [article no.] 5585 |
|---|
| PID: | 20.500.12556/DKUM-92857  |
|---|
| UDK: | 004.8:61 |
|---|
| COBISS.SI-ID: | 236281347  |
|---|
| DOI: | 10.3390/app15105585  |
|---|
| ISSN pri članku: | 2076-3417 |
|---|
| Avtorske pravice: | © 2025 by the authors
|
|---|
| Datum objave v DKUM: | 19.05.2025 |
|---|
| Število ogledov: | 189 |
|---|
| Število prenosov: | 6 |
|---|
| Metapodatki: |  |
|---|
| Področja: | Ostalo
|
|---|
|
:
|
Kopiraj citat |
|---|
| | | | Skupna ocena: | (0 glasov) |
|---|
| Vaša ocena: | Ocenjevanje je dovoljeno samo prijavljenim uporabnikom. |
|---|
| Objavi na: |  |
|---|
Postavite miškin kazalec na naslov za izpis povzetka. Klik na naslov izpiše
podrobnosti ali sproži prenos. |