| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:Korpusne oznake za opis konteksta govornih dogodkov v slovenskih govornih korpusih
Authors:ID Bizjak, Andreja (Author)
Files:.pdf Bizjak.pdf (803,84 KB)
MD5: 3D2A02F1B4BDA71B099A9D5E87A5641B
 
URL https://journals.uni-lj.si/slovenscina2/article/view/18015/16273
 
Language:Slovenian
Work type:Article
Typology:1.01 - Original Scientific Article
Organization:FERI - Faculty of Electrical Engineering and Computer Science
Abstract:Zaradi časovno in finančno zahtevne priprave govornega korpusa je ob zasnovi potreben temeljit razmislek o njegovi sestavi in kategorizaciji beleženih metapodatkov. Raznoliki govorni dogodki, vključeni v nacionalni referenčni korpus, naj bi v čim večji meri odražali raznolikost sodobnega govorjenega jezika. Zanimalo nas bo, na kakšen način kategorizirati oznake za opis konteksta govornih dogodkov, da bi to reprezentativnost dosegli, ne da bi se popolnoma odrekli medsebojni primerljivosti podatkov. Premišljena zasnova nam omogoča, da je ob kasnejših korpusnih nadgradnjah potrebnih čim manj časovno zamudnih prilagoditev oznak. Izvedli bomo primerjalno analizo domačih in tujih govornih korpusov, s katero bomo kritično ovrednotili štiri temeljne kategorije oznak za opis konteksta govorne situacije. Pregledali bomo zasnovo tujih referenčnih govornih korpusov FOLK, BNC2014, ORAL2013, Nizozemskega govornega korpusa in C-ORAL-ROM ter jih primerjali z aktualnim referenčnim korpusom govorjene slovenščine Gos 2.1. Problematizirali bomo izbrane oznake in izpostavili težavnejša mesta, ki bi zahtevala dodatne premisleke in potencialno prekategorizacijo v prihodnje.
Keywords:govorni korpusi, zasnova korpusa, govorni dogodki, kategorizacija oznak
Publication status:Published
Publication version:Version of Record
Publication date:30.08.2024
Publisher:University of Ljubljana Press, Slovenia (Založba Univerze v Ljubljani)
Year of publishing:2024
Number of pages:str. 54-94
Numbering:Letn. 12, št. 1
PID:20.500.12556/DKUM-90557-3212b11b-a506-210d-d38f-075f6b692399 New window
UDC:004.8:808
ISSN on article:2335-2736
COBISS.SI-ID:206755331 New window
DOI:10.4312/slo2.0.2024.1.54-94 New window
Copyright:https://creativecommons.org/licenses/by-sa/4.0
Publication date in DKUM:09.09.2024
Views:132
Downloads:9
Metadata:XML DC-XML DC-RDF
Categories:Misc.
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Record is a part of a journal

Title:Slovenščina 2.0. empirične, aplikativne in interdisciplinarne raziskave
Publisher:Trojina, zavod za uporabno slovenistiko, Trojina, zavod za uporabno slovenistiko, Trojina, zavod za uporabno slovenistiko, Znanstvena založba Filozofske fakultete, Znanstvena založba Filozofske fakultete, Založba Univerze v Ljubljani
ISSN:2335-2736
COBISS.SI-ID:264547328 New window

Document is financed by a project

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:J7-4642-2022
Name:Temeljne raziskave za razvoj govornih virov in tehnologij za slovenščino

Licences

License:CC BY-SA 4.0, Creative Commons Attribution-ShareAlike 4.0 International
Link:http://creativecommons.org/licenses/by-sa/4.0/
Description:This Creative Commons license is very similar to the regular Attribution license, but requires the release of all derivative works under this same license.

Secondary language

Language:English
Title:Corpus annotations for describing the context of speech events in Slovene speech corpora
Abstract:The time-consuming and costly preparation of a speech corpus requires ca-reful consideration of its composition and the categorization of the recorded metadata at the time of its design. The variety of speech events included in the national reference corpus should reflect the diversity of contemporary spoken language as much as possible. We will be interested in how to categorize the annotations used to describe the context of the speech events in order to achi-eve this representativeness without completely giving up the intercomparabi-lity of the data. A thoughtful design allows us to minimize the time-consuming tag adjustments required in subsequent corpus upgrades. We will carry out a comparative analysis of domestic and foreign speech corpora to critically eva-luate the four basic categories of annotations used to describe the context of a speech situation. We will review the design of the foreign reference speech corpora FOLK, BNC2014, ORAL2013, the Spoken Dutch Corpus and C-ORAL--ROM and compare them with the current reference corpus of spoken Slovene Gos 2.1. We will problematize the selected annotations and highlight the more problematic areas that would require further consideration and potential futu-re re-categorization.
Keywords:speech corpora, corpus design, speech events, categorization of annotations


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica