| | SLO | ENG | Cookies and privacy

Bigger font | Smaller font

Show document Help

Title:The OpenScience Slovenia metadata dataset
Authors:ID Borovič, Mladen (Author)
ID Ferme, Marko (Author)
ID Brezovnik, Janez (Author)
ID Majninger, Sandi (Author)
ID Bregant, Albin (Author)
ID Hrovat, Goran (Author)
ID Ojsteršek, Milan (Author)
Files:.pdf 1-s2.0-S2352340919312971-main.pdf (187,50 KB)
MD5: 96124625CAF8FA1FA4748950BA2D1379
Description: Data paper
 
URL https://hdl.handle.net/20.500.12556/DKUM-92889
Description: The research data is available in a digital object at DKUM.
 
URL https://doi.org/10.17632/7wh9xvvmgk.1
Description: The research data is available in a digital object at Mendeley.
 
Language:English
Work type:Unknown
Typology:1.03 - Other scientific articles
Organization:FERI - Faculty of Electrical Engineering and Computer Science
Abstract:The OpenScience Slovenia metadata dataset contains metadata entries for Slovenian public domain academic documents which include undergraduate and postgraduate theses, research and professional articles, along with other academic document types. The data within the dataset was collected as a part of the establishment of the Slovenian Open-Access Infrastructure which defined a unified document collection process and cataloguing for universities in Slovenia within the infrastructure repositories. The data was collected from several already established but separate library systems in Slovenia and merged into a single metadata scheme using metadata deduplication and merging techniques. It consists of text and numerical fields, representing attributes that describe documents. These attributes include document titles, keywords, abstracts, typologies, authors, issue years and other identifiers such as URL and UDC. The potential of this dataset lies especially in text mining and text classification tasks and can also be used in development or benchmarking of content-based recommender systems on real-world data.
Keywords:metadata, real world data, text data, text mining, text identification, natural language processing
Publication status:Published
Publication version:Version of Record
Publication date:01.02.2020
Year of publishing:2020
Number of pages:str. 1-5
Numbering:Vol. 28
PID:20.500.12556/DKUM-92888 New window
UDC:004.4
ISSN on article:2352-3409
COBISS.SI-ID:23110934 New window
DOI:10.1016/j.dib.2019.104942 New window
Publication date in DKUM:22.05.2025
Views:188
Downloads:13
Metadata:XML DC-XML DC-RDF
Categories:Misc.
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share



Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Record is a part of a journal

Title:Data in brief
Publisher:Elsevier
ISSN:2352-3409
COBISS.SI-ID:32117977 New window

Licences

License:CC BY-NC-ND 4.0, Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International
Link:http://creativecommons.org/licenses/by-nc-nd/4.0/
Description:The most restrictive Creative Commons license. This only allows people to download and share the work for no commercial gain and for no other purposes.

Secondary language

Language:Slovenian
Abstract:Nabor metapodatkov OpenScience Slovenija vsebuje metapodatkovne vnose za slovenske javno dostopne akademske dokumente, ki vključujejo dodiplomska in podiplomska dela, raziskovalne in strokovne članke ter druge vrste akademskih dokumentov. Podatki v podatkovnem nizu so bili zbrani v okviru vzpostavitve slovenske infrastrukture odprtega dostopa, ki je opredelila enoten postopek zbiranja in katalogizacije dokumentov za univerze v Sloveniji v okviru infrastrukturnih repozitorijev. Podatki so bili zbrani iz več že vzpostavljenih, vendar ločenih knjižničnih sistemov v Sloveniji in združeni v enotno metapodatkovno shemo z uporabo tehnik deduplikacije in združevanja metapodatkov. Sestavljajo jih besedilna in številčna polja, ki predstavljajo atribute, ki opisujejo dokumente. Ti atributi vključujejo naslove dokumentov, ključne besede, povzetke, tipologije, avtorje, letnice izdaje in druge identifikatorje, kot sta URL in UDC. Potencial tega nabora podatkov je zlasti v nalogah rudarjenja po besedilu in razvrščanja besedil, uporablja pa se lahko tudi pri razvoju ali primerjalni analizi priporočilnih sistemov, ki temeljijo na vsebini, na realnih podatkih.
Keywords:meta podatki, realni podatki, identifikacija teksta


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica