| Challenge: | Linguistic Linked Open Data (LLOD) is a flourishing line of research in the language resource community . existing LLOD standards and vocabularies are not widely used in this community despite its popularity . |
| Approach: | They propose to use Linguistic Linked Open Data to link a Sumerian corpus with lexical resources . they use a linguistically annotated archive to create a corpus of cuneiform texts . |
| Outcome: | The proposed LLOD framework is used in assyriology, with philological resources underrepresented . the proposed framework is based on a linguistically annotated corpus of Sumerian texts . |
Similar Papers
From Linguistic Linked Data to Big Data (2024.lrec-main)
Copied to clipboard
Dimitar Trajanov, Elena Apostol, Radovan Garabik, Katerina Gkirtzou, Dagmar Gromann, Chaya Liebeskind, Cosimo Palma, Michael Rosner, Alexia Sampri, Gilles Sérasset, Blerina Spahiu, Ciprian-Octavian Truică, Giedre Valunaite Oleskeviciene
| Challenge: | Language data on the LOD cloud has grown in number, size, and variety . Linked (Open) Data (LLOD) is a standardized way of representing and sharing linguistic datasets . |
| Approach: | They propose to combine LLOD and Big Data to improve interoperability of linguistic datasets . they propose to use a machine-readable format to represent and share linguistic data . |
| Outcome: | This paper examines the use cases of Linked (Open) Data and Big Data in language data. |
Recent Developments for the Linguistic Linked Open Data Infrastructure (2020.lrec-1)
Copied to clipboard
Thierry Declerck, John Philip McCrae, Matthias Hartung, Jorge Gracia, Christian Chiarcos, Elena Montiel-Ponsoda, Philipp Cimiano, Artem Revenko, Roser Saurí, Deirdre Lee, Stefania Racioppa, Jamal Abdul Nasir, Matthias Orlikowsk, Marta Lanau-Coronas, Christian Fäth, Mariano Rico, Mohammad Fazleh Elahi, Maria Khvalchik, Meritxell Gonzalez, Katharine Cooney
| Challenge: | Language data is rarely 'ready-to-use' and language technology specialists spend over 80% of their time cleaning, organizing and collecting language datasets. |
| Approach: | They propose a methodology for building data value chains based around language resources and language technologies that can be integrated by means of semantic technologies. |
| Outcome: | The proposed methodology is based on language resources and language technologies that can be integrated by means of semantic technologies. |
The Index Thomisticus Treebank as Linked Data in the LiLa Knowledge Base (2022.lrec-1)
Copied to clipboard
| Challenge: | a series of Latin treebanks with word-by-word account of syntax and morphology of Latin texts have been published only in recent years. |
| Approach: | They propose to publish Latin treebanks that contain morphology and syntax annotations . they propose to use principles of the Linguistic Linked Open Data community . |
| Outcome: | The proposed approach enables interoperability between corpora and lexical resources for Latin . language learning and corpus-based research are the most obvious applications . |
LiDo RDF: From a Relational Database to a Linked Data Graph of Linguistic Terms and Bibliographic Data (L18-1)
Copied to clipboard
| Challenge: | linguists and researchers benefit from the data by looking it up on the Web . a new approach allows the direct use and reuse of the data for scientific research and machine processing . |
| Approach: | They propose to convert LiDo TBD database into Linked Data graph using Semantic Web . goal is to enable direct use and reuse of data for scientific research community . |
| Outcome: | The proposed dataset is based on the framework developed by linguist Dr. Christian Lehmann 40 years ago and is available on the LiDo website since 2006. |
WeDH - a Friendly Tool for Building Literary Corpora Enriched with Encyclopedic Metadata (2020.lrec-1)
Copied to clipboard
| Challenge: | Linked Open Data repositories are difficult to use for text corpora enriched with metadata . a collaborative project aims to fill the access to textual resources available on the web and the possibility of combining these resources with sources of metadata extending the life and maintenance of the data itself. |
| Approach: | They propose a web interface that allows users to leverage encyclopedic knowledge from DBpedia, wikidata and VIAF to enrich texts with bibliographical and exegetical knowledge. |
| Outcome: | WeDH aims to fill the access to textual resources available on the web and the possibility of combining these resources with sources of metadata that can enrich the texts with useful information. |
Towards the First Machine Translation System for Sumerian Transliterations (2020.coling-main)
Copied to clipboard
| Challenge: | Sumerian cuneiform script was invented more than 5,000 years ago and is one of the oldest in history. |
| Approach: | They propose to translate Sumerian texts into English automatically using supervised, phrase-based, and transfer learning techniques. |
| Outcome: | The proposed method accelerates the costly and time-consuming manual translation process and helps researchers better explore the relationships between Sumerian and Mesopotamian culture. |
The Abkhaz National Corpus (L18-1)
Copied to clipboard
| Challenge: | Abkhaz National Corpus is a comprehensive and open, grammatically annotated text corpus . it is currently growing and is being extended to include all important texts written in the language . |
| Approach: | They propose to use the Abkhaz National Corpus to annotate Abkhhaz texts . the corpus is a comprehensive and open, grammatically annotated text corpus . |
| Outcome: | The proposed corpus is a grammatically annotated text corpus which makes the language accessible to scientific investigations from various perspectives. |
The ACL OCL Corpus: Advancing Open Science in Computational Linguistics (2023.emnlp-main)
Copied to clipboard
| Challenge: | ACL OCL is a scholarly corpus derived from the ACL Anthology . it provides metadata, PDF files, citation graphs and additional structured full texts . |
| Approach: | They present ACL OCL, a scholarly corpus derived from the ACL Anthology . it integrates metadata, PDF files, citation graphs and additional structured full texts . they highlight how it applies to observe trends in computational linguistics . |
| Outcome: | The ACL OCL spans seven decades and contains 73,285 papers . the scholarly corpus is based on the ACL Anthology and is available from HuggingFace . |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Praaline: An Open-Source System for Managing, Annotating, Visualising and Analysing Speech Corpora (P18-4)
Copied to clipboard
| Challenge: | Praaline is an open-source software system for constituting and managing spoken language and multimodal corpora. |
| Approach: | They present the latest developments of Praaline, an open-source software system for constituting and managing spoken language and multimodal corpora. |
| Outcome: | The proposed system can be used for creating, managing, visualising and analysing spoken language and multimodal corpora. |