Papers by Christian Chiarcos
Interoperability of Language-related Information: Mapping the BLL Thesaurus to Lexvo and Glottolog (L18-1)
Copied to clipboard
| Challenge: | The Bibliography of Linguistic Literature (BLL Thesaurus) has been used since 2013 in the context of the Lin gu is tik portal, a hub for linguistically relevant information. |
| Approach: | They propose to use Lexvo and Glottolog to facilitate interoperability between the BLL Thesaurus and terminological repositories in the Linguistic Linked Open Data cloud. |
| Outcome: | The proposed model is based on Lexvo and Glottolog and is able to connect to the Linguistic Linked Open Data cloud. |
Fintan - Flexible, Integrated Transformation and Annotation eNgineering (2020.lrec-1)
Copied to clipboard
| Challenge: | Fintan is a platform for converting heterogeneous linguistic resources to RDF. |
| Approach: | They introduce Fintan for converting heterogeneous linguistic resources to RDF with its modular architecture, workflow management and visualization features. |
| Outcome: | The Fintan platform is designed to transform linguistic resources to graphs and graphs. |
Recent Developments for the Linguistic Linked Open Data Infrastructure (2020.lrec-1)
Copied to clipboard
Thierry Declerck, John Philip McCrae, Matthias Hartung, Jorge Gracia, Christian Chiarcos, Elena Montiel-Ponsoda, Philipp Cimiano, Artem Revenko, Roser Saurí, Deirdre Lee, Stefania Racioppa, Jamal Abdul Nasir, Matthias Orlikowsk, Marta Lanau-Coronas, Christian Fäth, Mariano Rico, Mohammad Fazleh Elahi, Maria Khvalchik, Meritxell Gonzalez, Katharine Cooney
| Challenge: | Language data is rarely 'ready-to-use' and language technology specialists spend over 80% of their time cleaning, organizing and collecting language datasets. |
| Approach: | They propose a methodology for building data value chains based around language resources and language technologies that can be integrated by means of semantic technologies. |
| Outcome: | The proposed methodology is based on language resources and language technologies that can be integrated by means of semantic technologies. |
A Tree Extension for CoNLL-RDF (2020.lrec-1)
Copied to clipboard
| Challenge: | CoNLL-RDF provides a bridge for popular oneword-per-line formats . main reasons for their popularity are the simplicity of tables and tab-separated values . |
| Approach: | They propose a technology that provides a bridge between knowledge graphs and natural language processing. |
| Outcome: | The proposed technology provides a bridge for popular one-word-per-line formats . it provides native support for word-level annotations, but not phrase structures or text structure . |
The ACoLi CoNLL Libraries: Beyond Tab-Separated Values (L18-1)
Copied to clipboard
| Challenge: | a new set of Java archives facilitates advanced manipulations of corpora annotated in TSV formats. |
| Approach: | They propose to use Java archives to facilitate advanced manipulations of corpora annotated in TSV formats. |
| Outcome: | The proposed libraries support all members of the CoNLL format family. |
Towards a Linked Open Data Edition of Sumerian Corpora (L18-1)
Copied to clipboard
| Challenge: | Linguistic Linked Open Data (LLOD) is a flourishing line of research in the language resource community . existing LLOD standards and vocabularies are not widely used in this community despite its popularity . |
| Approach: | They propose to use Linguistic Linked Open Data to link a Sumerian corpus with lexical resources . they use a linguistically annotated archive to create a corpus of cuneiform texts . |
| Outcome: | The proposed LLOD framework is used in assyriology, with philological resources underrepresented . the proposed framework is based on a linguistically annotated corpus of Sumerian texts . |
Bridging Computational Lexicography and Corpus Linguistics: A Query Extension for OntoLex-FrAC (2024.lrec-main)
Copied to clipboard
| Challenge: | OntoLex is the dominant community standard for machine-readable lexical resources . it is currently extended with a designated module for Frequency, Attestations and Corpus-based Information . |
| Approach: | They propose a module for Frequency, Attestations and Corpus-based Information for OntoLex . the module enables RDF-based web services to exchange corpus queries dynamically . |
| Outcome: | The proposed module addresses the incorporation of corpus queries for linking dictionaries with corpus engines and enabling RDF-based web services to exchange corpus query data dynamically. |
Cross-Lingual Link Discovery for Under-Resourced Languages (2022.lrec-1)
Copied to clipboard
Michael Rosner, Sina Ahmadi, Elena-Simona Apostol, Julia Bosque-Gil, Christian Chiarcos, Milan Dojchinovski, Katerina Gkirtzou, Jorge Gracia, Dagmar Gromann, Chaya Liebeskind, Giedrė Valūnaitė Oleškevičienė, Gilles Sérasset, Ciprian-Octavian Truică
| Challenge: | Linked data paradigms can be used to solve under-resourced languages' problem of under-utilization of resources. |
| Approach: | They propose a paradigm for cross-lingual link discovery that can be applied to under-resourced languages . they argue that techniques for cross language linking can be readily applied . |
| Outcome: | The proposed technologies can be applied to under-resourced languages, the authors argue . the authors show that the Linked Data paradigm can be used to solve the problem . |
Universal Morphologies for the Caucasus region (L18-1)
Copied to clipboard
Christian Chiarcos, Kathrin Donandt, Maxim Ionov, Monika Rind-Pawlowski, Hasmik Sargsian, Jesse Wichers Schreur, Frank Abromeit, Christian Fäth
| Challenge: | Caucasus region is famed for its rich and diverse arrays of languages and language families . authors describe efforts to improve the coverage of Universal Morphologies for languages of the region . |
| Approach: | They propose to improve the coverage of Universal Morphologies for Caucasus languages . they propose to complement the Universal Dependencies which focus on morphosyntax and syntax. |
| Outcome: | The proposed framework improves the coverage of languages of the Caucasus region . the proposed framework criticizes the UniMorph TSV format for its limited expressiveness . |
Inducing Discourse Marker Inventories from Lexical Knowledge Graphs (2022.lrec-1)
Copied to clipboard
| Challenge: | Discourse marker inventories are important tools for the development of discourse parsers and corpora with discourse annotations. |
| Approach: | They explore the potential of multilingual lexical knowledge graphs to induce multilingual discourse marker lexicons using concept propagation methods previously developed in translation inference across dictionaries. |
| Outcome: | The proposed method can induce multilingual discourse marker lexicons using multilingual knowledge graphs. |
Modelling Frequency, Attestation, and Corpus-Based Information with OntoLex-FrAC (2022.coling-1)
Copied to clipboard
| Challenge: | OntoLex-Lemon has become a de facto standard for lexical resources in the web of data. |
| Approach: | This paper provides the first overall description of the emerging OntoLex module for Frequency, Attestations, and Corpus-Based Information. |
| Outcome: | This paper provides the first overall description of the emerging OntoLex module for Frequency, Attestations, and Corpus-Based Information (OntoLx-FrAC) it is intended to complement OntoLemon with the vocabulary to represent major types of information found in or automatically derived from corpora, for applications in both language technology and the language sciences. |
The ACoLi Dictionary Graph (2020.lrec-1)
Copied to clipboard
| Challenge: | ACoLi Dictionary Graph is a collection of multilingual open source dictionaries available in two machine-readable formats. |
| Approach: | They propose to map and harmonize ACoLi Dictionary Graph into a unified representation and a tabular data format to facilitate their use in NLP tasks. |
| Outcome: | The ACoLi Dictionary Graph is a collection of multilingual open source dictionaries available in two machine-readable formats. |
Annotation Interoperability for the Post-ISOCat Era (2020.lrec-1)
Copied to clipboard
| Challenge: | Using ISOCat successor solutions, annotation standards have been developed since 2010 . |
| Approach: | They describe ISOCat successor solutions and annotation standardization efforts since 2010 . they describe low-cost harmonization of post-ISOCat vocabularies by means of linked ontologies . |
| Outcome: | The proposed ontologies are linked with the Ontologie of Linguistic Annotation and ISOCat, the GOLD ontology, the Typological Database Systems ontological and a large number of annotation schemes. |
On Modelling Corpus Citations in Computational Lexical Resources (2024.lrec-main)
Copied to clipboard
| Challenge: | TEI and OntoLex deal with corpus citations in lexicons. |
| Approach: | They argue that TEI and OntoLex can be used to model corpus citations in lexicons . they also argue that they should be combined to achieve a more accurate encoding . |
| Outcome: | The proposed approach favours a combination of TEI and OntoLex . the proposed approach is based on a model of an example entry from a legacy dictionary . |
Towards the First Machine Translation System for Sumerian Transliterations (2020.coling-main)
Copied to clipboard
| Challenge: | Sumerian cuneiform script was invented more than 5,000 years ago and is one of the oldest in history. |
| Approach: | They propose to translate Sumerian texts into English automatically using supervised, phrase-based, and transfer learning techniques. |
| Outcome: | The proposed method accelerates the costly and time-consuming manual translation process and helps researchers better explore the relationships between Sumerian and Mesopotamian culture. |
Unifying Morphology Resources with OntoLex-Morph. A Case Study in German (2022.lrec-1)
Copied to clipboard
| Challenge: | OntoLex is a widely used community standard for machine-readable lexical resources on the web. |
| Approach: | They propose a module for representing morphology that can be used to encode and integrate morphological resources on a unified basis. |
| Outcome: | The proposed module can be used to represent morphological resources on a unified basis. |
Analyzing Middle High German Syntax with RDF and SPARQL (L18-1)
Copied to clipboard
| Challenge: | Using CoNLL-RDF and SPARQL Update, we analyze the diachronic changes of Middle High German syntax. |
| Approach: | They propose a rule-based shallow parser and an enrichment pipeline grounded in CoNLL-RDF and SPARQL Update for parsing. |
| Outcome: | The proposed pipeline is based on CoNLL-RDF and SPARQL Update for syntactic annotation and semantic enrichment of Middle High German. |
Querying a Dozen Corpora and a Thousand Years with Fintan (2022.lrec-1)
Copied to clipboard
| Challenge: | Large-scale quantitative diachronic corpus studies are difficult if multiple corpus are to be consulted . multi-layer corpus technology can solve the problem, but it requires the user to run queries manually. |
| Approach: | They propose a platform for studying word order in German using syntactically annotated corpora . fintan is a flexible integrated transformation and annotation platform . |
| Outcome: | The proposed platform can be used to study word order in German . it hints at two major phases in the development of scrambling in modern german . |
ISO-based Annotated Multilingual Parallel Corpus for Discourse Markers (2022.lrec-1)
Copied to clipboard
Purificação Silvano, Mariana Damova, Giedrė Valūnaitė Oleškevičienė, Chaya Liebeskind, Christian Chiarcos, Dimitar Trajanov, Ciprian-Octavian Truică, Elena-Simona Apostol, Anna Baczkowska
| Challenge: | Discourse markers carry information about the discourse structure and organization, and also signal local dependencies or epistemic stance of speaker. |
| Approach: | They propose an ISO-based annotated multilingual parallel corpus for discourse markers . they propose an annotation scheme for discourse relations with a plug-in to ISO 24617-2 . |
| Outcome: | The proposed language resource is based on an ISO-based annotated multilingual parallel corpus of discourse markers. |