Papers by Christian Fäth
Interoperability of Language-related Information: Mapping the BLL Thesaurus to Lexvo and Glottolog (L18-1)
Copied to clipboard
| Challenge: | The Bibliography of Linguistic Literature (BLL Thesaurus) has been used since 2013 in the context of the Lin gu is tik portal, a hub for linguistically relevant information. |
| Approach: | They propose to use Lexvo and Glottolog to facilitate interoperability between the BLL Thesaurus and terminological repositories in the Linguistic Linked Open Data cloud. |
| Outcome: | The proposed model is based on Lexvo and Glottolog and is able to connect to the Linguistic Linked Open Data cloud. |
Fintan - Flexible, Integrated Transformation and Annotation eNgineering (2020.lrec-1)
Copied to clipboard
| Challenge: | Fintan is a platform for converting heterogeneous linguistic resources to RDF. |
| Approach: | They introduce Fintan for converting heterogeneous linguistic resources to RDF with its modular architecture, workflow management and visualization features. |
| Outcome: | The Fintan platform is designed to transform linguistic resources to graphs and graphs. |
Recent Developments for the Linguistic Linked Open Data Infrastructure (2020.lrec-1)
Copied to clipboard
Thierry Declerck, John Philip McCrae, Matthias Hartung, Jorge Gracia, Christian Chiarcos, Elena Montiel-Ponsoda, Philipp Cimiano, Artem Revenko, Roser Saurí, Deirdre Lee, Stefania Racioppa, Jamal Abdul Nasir, Matthias Orlikowsk, Marta Lanau-Coronas, Christian Fäth, Mariano Rico, Mohammad Fazleh Elahi, Maria Khvalchik, Meritxell Gonzalez, Katharine Cooney
| Challenge: | Language data is rarely 'ready-to-use' and language technology specialists spend over 80% of their time cleaning, organizing and collecting language datasets. |
| Approach: | They propose a methodology for building data value chains based around language resources and language technologies that can be integrated by means of semantic technologies. |
| Outcome: | The proposed methodology is based on language resources and language technologies that can be integrated by means of semantic technologies. |
Universal Morphologies for the Caucasus region (L18-1)
Copied to clipboard
Christian Chiarcos, Kathrin Donandt, Maxim Ionov, Monika Rind-Pawlowski, Hasmik Sargsian, Jesse Wichers Schreur, Frank Abromeit, Christian Fäth
| Challenge: | Caucasus region is famed for its rich and diverse arrays of languages and language families . authors describe efforts to improve the coverage of Universal Morphologies for languages of the region . |
| Approach: | They propose to improve the coverage of Universal Morphologies for Caucasus languages . they propose to complement the Universal Dependencies which focus on morphosyntax and syntax. |
| Outcome: | The proposed framework improves the coverage of languages of the Caucasus region . the proposed framework criticizes the UniMorph TSV format for its limited expressiveness . |
The ACoLi Dictionary Graph (2020.lrec-1)
Copied to clipboard
| Challenge: | ACoLi Dictionary Graph is a collection of multilingual open source dictionaries available in two machine-readable formats. |
| Approach: | They propose to map and harmonize ACoLi Dictionary Graph into a unified representation and a tabular data format to facilitate their use in NLP tasks. |
| Outcome: | The ACoLi Dictionary Graph is a collection of multilingual open source dictionaries available in two machine-readable formats. |
Annotation Interoperability for the Post-ISOCat Era (2020.lrec-1)
Copied to clipboard
| Challenge: | Using ISOCat successor solutions, annotation standards have been developed since 2010 . |
| Approach: | They describe ISOCat successor solutions and annotation standardization efforts since 2010 . they describe low-cost harmonization of post-ISOCat vocabularies by means of linked ontologies . |
| Outcome: | The proposed ontologies are linked with the Ontologie of Linguistic Annotation and ISOCat, the GOLD ontology, the Typological Database Systems ontological and a large number of annotation schemes. |
Unifying Morphology Resources with OntoLex-Morph. A Case Study in German (2022.lrec-1)
Copied to clipboard
| Challenge: | OntoLex is a widely used community standard for machine-readable lexical resources on the web. |
| Approach: | They propose a module for representing morphology that can be used to encode and integrate morphological resources on a unified basis. |
| Outcome: | The proposed module can be used to represent morphological resources on a unified basis. |
Analyzing Middle High German Syntax with RDF and SPARQL (L18-1)
Copied to clipboard
| Challenge: | Using CoNLL-RDF and SPARQL Update, we analyze the diachronic changes of Middle High German syntax. |
| Approach: | They propose a rule-based shallow parser and an enrichment pipeline grounded in CoNLL-RDF and SPARQL Update for parsing. |
| Outcome: | The proposed pipeline is based on CoNLL-RDF and SPARQL Update for syntactic annotation and semantic enrichment of Middle High German. |
Querying a Dozen Corpora and a Thousand Years with Fintan (2022.lrec-1)
Copied to clipboard
| Challenge: | Large-scale quantitative diachronic corpus studies are difficult if multiple corpus are to be consulted . multi-layer corpus technology can solve the problem, but it requires the user to run queries manually. |
| Approach: | They propose a platform for studying word order in German using syntactically annotated corpora . fintan is a flexible integrated transformation and annotation platform . |
| Outcome: | The proposed platform can be used to study word order in German . it hints at two major phases in the development of scrambling in modern german . |