| Challenge: | a specialized web corpus and robust pattern-based extraction methods are used to detect definitory contexts. |
| Approach: | They propose to use a web corpus and a database to detect definitory contexts . they describe an experimental setting and front-end for pattern-based definition extraction . |
| Outcome: | The proposed method is based on a web corpus and a robust pattern-based extraction method. |
Similar Papers
Know thy Corpus! Robust Methods for Digital Curation of Web corpora (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for estimating the lexicon of Web corpora have not been used to train pre-trained models. |
| Approach: | They propose a framework for digital curation of Web corpora to provide robust estimation of their parameters. |
| Outcome: | The proposed framework provides robust estimation of Web corpora's composition and lexicon . the proposed framework is similar to the BNC and ELMO models, but lacks curated categories . |
Annotation and Automatic Classification of Aspectual Categories (P19-1)
Copied to clipboard
| Challenge: | Annotated resource for aspectual classification of German verb tokens in context. |
| Approach: | They present a resource for aspectual classification of German verb tokens in their clausal context. |
| Outcome: | The proposed resource is compared with previous work on German verb tokens using aspectual features compatible with the plurality of aspectual classifications. |
Text Mining for History: first steps on building a large dataset (L18-1)
Copied to clipboard
| Challenge: | a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way . |
| Approach: | They propose to use a Brazilian historical-biographical dictionary as a resource for text mining. |
| Outcome: | The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated . |
Textual Coverage of Eventive Entries in Lexical Semantic Resources (2024.lrec-main)
Copied to clipboard
| Challenge: | Several English, German, Spanish and Czech lexical semantic resources (which, for the most part, focus on verbs and predicates) have been selected for this experiment. |
| Approach: | They propose to quantify coverage gaps in lexical semantic resources when applied to running texts taken from the internet. |
| Outcome: | The proposed resources cover eventive entries (verbs, predicates, etc.) of well-known lexical semantic resources when applied to running texts taken from the internet. |
FEIDEGGER: A Multi-modal Corpus of Fashion Images and Descriptions in German (L18-1)
Copied to clipboard
| Challenge: | Recent years have seen a renewed interest in text-image multi-modality . paired text-picture datasets are often limited to English language text . |
| Approach: | They propose a multi-modal corpus that pairs images and textual descriptions of their content in German to enable study of these challenges. |
| Outcome: | The proposed dataset focuses on the domain of fashion items and their visual descriptions in German. |
GerEO: A Large-Scale Resource on the Syntactic Distribution of German Experiencer-Object Verbs (2022.lrec-1)
Copied to clipboard
| Challenge: | Psych verbs and their properties in multiple languages have ignited discussions among linguists for several decades . Psych-verbs are often considered syntactically deviant, although this has occasionally been called into question . |
| Approach: | They propose to use a large-scale database of more than 10,000 examples for 64 verbs from a newspaper corpus annotated for several syntactic and semantic features relevant for their analysis. |
| Outcome: | The proposed database contains 10,000 examples for 64 verbs from a newspaper corpus and includes syntactic construction, semantic stimulus type, and form of possible stimulus preposition. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Universal Dependencies and Quantitative Typological Trends. A Case Study on Word Order (L18-1)
Copied to clipboard
| Challenge: | a new method is proposed to acquire typological evidence from "gold" treebanks for different languages. |
| Approach: | They propose a method for acquiring typological evidence from "gold" treebanks for different languages. |
| Outcome: | The proposed method can shed light on key issues of the linguistic typological literature. |
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)
Copied to clipboard
| Challenge: | Using monolingual-only data, we can automate readability assessment and text simplification of simplified language. |
| Approach: | They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data. |
| Outcome: | The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images. |
Predicting Degrees of Technicality in Automatic Terminology Extraction (2020.acl-main)
Copied to clipboard
| Challenge: | a recent study has focused on term technicality, but there are still few studies on it. |
| Approach: | They semi-automatically create a German gold standard of technicality across four domains . they propose two new models to exploit general- vs. domain-specific comparisons based on vector spaces . |
| Outcome: | The proposed model outperforms previous methods in terms of general- vs. domain-specific comparisons. |