A database of German definitory contexts from selected web sources (L18-1)

Copied to clipboard

Challenge: a specialized web corpus and robust pattern-based extraction methods are used to detect definitory contexts.
Approach: They propose to use a web corpus and a database to detect definitory contexts . they describe an experimental setting and front-end for pattern-based definition extraction .
Outcome: The proposed method is based on a web corpus and a robust pattern-based extraction method.

Similar Papers

Know thy Corpus! Robust Methods for Digital Curation of Web corpora (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for estimating the lexicon of Web corpora have not been used to train pre-trained models.
Approach: They propose a framework for digital curation of Web corpora to provide robust estimation of their parameters.
Outcome: The proposed framework provides robust estimation of Web corpora's composition and lexicon . the proposed framework is similar to the BNC and ELMO models, but lacks curated categories .
Annotation and Automatic Classification of Aspectual Categories (P19-1)

Copied to clipboard

Challenge: Annotated resource for aspectual classification of German verb tokens in context.
Approach: They present a resource for aspectual classification of German verb tokens in their clausal context.
Outcome: The proposed resource is compared with previous work on German verb tokens using aspectual features compatible with the plurality of aspectual classifications.
Text Mining for History: first steps on building a large dataset (L18-1)

Copied to clipboard

Challenge: a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way .
Approach: They propose to use a Brazilian historical-biographical dictionary as a resource for text mining.
Outcome: The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated .
Textual Coverage of Eventive Entries in Lexical Semantic Resources (2024.lrec-main)

Copied to clipboard

Challenge: Several English, German, Spanish and Czech lexical semantic resources (which, for the most part, focus on verbs and predicates) have been selected for this experiment.
Approach: They propose to quantify coverage gaps in lexical semantic resources when applied to running texts taken from the internet.
Outcome: The proposed resources cover eventive entries (verbs, predicates, etc.) of well-known lexical semantic resources when applied to running texts taken from the internet.
FEIDEGGER: A Multi-modal Corpus of Fashion Images and Descriptions in German (L18-1)

Copied to clipboard

Challenge: Recent years have seen a renewed interest in text-image multi-modality . paired text-picture datasets are often limited to English language text .
Approach: They propose a multi-modal corpus that pairs images and textual descriptions of their content in German to enable study of these challenges.
Outcome: The proposed dataset focuses on the domain of fashion items and their visual descriptions in German.
GerEO: A Large-Scale Resource on the Syntactic Distribution of German Experiencer-Object Verbs (2022.lrec-1)

Copied to clipboard

Challenge: Psych verbs and their properties in multiple languages have ignited discussions among linguists for several decades . Psych-verbs are often considered syntactically deviant, although this has occasionally been called into question .
Approach: They propose to use a large-scale database of more than 10,000 examples for 64 verbs from a newspaper corpus annotated for several syntactic and semantic features relevant for their analysis.
Outcome: The proposed database contains 10,000 examples for 64 verbs from a newspaper corpus and includes syntactic construction, semantic stimulus type, and form of possible stimulus preposition.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Universal Dependencies and Quantitative Typological Trends. A Case Study on Word Order (L18-1)

Copied to clipboard

Challenge: a new method is proposed to acquire typological evidence from "gold" treebanks for different languages.
Approach: They propose a method for acquiring typological evidence from "gold" treebanks for different languages.
Outcome: The proposed method can shed light on key issues of the linguistic typological literature.
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)

Copied to clipboard

Challenge: Using monolingual-only data, we can automate readability assessment and text simplification of simplified language.
Approach: They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data.
Outcome: The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images.
Predicting Degrees of Technicality in Automatic Terminology Extraction (2020.acl-main)

Copied to clipboard

Challenge: a recent study has focused on term technicality, but there are still few studies on it.
Approach: They semi-automatically create a German gold standard of technicality across four domains . they propose two new models to exploit general- vs. domain-specific comparisons based on vector spaces .
Outcome: The proposed model outperforms previous methods in terms of general- vs. domain-specific comparisons.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations