Acquiring a Formality-Informed Lexical Resource for Style Analysis (2021.eacl-main)
Copied to clipboard
| Challenge: | lexico-statistics analysis of formality levels in written communication has long been dominated by application concerns, such as authorship and plagiarism assignment problems. |
| Approach: | They propose a lexicon with entries ordered by their degree of (in)formality and let crowdworkers assess the enlarged set of lexical items on a continuous informal-formal scale as a gold standard for evaluation. |
| Outcome: | The proposed lexicon is evaluated on a German-language email corpus and is then evaluated by crowdworkers. |
Similar Papers
A Question of Style: A Dataset for Analyzing Formality on Different Levels (2023.findings-eacl)
Copied to clipboard
| Challenge: | Using machine learning, we can produce contextually appropriate language. |
| Approach: | They present a dataset of German sentence-level formality assessed on a continuous informal-formal scale. |
| Outcome: | The proposed dataset compares sentences from a wide range of genres assessed on a continuous informal-formal scale. |
Predicting Degrees of Technicality in Automatic Terminology Extraction (2020.acl-main)
Copied to clipboard
| Challenge: | a recent study has focused on term technicality, but there are still few studies on it. |
| Approach: | They semi-automatically create a German gold standard of technicality across four domains . they propose two new models to exploit general- vs. domain-specific comparisons based on vector spaces . |
| Outcome: | The proposed model outperforms previous methods in terms of general- vs. domain-specific comparisons. |
Exploiting a lexical resource for discourse connective disambiguation in German (2020.coling-main)
Copied to clipboard
| Challenge: | a connective lexicon can be a valuable resource for languages with limited PDTB-style annotations . connectives are usually understood to be ambiguous in two different ways . |
| Approach: | They propose to augment a purely-empirical approach to connective identification and sense classification in German . they find that a connective lexicon can be a valuable resource for those languages with a large PDTB-style-annotated coprus . |
| Outcome: | The proposed approach improves on published results and achieves an F1 score for German sense classification. |
Enriching a Lexicon of Discourse Connectives with Corpus-based Data (L18-1)
Copied to clipboard
| Challenge: | Existing annotation efforts for multiple languages have focused on discourse connectives, but we have limited it to the class of connectives marking contrast and the additional relations such connectives might convey. |
| Approach: | They enrich a lexicon of italian COnnectives with real corpus data for connectives marking contrast relations in text. |
| Outcome: | The proposed resource is a valuable tool for linguistic analyses of discourse relations and the training of a classifier for NLP applications. |
Introducing Lexical Masks: a New Representation of Lexical Entries for Better Evaluation and Exchange of Lexicons (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing standards for lexicon format and features are inadequate for evaluation and exchange . lexical masks are a powerful tool used to evaluate and exchange large lexiconic databases . |
| Approach: | They propose a tool to evaluate and exchange lexicon databases in many languages . they propose lexical masks which represent the expected internal structure of a lexico . |
| Outcome: | The proposed lexical masks can be used to evaluate and exchange lexicon databases in many languages. |
A database of German definitory contexts from selected web sources (L18-1)
Copied to clipboard
| Challenge: | a specialized web corpus and robust pattern-based extraction methods are used to detect definitory contexts. |
| Approach: | They propose to use a web corpus and a database to detect definitory contexts . they describe an experimental setting and front-end for pattern-based definition extraction . |
| Outcome: | The proposed method is based on a web corpus and a robust pattern-based extraction method. |
Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality Transfer (2021.emnlp-main)
Copied to clipboard
| Challenge: | a lack of standardized and reliable methods for automatic evaluation hinders ST . prior work has employed as many as nine different automatic systems to rate formality alone . |
| Approach: | They evaluate automatic metrics on the oft-researched task of formality style transfer . they outline best practices for automatic evaluation in (formality) style transfer and identify models that correlate well with human judgments. |
| Outcome: | The proposed models correlate well with human judgments and are robust across languages. |
Dear Sir or Madam, May I Introduce the GYAFC Dataset: Corpus, Benchmarks and Metrics for Formality Style Transfer (N18-1)
Copied to clipboard
| Challenge: | a lack of training and evaluation datasets, benchmarks and automatic metrics has blocked progress in this field. |
| Approach: | They propose to use a grammarly's Yahoo Answers Formality corpus to create the largest corpus for a particular style . they also propose to apply machine translation metrics to the task . |
| Outcome: | The proposed model can be used to train and evaluate a text in a particular style . the proposed model is based on the existing model and can be applied to other tasks . |
Textual Coverage of Eventive Entries in Lexical Semantic Resources (2024.lrec-main)
Copied to clipboard
| Challenge: | Several English, German, Spanish and Czech lexical semantic resources (which, for the most part, focus on verbs and predicates) have been selected for this experiment. |
| Approach: | They propose to quantify coverage gaps in lexical semantic resources when applied to running texts taken from the internet. |
| Outcome: | The proposed resources cover eventive entries (verbs, predicates, etc.) of well-known lexical semantic resources when applied to running texts taken from the internet. |
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)
Copied to clipboard
| Challenge: | Using monolingual-only data, we can automate readability assessment and text simplification of simplified language. |
| Approach: | They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data. |
| Outcome: | The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images. |