Challenge: Using machine learning, we can produce contextually appropriate language.
Approach: They present a dataset of German sentence-level formality assessed on a continuous informal-formal scale.
Outcome: The proposed dataset compares sentences from a wide range of genres assessed on a continuous informal-formal scale.

Similar Papers

Acquiring a Formality-Informed Lexical Resource for Style Analysis (2021.eacl-main)

Copied to clipboard

Challenge: lexico-statistics analysis of formality levels in written communication has long been dominated by application concerns, such as authorship and plagiarism assignment problems.
Approach: They propose a lexicon with entries ordered by their degree of (in)formality and let crowdworkers assess the enlarged set of lexical items on a continuous informal-formal scale as a gold standard for evaluation.
Outcome: The proposed lexicon is evaluated on a German-language email corpus and is then evaluated by crowdworkers.
Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality Transfer (2021.emnlp-main)

Copied to clipboard

Challenge: a lack of standardized and reliable methods for automatic evaluation hinders ST . prior work has employed as many as nine different automatic systems to rate formality alone .
Approach: They evaluate automatic metrics on the oft-researched task of formality style transfer . they outline best practices for automatic evaluation in (formality) style transfer and identify models that correlate well with human judgments.
Outcome: The proposed models correlate well with human judgments and are robust across languages.
Dear Sir or Madam, May I Introduce the GYAFC Dataset: Corpus, Benchmarks and Metrics for Formality Style Transfer (N18-1)

Copied to clipboard

Challenge: a lack of training and evaluation datasets, benchmarks and automatic metrics has blocked progress in this field.
Approach: They propose to use a grammarly's Yahoo Answers Formality corpus to create the largest corpus for a particular style . they also propose to apply machine translation metrics to the task .
Outcome: The proposed model can be used to train and evaluate a text in a particular style . the proposed model is based on the existing model and can be applied to other tasks .
Formality Style Transfer for Noisy, User-generated Conversations: Extracting Labeled, Parallel Data from Unlabeled Corpora (D19-55)

Copied to clipboard

Challenge: Typical datasets used for style transfer in NLP contain aligned pairs of two opposite extremes of a style.
Approach: They propose a technique to derive a dataset of aligned pairs from an unlabeled corpus by using an auxiliary dataset, allowing for in-domain training.
Outcome: The proposed method significantly outperforms OpenNMT’s Seq2Seq model trained on the Yahoo Formality Dataset and 6 novel datasets.
Predicting Degrees of Technicality in Automatic Terminology Extraction (2020.acl-main)

Copied to clipboard

Challenge: a recent study has focused on term technicality, but there are still few studies on it.
Approach: They semi-automatically create a German gold standard of technicality across four domains . they propose two new models to exploit general- vs. domain-specific comparisons based on vector spaces .
Outcome: The proposed model outperforms previous methods in terms of general- vs. domain-specific comparisons.
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)

Copied to clipboard

Challenge: Using monolingual-only data, we can automate readability assessment and text simplification of simplified language.
Approach: They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data.
Outcome: The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images.
Annotation and Automatic Classification of Aspectual Categories (P19-1)

Copied to clipboard

Challenge: Annotated resource for aspectual classification of German verb tokens in context.
Approach: They present a resource for aspectual classification of German verb tokens in their clausal context.
Outcome: The proposed resource is compared with previous work on German verb tokens using aspectual features compatible with the plurality of aspectual classifications.
The Syntactic Acceptability Dataset (Preview): A Resource for Machine Learning and Linguistic Analysis of English (2024.lrec-main)

Copied to clipboard

Challenge: Syntactic acceptance dataset is a resource being designed for syntax and computational linguistics research.
Approach: They propose to use the Syntactic Acceptability Dataset to examine the syntactical discourse.
Outcome: The proposed dataset is the largest of its kind that is publicly accessible.
A unified approach to sentence segmentation of punctuated text in many languages (2021.acl-long)

Copied to clipboard

Challenge: Existing tools for segmenting punctuated text in many languages are limited in their language coverage and evaluation is ad hoc.
Approach: They propose a new context-based modeling approach that can be trained on noisily-annotated data.
Outcome: The proposed model exceeds baselines set by existing methods on English corpora and performs well on average on new multilingual evaluation set.
Does It Capture STEL? A Modular, Similarity-based Linguistic Style Evaluation Framework (2021.emnlp-main)

Copied to clipboard

Challenge: linguistic style is an integral part of natural language, but evaluation methods for style measures are rare, often task-specific and usually do not control for content.
Approach: They propose a modular, fine-grained and content-controlled similarity-based STyle EvaLuation framework to test the performance of any model that can compare two sentences on style.
Outcome: The proposed model outperforms simple versions of commonly used style measures like 3-grams, punctuation frequency and LIWC-based approaches.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations