Unsupervised Neologism Normalization Using Embedding Space Mapping (D19-55)

Copied to clipboard

Challenge: Neologisms refer to recent expressions that are specific to certain entities or events, but have not yet been accepted into mainstream language.
Approach: They propose an unsupervised approach for detecting and normalizing neologisms in social media content without relying on parallel training data.
Outcome: The proposed method detects neologisms and normalizes them to canonical words without training data.

Similar Papers

A Taxonomy for In-depth Evaluation of Normalization for User Generated Content (L18-1)

Copied to clipboard

Challenge: Existing taxonomies for lexical normalization are not suitable for the task of normalization since the categories are substantially different.
Approach: They propose a taxonomy of error categories for lexical normalization . they annotate a recent normalization dataset and read a near-perfect agreement .
Outcome: The proposed taxonomy is based on a recent normalization dataset and it performs well.
Handling Normalization Issues for Part-of-Speech Tagging of Online Conversational Text (L18-1)

Copied to clipboard

Challenge: a new approach to POS tagging noisy user generated text is proposed . word embeddings are trained on a noisy corpus to address both normalization and POS.
Approach: They propose to use word embeddings to normalize text before tagging it, while a gated neural network based tagger handles the remaining errors.
Outcome: The proposed approach normalizes some errors before tagging, while a gated neural network handles the remaining errors.
Synthetic Data for English Lexical Normalization: How Close Can We Get to Manually Annotated Data? (2020.lrec-1)

Copied to clipboard

Challenge: Social media data is a valuable data resource for natural language processing tasks.
Approach: They propose to adapt input text to a more standard form, a task also referred to as normalization.
Outcome: The proposed system scores 94.29 accuracy on the test data compared to 95.22 when trained on human-annotated data.
An In-depth Analysis of the Effect of Lexical Normalization on the Dependency Parsing of Social Media (D19-55)

Copied to clipboard

Challenge: Existing natural language processing tools are focused on standard texts, but performance drops when used on a different domain.
Approach: They analyze the effect of manual and automatic lexical normalization for dependency parsing . they conclude that automatic normalization scores close to manually annotated normalization .
Outcome: The proposed approach improves performance on social media data for many tasks . it is unclear which replacements have the most impact and what weaknesses exist in the system .
Sentence Meta-Embeddings for Unsupervised Semantic Textual Similarity (2020.acl-main)

Copied to clipboard

Challenge: Existing word embeddings combine complementary strengths of their components to achieve unsupervised semantic similarity (STS).
Approach: They propose to ensemble pre-trained sentence encoders into sentence meta-embeddings to achieve unsupervised Semantic Textual Similarity (STS) they adapt dimensionality reduction, generalized Canonical Correlation Analysis and cross-view auto-encoders to their work.
Outcome: The proposed method achieves 3.7% to 6.4% Pearson’s r over single-source word embeddings on the STS Benchmark and on the StS12-STS16 datasets.
MoNoise: A Multi-lingual and Easy-to-use Lexical Normalization Tool (P19-3)

Copied to clipboard

Challenge: In this paper, we demonstrate the online demo and command line interface of a lexical normalization system (MoNoise) for a variety of languages.
Approach: They propose to bundle seven datasets in six languages to form a new benchmark and a novel evaluation metric which is particularly suitable for cross-dataset comparisons.
Outcome: The proposed model is based on the original word and features from the original language for each normalization candidate.
Normalizing Non-canonical Turkish Texts Using Machine Translation Approaches (P19-2)

Copied to clipboard

Challenge: a study using non-canonical text normalization shows that it can surpass the current best performing system by a large margin.
Approach: They propose a fully automated, context-aware machine translation approach with fewer stages of processing.
Outcome: The proposed approach surpasses the current best-performing system by a large margin . the proposed method is more data-hungry and more data sensitive than other methods .
Diachronic word embeddings and semantic shifts: a survey (C18-1)

Copied to clipboard

Challenge: Existing methods for tracing time-related semantic shifts with word embedding models lack the cohesion, common terminology and shared practices of more established areas of natural language processing.
Approach: They propose several axes along which these methods can be compared and propose a framework for comparison.
Outcome: The proposed methods are compared with existing methods and outline their main challenges and potential applications.
Interpretable Text Embeddings and Text Similarity Explanation: A Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Text embeddings are a fundamental component in many NLP tasks, but their interpretation and explanation remain challenging.
Approach: They propose a framework for interpretable text embeddings and text similarity explanation . they characterize the main ideas, approaches, and trade-offs and discuss lessons learned .
Outcome: The proposed methods are compared with existing models and compare them with existing ones.
Semi-supervised Contextual Historical Text Normalization (2020.acl-main)

Copied to clipboard

Challenge: Historical text normalization is the task of mapping historical word forms to their modern counterparts.
Approach: They propose to use a generative normalization model to obtain contextualization from the target-side language model.
Outcome: et al., 2018) show that the most effective approach reduces manual normalization time and manual training costs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations