Challenge: In this paper, the problem of recovery of morphological information lost in abbreviated forms is addressed . correct inflected form of expanded abbrevation can be deduced from context words .
Approach: They propose a deep bidirectional LSTM network with tag embedding to predict abbreviated words . they train on 10 million words from the Polish Sejm Corpus and achieve 74.2% prediction accuracy .
Outcome: The proposed model achieves 74.2% accuracy on a smaller but more general corpus of Polish words.

Similar Papers

Experiments with ad hoc ambiguous abbreviation expansion (D19-62)

Copied to clipboard

Challenge: ad hoc abbreviations are difficult to interpret for patients and nonspecialists.
Approach: They propose to use morphologically annotated medical notes to expand ad hoc abbreviations without using additional domain resources.
Outcome: The proposed methods outperform the previously proposed methods on Polish data but can be used for other languages.
Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages (2020.lrec-1)

Copied to clipboard

Challenge: Using morphosyntactic tools, we train and distribute tools for approximately one thousand languages.
Approach: They train and distribute morphosyntactic tools for approximately one thousand languages.
Outcome: The results show that the tools generalize well across rare and common forms alike.
Morphological Inflection: A Reality Check (2023.acl-long)

Copied to clipboard

Challenge: Morphological inflection is a popular task in sub-word NLP with practical and cognitive applications.
Approach: They propose new methods to analyze data sets and evaluate their generalization abilities to better reflect likely use-cases.
Outcome: The proposed methods improve generalizability and reliability of results and improve generalization abilities.
An Extended Sequence Tagging Vocabulary for Grammatical Error Correction (2023.findings-eacl)

Copied to clipboard

Challenge: Current sequence-to-sequence and sequence-tagging approaches treat GEC as a machine-translation problem.
Approach: They propose to introduce specialised tags for spelling correction and morphological inflection using the SymSpell and LemmInflect algorithms.
Outcome: The proposed approach outperforms existing methods on the BEA benchmark.
Structured abbreviation expansion in context (2021.findings-emnlp)

Copied to clipboard

Challenge: Ad hoc abbreviations are commonly found in informal communication channels that favor shorter messages.
Approach: They propose to reverse ad hoc abbreviations in context to recover normalized, expanded versions of abbrevated messages.
Outcome: The proposed method can recover normalized, expanded abbreviations from text . it is similar to spelling correction, but requires more extensive work .
Wikinflection Corpus: A (Better) Multilingual, Morpheme-Annotated Inflectional Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Inflectional corpora with annotated morpheme boundaries are scarce in the NLP community . a generated, multilingual inflectional lexicon with morphological features is not as good as UniMorph's .
Approach: They evaluate a multilingual inflectional corpus with morpheme boundaries from the English Wiktionary and the UniMorph project's inflection corpus.
Outcome: The generated Wikinflection corpus is not as good as UniMorph's, but extracts significant amount of words from the intersection of the two corpora.
Inflecting When There’s No Majority: Limitations of Encoder-Decoder Neural Networks as Cognitive Models for German Plurals (2020.acl-main)

Copied to clipboard

Challenge: Encoder-decoder models can be used to generalize to inflectional morphology and generalize new words, but they fail on tasks like German number inflection, where infrequent suffixes like /-s/ can still be productively generalized.
Approach: They propose to use a dataset to collect data from German speakers to examine whether ED models can generalize the most frequently produced plural class.
Outcome: The proposed model does not show human-like variability or ‘regular’ extension of other plural markers.
When do Generative Query and Document Expansions Fail? A Comprehensive Study Across Methods, Retrievers, and Datasets (2024.findings-eacl)

Copied to clipboard

Challenge: Using large language models (LMs) for query or document expansion can improve generalization in information retrieval.
Approach: They conduct the first comprehensive analysis of large language models (LMs) for query or document expansion.
Outcome: The proposed expansions improve retrieval performance for weaker models but harm stronger models.
Morphology Matters: A Multilingual Language Modeling Analysis (2021.tacl-1)

Copied to clipboard

Challenge: Existing studies on inflectional morphology disagree on whether or not it makes languages harder to model.
Approach: They propose to use a corpus of 145 Bible translations in 92 languages to investigate whether inflectional morphology makes languages harder to model.
Outcome: The proposed model trains with linguistically motivated subword segmentation strategies and reduces the impact of morphology on language modeling.
Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models (2026.acl-long)

Copied to clipboard

Challenge: Prior work suggests hierarchical organization where different layers specialize in capturing distinct levels of linguistic structure.
Approach: They probe 25 models from BERT Base to Qwen2.5-7B focusing on linguistic properties: lexical identity and inflectional features.
Outcome: The proposed model maintains inflectional features across layers while trading off lexical identity for compact, predictive representations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations