Expanding Abbreviations in a Strongly Inflected Language: Are Morphosyntactic Tags Sufficient? (L18-1)
Copied to clipboard
| Challenge: | In this paper, the problem of recovery of morphological information lost in abbreviated forms is addressed . correct inflected form of expanded abbrevation can be deduced from context words . |
| Approach: | They propose a deep bidirectional LSTM network with tag embedding to predict abbreviated words . they train on 10 million words from the Polish Sejm Corpus and achieve 74.2% prediction accuracy . |
| Outcome: | The proposed model achieves 74.2% accuracy on a smaller but more general corpus of Polish words. |
Similar Papers
Experiments with ad hoc ambiguous abbreviation expansion (D19-62)
Copied to clipboard
| Challenge: | ad hoc abbreviations are difficult to interpret for patients and nonspecialists. |
| Approach: | They propose to use morphologically annotated medical notes to expand ad hoc abbreviations without using additional domain resources. |
| Outcome: | The proposed methods outperform the previously proposed methods on Polish data but can be used for other languages. |
Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Using morphosyntactic tools, we train and distribute tools for approximately one thousand languages. |
| Approach: | They train and distribute morphosyntactic tools for approximately one thousand languages. |
| Outcome: | The results show that the tools generalize well across rare and common forms alike. |
Morphological Inflection: A Reality Check (2023.acl-long)
Copied to clipboard
| Challenge: | Morphological inflection is a popular task in sub-word NLP with practical and cognitive applications. |
| Approach: | They propose new methods to analyze data sets and evaluate their generalization abilities to better reflect likely use-cases. |
| Outcome: | The proposed methods improve generalizability and reliability of results and improve generalization abilities. |
An Extended Sequence Tagging Vocabulary for Grammatical Error Correction (2023.findings-eacl)
Copied to clipboard
| Challenge: | Current sequence-to-sequence and sequence-tagging approaches treat GEC as a machine-translation problem. |
| Approach: | They propose to introduce specialised tags for spelling correction and morphological inflection using the SymSpell and LemmInflect algorithms. |
| Outcome: | The proposed approach outperforms existing methods on the BEA benchmark. |
Structured abbreviation expansion in context (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Ad hoc abbreviations are commonly found in informal communication channels that favor shorter messages. |
| Approach: | They propose to reverse ad hoc abbreviations in context to recover normalized, expanded versions of abbrevated messages. |
| Outcome: | The proposed method can recover normalized, expanded abbreviations from text . it is similar to spelling correction, but requires more extensive work . |
Wikinflection Corpus: A (Better) Multilingual, Morpheme-Annotated Inflectional Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Inflectional corpora with annotated morpheme boundaries are scarce in the NLP community . a generated, multilingual inflectional lexicon with morphological features is not as good as UniMorph's . |
| Approach: | They evaluate a multilingual inflectional corpus with morpheme boundaries from the English Wiktionary and the UniMorph project's inflection corpus. |
| Outcome: | The generated Wikinflection corpus is not as good as UniMorph's, but extracts significant amount of words from the intersection of the two corpora. |
Inflecting When There’s No Majority: Limitations of Encoder-Decoder Neural Networks as Cognitive Models for German Plurals (2020.acl-main)
Copied to clipboard
| Challenge: | Encoder-decoder models can be used to generalize to inflectional morphology and generalize new words, but they fail on tasks like German number inflection, where infrequent suffixes like /-s/ can still be productively generalized. |
| Approach: | They propose to use a dataset to collect data from German speakers to examine whether ED models can generalize the most frequently produced plural class. |
| Outcome: | The proposed model does not show human-like variability or ‘regular’ extension of other plural markers. |
When do Generative Query and Document Expansions Fail? A Comprehensive Study Across Methods, Retrievers, and Datasets (2024.findings-eacl)
Copied to clipboard
| Challenge: | Using large language models (LMs) for query or document expansion can improve generalization in information retrieval. |
| Approach: | They conduct the first comprehensive analysis of large language models (LMs) for query or document expansion. |
| Outcome: | The proposed expansions improve retrieval performance for weaker models but harm stronger models. |
Morphology Matters: A Multilingual Language Modeling Analysis (2021.tacl-1)
Copied to clipboard
| Challenge: | Existing studies on inflectional morphology disagree on whether or not it makes languages harder to model. |
| Approach: | They propose to use a corpus of 145 Bible translations in 92 languages to investigate whether inflectional morphology makes languages harder to model. |
| Outcome: | The proposed model trains with linguistically motivated subword segmentation strategies and reduces the impact of morphology on language modeling. |
Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Prior work suggests hierarchical organization where different layers specialize in capturing distinct levels of linguistic structure. |
| Approach: | They probe 25 models from BERT Base to Qwen2.5-7B focusing on linguistic properties: lexical identity and inflectional features. |
| Outcome: | The proposed model maintains inflectional features across layers while trading off lexical identity for compact, predictive representations. |