Papers with Lemmatization

5 papers
How low is too low? A monolingual take on lemmatisation in Indian languages (2021.naacl-main)

Copied to clipboard

Challenge: Prior work on ML based lemmatization focused on high resource languages, where data sets (word forms) are readily available.
Approach: They propose to use neural methods to relate inflected forms of words to their dictionary form to reduce the sparse data problem.
Outcome: The proposed methods can give competitive accuracy even in low resource setting.
Data Augmentation for Context-Sensitive Neural Lemmatization Using Inflection Tables and Raw Text (N19-1)

Copied to clipboard

Challenge: Using context-sensitive approaches to lemmatization can improve accuracy on unseen and unseense words.
Approach: They propose to use inflection tables and Wikipedia sentences to train a lemmatizer with little or no labeled corpus data to combine type-based learning with context.
Outcome: The proposed model generalizes from unambiguous examples, improving overall and especially on unseen words.
Analysing cross-lingual transfer in lemmatisation for Indian languages (2020.coling-main)

Copied to clipboard

Challenge: Inference-based scripts such as Abjad are difficult for cross-lingual models to learn in extremely low resource scenarios.
Approach: They evaluate cross-lingual approaches for low resource languages and compare their performance against other models using different linguistic factors.
Outcome: The proposed model on six low resource languages from two different families is compared with monolingual models on morphologically rich Indian languages.
Lemma Dilemma: On Lemma Generation Without Domain- or Language-Specific Training Data (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) can generate lemmas in context without prior fine-tuning.
Approach: They compare in-context lemma generation with traditional fully supervised approaches . they use encoder-only supervised methods and cross-lingual methods .
Outcome: The proposed model outperforms the traditional fully supervised approach in the context of lemmatization tasks.
Lemmatization as a Classification Task: Results from Arabic across Multiple Genres (2025.emnlp-main)

Copied to clipboard

Challenge: Existing tools for lemmatization in morphologically rich languages with ambiguous orthography face inconsistent standards and limited genre coverage.
Approach: They propose two new approaches that frame lemmatization as classification into a Lemma-POS-Gloss tagset, leveraging machine translation and semantic clustering.
Outcome: The proposed models perform better than existing models and are more interpretable, the authors show.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations