Papers by Constantine Lignos

18 papers
Macro-Average: Rare Types Are Important Too (2021.naacl-main)

Copied to clipboard

Challenge: MT metrics trained on segment-level human judgments are inherently non-transparent and reflect undesirable biases.
Approach: They propose to use a type-based classifier metric to evaluate machine translation and compare it with a supervised and unsupervised one.
Outcome: The proposed model outperforms other models in indicating cross-lingual information retrieval task performance and shows that it can be used to compare supervised and unsupervised neural machine translation.
LR-Sum: Summarization for Less-Resourced Languages (2023.findings-acl)

Copied to clipboard

Challenge: LR-Sum contains human-written summaries for 40 languages, many of which are less-resourced.
Approach: They propose to use a permissively-licensed dataset to analyze human-written summaries for 40 languages.
Outcome: The proposed dataset contains human-written summaries for 40 languages . authors describe abstractive and extractive summarization experiments .
MasakhaNER: Named Entity Recognition for African Languages (2021.tacl-1)

Copied to clipboard

Challenge: (2020) African languages are underrepresented in existing natural language processing datasets, research, and tools due to lack of datasets and reproducible results.
Approach: They propose to create a dataset for named entity recognition (NER) in ten African languages.
Outcome: The results of the first large dataset for named entity recognition (NER) in ten African languages are released to inform future research on African NLP.
ParaNames 1.0: Creating an Entity Name Corpus for 400+ Languages Using Wikidata (2024.lrec-main)

Copied to clipboard

Challenge: ParaNames is a massively multilingual parallel name resource . it provides names for 16.8 million entities in over 400 languages .
Approach: They propose a massively multilingual parallel name resource with 140 million names . they use Wikidata to standardize the data and perform canonical name translation .
Outcome: The proposed resource is the largest of its type to date and performs well on 10 languages.
CoNLL#: Fine-grained Error Analysis and a Corrected Test Set for CoNLL-03 English (2024.lrec-main)

Copied to clipboard

Challenge: a glass ceiling for named entity recognition systems has been suggested for 2021 . however, the performance of the most popular NER benchmarks has plateaued since then . we investigate what NER models are still struggling with .
Approach: They perform a fine-grained evaluation of the model outputs by adding document annotations to the CoNLL-03 English dataset to identify lingering errors.
Outcome: The proposed model is able to correct errors and guide future work.
Toward More Meaningful Resources for Lower-resourced Languages (2022.findings-acl)

Copied to clipboard

Challenge: a new position paper examines how meaningful resources for lower-resourced languages should be developed in connection with the speakers of those languages.
Approach: They propose a position paper on how meaningful resources should be developed for lower-resourced languages . they examine the contents of Wikidata for a few lower-rsourced languages and examine quality issues .
Outcome: The proposed approach is based on the findings of a recent study on the use of multilingual resources in language technology development.
The Effectiveness of Morphology-aware Segmentation in Low-Resource Neural Machine Translation (2021.eacl-srw)

Copied to clipboard

Challenge: Current NMT systems typically operate at the level of subwords, causing problems of vocabulary sparsity.
Approach: They compare subword segmentation methods with morphologically-based methods in a low-resource setting . they find that no consistent and reliable differences emerge between the methods .
Outcome: The proposed methods outperform BPE in a low-resource translation setting.
Language Model Priors and Data Augmentation Strategies for Low-resource Machine Translation: A Case Study Using Finnish to Northern Sámi (2024.findings-acl)

Copied to clipboard

Challenge: a new study examines the use of monolingual data for improving low-resource machine translation.
Approach: They investigate ways of using monolingual data for improving low-resource machine translation.
Outcome: The proposed model can perform better on the target-side data without augmentation of parallel data.
TMR: Evaluating NER Recall on Tough Mentions (2021.eacl-srw)

Copied to clipboard

Challenge: a NER evaluation tool is available via a repository.
Approach: They propose to use Tough Mentions Recall to supplement traditional named entity recognition evaluation by examining recall on specific subsets of ”tough” mentions.
Outcome: The proposed metrics enable differentiation between otherwise similar-scoring systems and identify patterns in performance that would go unnoticed from overall precision, recall, and F1.
MetaMeme: A Dataset for Meme Template and Meta-Category Classification (2025.naacl-srw)

Copied to clipboard

Challenge: a new dataset for classifying memes by their template and communicative intent is presented.
Approach: They propose a new dataset for classifying memes by their template and communicative intent.
Outcome: The proposed method outperforms existing methods in classifying memes by their template and communicative intent.
Multilingual Open Text Release 1: Public Domain News in 44 Languages (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of permissively licensed text is being developed in 44 languages, many of which have limited existing text resources for natural language processing.
Approach: They propose to create a multilingual corpus containing text in 44 languages . they describe their process for collecting, filtering, and processing the data .
Outcome: The first release of the corpus contains over 2.8 million news articles and an additional 1 million short snippets published between 2001–2022 and collected from Voice of America news websites.
The Challenges of Optimizing Machine Translation for Low Resource Cross-Language Information Retrieval (D19-1)

Copied to clipboard

Challenge: Existing studies do not investigate the effectiveness of MT metrics in predicting performance of downstream IR models.
Approach: They examine the relationship between MT performance and IR quality in a CLIR-based system . they find that the choice of IR collection can significantly affect MT tuning decisions .
Outcome: The proposed model can predict CLIR performance better from MT quality, the authors show . the proposed model is based on a BLEU-based model with a bag of words constraint .
Borrowing or Codeswitching? Annotating for Finer-Grained Distinctions in Language Mixing (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of tweets annotated for codeswitching and borrowing between Spanish and English is presented . the annotation does not treat common “internet-speak” as codeswitched when used in an otherwise monolingual context.
Approach: They present a new corpus of tweets annotated for codeswitching and borrowing between Spanish and English.
Outcome: The proposed corpus contains 9,500 tweets annotated with codeswitches, borrowings, and named entities.
Detecting Unassimilated Borrowings in Spanish: An Annotated Corpus and Approaches to Modeling (2022.acl-long)

Copied to clipboard

Challenge: a corpus of Spanish newswire rich in unassimilated lexical borrowings is used to identify the language of a word.
Approach: They propose to annotate a corpus of Spanish newswire rich in unassimilated lexical borrowings and evaluate how models perform on this task.
Outcome: The proposed model outperforms models fed with subword embeddings and Transformer-based embeddables on the Spanish newswire corpus.
OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages (2025.emnlp-main)

Copied to clipboard

Challenge: Existing datasets are not consistently formatted and use a variety of chunk encodings (IOB, BIO, etc.), often without documentation.
Approach: They present OpenNER 1.0, a standardized collection of openly-available named entity recognition (NER) datasets.
Outcome: The proposed datasets correct annotation format issues and provide a structure that enables research in multilingual and multi-ontology NER.
QueryNER: Segmentation of E-commerce Queries (2024.lrec-main)

Copied to clipboard

Challenge: Prior work on aspect-value extraction has focused on extracting portions of a product title or query for narrowly defined aspects.
Approach: They propose a manually-annotated dataset and model for e-commerce query segmentation.
Outcome: The proposed model can recover from null and low recall queries with token and entity dropping.
SARAL: A Low-Resource Cross-Lingual Domain-Focused Information Retrieval System for Effective Rapid Document Triage (P19-3)

Copied to clipboard

Challenge: a new cross-lingual information retrieval system for low-resource languages is available in less-frequently-taught languages . a multilingual system can search for relevant information in a haystack of documents in swahili or Somali . human-driven approaches to this problem are complicated in 'low-resourced' languages aaron sagar: "the key role played by humans in triaging results is complicated"
Approach: They propose an end-to-end cross-lingual information retrieval system for low-resource languages . the system enables English speakers to search foreign language repositories using English queries . it summarizes the retrieved documents in English with respect to a particular information need .
Outcome: The proposed system achieves top performance in the most recent IARPA MATERIAL CLIR+summarization evaluations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations