Papers by Jose Camacho-Collados

46 papers
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding.
Approach: They propose to use sense-annotated corpora for supervised Word Sense Disambiguation.
Outcome: The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available.
Construction Artifacts in Metaphor Identification Datasets (2023.emnlp-main)

Copied to clipboard

Challenge: Existing metaphor identification datasets can be gamed by completely ignoring the potential metaphorical expression or the context in which it occurs.
Approach: They show that existing metaphor identification datasets can be gamed by fully ignoring the potential metaphorical expression or the context in which it occurs.
Outcome: The proposed system can be gamed by fully ignoring the potential metaphorical expression or the context in which it occurs.
Improving Cross-Lingual Word Embeddings by Meeting in the Middle (D18-1)

Copied to clipboard

Challenge: Cross-lingual word embeddings are becoming increasingly important in multilingual NLP.
Approach: They propose to apply an additional transformation after initial alignment to align two disjoint monolingual vector spaces.
Outcome: The proposed approach outperforms state-of-the-art models in monolingual and cross-lingual evaluation tasks.
Named Entity Recognition in Twitter: A Dataset and Analysis on Short-Term Temporal Shifts (2022.aacl-main)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a longstanding NLP task that consists of identifying an entity in a sentence or document.
Approach: They construct a dataset of seven entity types annotated over 11,382 tweets . they provide a set of language model baselines and analyze the performance of the model .
Outcome: The proposed dataset contains seven entity types annotated over 11,382 tweets . the authors focus on short-term degradation of NER models over time and strategies to fine-tune a language model over different periods .
On the Robustness of Unsupervised and Semi-supervised Cross-lingual Word Embedding Learning (2020.lrec-1)

Copied to clipboard

Challenge: Cross-lingual word embeddings are vector representations of words in different languages where words with similar meaning are represented by similar vectors, regardless of the language.
Approach: They propose to evaluate multiple cross-lingual word embedding models and compare their strengths and limitations to evaluate their effectiveness.
Outcome: The proposed models perform well with noisy text and language pairs with major differences.
XL-WiC: A Multilingual Benchmark for Evaluating Semantic Contextualization (2020.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation benchmarks for assessing distinct meanings of words are tied to sense inventories, restricting their usage to knowledge-based representation techniques.
Approach: They propose a multilingual benchmark that models distinct meanings of words in English . they use a binary disambiguation task with gold standards in 12 new languages .
Outcome: The proposed model can model distinct meanings of words in English even when no tagged instances are available for a target language.
The interplay between lexical resources and Natural Language Processing (N18-6)

Copied to clipboard

Challenge: linguistic, world and common sense knowledge is an important research area, but processing and storing it in lexical resources is not a straightforward task.
Approach: They propose to use NLP methods to help process of constructing and enriching lexical resources and the use of lexicals for improving NLP applications.
Outcome: The proposed approach aims to speed up and/or ease up the process of resource curation and enrichment.
Pun Unintended: LLMs and the Illusion of Humor Understanding (2025.emnlp-main)

Copied to clipboard

Challenge: Existing models for pun detection lack nuanced grasp typical of human interpretation.
Approach: They analyze existing pun detection benchmarks and human evaluation across recent LLMs to find subtle changes in puns that mislead LLM.
Outcome: The proposed models lack the nuance typical of human interpretation and lack the depth of their analysis to detect puns.
XLM-T: Multilingual Language Models in Twitter for Sentiment Analysis and Beyond (2022.lrec-1)

Copied to clipboard

Challenge: Language models are ubiquitous in NLP, but current analyses focus on (multilingual variants of) standard benchmarks and task-specific corpora as multilingual signals.
Approach: They propose a model to train and evaluate multilingual language models in Twitter using a set of Twitter datasets in eight different languages and a XLM-T model.
Outcome: The proposed model trains and evaluates multilingual models on Twitter.
Interpretable Emoji Prediction via Label-Wise Attention LSTMs (D18-1)

Copied to clipboard

Challenge: Emojis are the evolution of characterbased emoticons and are used to express ideas about a myriad of topics.
Approach: They propose a label-wise attention mechanism to better understand emoji prediction . they propose to model e-mails with eojis and then label them based on their meaning .
Outcome: The proposed model improves over baselines and does particularly well when predicting infrequent emojis.
T-NER: An All-Round Python Library for Transformer-based Named Entity Recognition (2021.eacl-demos)

Copied to clipboard

Challenge: Language model (LM) pretraining has led to consistent improvements in many downstream tasks, including named entity recognition (NER).
Approach: They propose a Python library for NER LM finetuning that facilitates cross-domain and cross-lingual generalization of LMs finetuned on NER.
Outcome: The proposed library outperforms LMs trained on NERs in cross-domain and cross-lingual generalization tests on nine datasets.
False Friends or Cognates? A Cross-lingual Semantic Ambiguity Evaluation for Galician, Portuguese and Spanish (2026.acl-long)

Copied to clipboard

Challenge: Closely related languages exhibit a high degree of lexical and orthographic similarity, which can facilitate cross-lingual understanding but also give rise to systematic semantic ambiguity.
Approach: They introduce six cross-lingual datasets that are manually or semi-automatically generated and are able to identify and process false friends among these languages.
Outcome: The proposed models can identify and process false friends among Galician, Portuguese, and Spanish.
TweetEval: Unified Benchmark and Comparative Evaluation for Tweet Classification (2020.findings-emnlp)

Copied to clipboard

Challenge: Modern NLP systems are typically ill-equipped when applied to noisy user-generated text.
Approach: They propose a new evaluation framework consisting of seven Twitter-specific classification tasks.
Outcome: The proposed framework is based on seven heterogeneous Twitter-specific classification tasks.
Understanding LLM Performance Degradation in Multi-Instance Processing: The Roles of Instance Count and Context Length (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are used for processing multiple documents or analysis over a number of instances.
Approach: They perform a comprehensive evaluation of the multi-instance processing ability of LLMs for tasks in which they excel individually.
Outcome: The proposed model performs well on tasks in which it excels individually.
TempoWiC: An Evaluation Benchmark for Detecting Meaning Shift in Social Media (2022.coling-1)

Copied to clipboard

Challenge: Language models are often clean and time-invariant, and do little to no account of social media usage.
Approach: They propose a benchmark to accelerate research in social media-based meaning shift.
Outcome: The proposed benchmark is aimed at accelerating research in social media-based meaning shift.
SuperTweetEval: A Challenging, Unified and Heterogeneous Benchmark for Social Media NLP Research (2023.findings-emnlp)

Copied to clipboard

Challenge: specialised language models (LMs) have shown to exhibit lower perplexity and higher downstream performance across the board.
Approach: They propose a benchmark for NLP evaluation in social media, SuperTweetEval.
Outcome: The proposed benchmark shows that social media models perform better when compared to general-purpose models, metrics and benchmarks.
BERT is to NLP what AlexNet is to CV: Can Pre-Trained Language Models Identify Analogies? (2021.acl-long)

Copied to clipboard

Challenge: Analogies play a central role in human commonsense reasoning.
Approach: They analyze the capabilities of transformer-based language models on an unsupervised task . they find off-the-shelf language models can identify analogies to a certain extent .
Outcome: The proposed language models outperform word embedding models on an unsupervised task . the best results were obtained with GPT-2 and RoBERTa .
Distilling Relation Embeddings from Pretrained Language Models (2021.emnlp-main)

Copied to clipboard

Challenge: Pre-trained language models capture a surprisingly rich amount of lexical knowledge, but it is unclear to what extent relation embeddings can be used to encode relational knowledge.
Approach: They found that word vector differences capture lexical relations . relationship embeddings can be used to encode relational knowledge .
Outcome: The results are highly competitive on analogy (unsupervised) and relation classification (supervised) benchmarks, even without any task-specific fine-tuning.
Embeddings in Natural Language Processing (2020.coling-tutorials)

Copied to clipboard

Challenge: Embeddings have been a key topic of interest in NLP for the past decade . a quick warm-up introduction to NLP and why it is important to have a semantic comprehension of texts .
Approach: This tutorial will provide a high-level synthesis of the main embedding techniques in NLP . it will start with word embedds and then move to other types of embeddable vectors .
Outcome: This tutorial will provide a high-level synthesis of the main embedding techniques in NLP . it will start with word embedds and move to other types of embeddable representations .
Back to the Basics: A Quantitative Analysis of Statistical and Graph-Based Term Weighting Schemes for Keyword Extraction (2021.emnlp-main)

Copied to clipboard

Challenge: Term weighting schemes are widely used in Natural Language Processing and Information Retrieval.
Approach: They perform an exhaustive and large-scale empirical comparison of term weighting methods in the context of keyword extraction using tf-idf.
Outcome: The proposed methods have advantages over tf-idf, and qualitative differences between them.
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis (2024.naacl-long)

Copied to clipboard

Challenge: Existing datasets for hate speech detection neglect the cultural diversity within a single language.
Approach: They propose a CR**oss-cultural **E**nglish **Hate* speech dataset that uses culturally hateful keywords to identify posts from four countries plus the United States.
Outcome: The proposed dataset shows that only 56.2% of the posts in CREHate achieve consensus among all countries, with the highest pairwise label difference rate of 26%.
Language Models for Text Classification: Is In-Context Learning Enough? (2024.lrec-main)

Copied to clipboard

Challenge: Existing research on text classification models with prompts is limited in scale and lacks understanding of how these methods compare to more established methods.
Approach: They compare the performance of large and smaller language models with prompts to achieve state-of-the-art performance in many NLP tasks.
Outcome: The proposed models outperform the more standard approaches in binary, multiclass, and multilabel tasks in a large scale evaluation of 16 text classification datasets.
WiC-TSV: An Evaluation Benchmark for Target Sense Verification of Words in Context (2021.eacl-main)

Copied to clipboard

Challenge: Existing benchmarks for Word Sense Disambiguation are limited to those systems in which sense distinctions are defined according to an underlying sense inventory.
Approach: They propose a framework for Target Sense Verification of Words in Context which grounds its uniqueness as binary classification task and independent of external sense inventories.
Outcome: The proposed framework is highly flexible for evaluation of diverse models and systems in and across domains.
Do Large Language Models Understand Mansplaining? Well, Actually... (2024.lrec-main)

Copied to clipboard

Challenge: Gender bias has been studied by the NLP community, but other variations of it, such as mansplaining, have received little attention.
Approach: They propose to analyze a corpus of 886 mansplaining stories experienced by women and examine how Large Language Models can understand and identify mansplaiting.
Outcome: The proposed models reproduce some of the social patterns behind mansplaining situations by praising men for giving unsolicited advice to women.
Twitter Topic Classification (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to identify topics from posts are difficult to interpret and can differ from corpus to corpus.
Approach: They propose a task based on tweet topic classification and release two datasets that can be used to train and test models.
Outcome: The proposed task is based on two datasets from recent time periods and provides training and testing data.
COVID-19 and Misinformation: A Large-Scale Lexical Analysis on Twitter (2021.acl-srw)

Copied to clipboard

Challenge: Social media is used by individuals and organisations as a platform to spread misinformation.
Approach: They compile a large corpus of tweets related to coronavirus and perform an analysis to discover patterns with respect to vocabulary usage.
Outcome: The proposed model based on lexical features is effective in identifying misinformation-related tweets with accuracy over 80%.
A RelEntLess Benchmark for Modelling Graded Relations between Named Entities (2024.eacl-long)

Copied to clipboard

Challenge: Existing Knowledge Graphs do not cover graded relations, yet they are difficult to draw a line between those that satisfy them and those that do not.
Approach: They propose a benchmark in which entity pairs have to be ranked according to how much they satisfy a given graded relation.
Outcome: The proposed model outperforms several publicly available LLMs and closed conversational models.
Probing Relational Knowledge in Language Models via Word Analogies (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have focused on probing relational knowledge by filling the blanks in pre-defined prompts such as “The capital of France is —” but these are affected by the co-occurrence of target relation words and entities in the pre-training corpus.
Approach: They extend probing methodologies by using analogical proportions as a proxy to probe relational knowledge in transformer-based PLMs without directly presenting the desired relation.
Outcome: The proposed methods are extremely accurate at (1) and (2), but have room for improvement for (3).
Go Simple and Pre-Train on Domain-Specific Corpora: On the Role of Training Data for Text Classification (2020.coling-main)

Copied to clipboard

Challenge: Pre-trained language models provide the foundations for state-of-the-art performance across a wide range of natural language processing tasks, including text classification.
Approach: They compare the performance of a linear classifier based on word embeddings with a pre-trained language model, i.e., BERT, across a wide range of datasets and classification tasks.
Outcome: The proposed method outperforms baselines in standard datasets with large training sets, but in settings with small training datasets it performs better.
WiC: the Word-in-Context Dataset for Evaluating Context-Sensitive Meaning Representations (N19-1)

Copied to clipboard

Challenge: Existing word embeddings cannot model the dynamic nature of words’ semantics, i.e., the property of words to correspond to potentially different meanings.
Approach: They propose a large-scale Word in Context dataset, called WiC, which is curated by experts and can be used to evaluate context-sensitive representations.
Outcome: The proposed models outperform the standard evaluation dataset for the purpose and highlight their shortcomings.
Morables: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables (2025.emnlp-main)

Copied to clipboard

Challenge: Literature-based benchmarks provide a compelling framework for evaluating LLMs' capacity for complex abstract reasoning and inference.
Approach: They propose a novel moral reasoning benchmark built from fables and short stories that uses adversarial variants to stress-test model robustness.
Outcome: The proposed model outperforms models on fables and short stories, but is susceptible to adversarial manipulation and rely on superficial patterns rather than true moral reasoning.
An Empirical Comparison of LM-based Question and Answer Generation Methods (2023.findings-acl)

Copied to clipboard

Challenge: Question and answer generation (QAG) is a task of generating question-answer pairs given a context.
Approach: They propose to leverage sequence-to-sequence language model fine-tuning to generate question-answer pairs given a context.
Outcome: The proposed model outperforms other more convoluted approaches in the end-to-end model and is computationally light at both training and inference times.
Relational Word Embeddings (P19-1)

Copied to clipboard

Challenge: Existing approaches to learn word embeddings rely on external knowledge bases . however, they are limited by the amount of available relational knowledge .
Approach: They propose to encode relational knowledge in a separate word embedding . this is complementary to a standard word embedded from co-occurrence statistics .
Outcome: The proposed word embedding is complementary to a standard word embed.
Question Answering over Tabular Data with DataBench: A Large-Scale Empirical Evaluation of LLMs (2024.lrec-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are showing emerging abilities, but they are not large enough to assess their capabilities.
Approach: They propose a benchmark that compares large language models with open and closed source models.
Outcome: The proposed benchmark compares open and closed-source models with open-source and closed source models.
The Interplay between Metaphors and NLP (2026.acl-tutorials)

Copied to clipboard

Challenge: This tutorial will provide an overview of the metaphor processing field.
Approach: This tutorial will provide an overview of the metaphor processing field . it will focus on recent directions opened by LLMs for metaphor interpretation .
Outcome: The tutorial will discuss the influence of various metaphor theories on the creation of annotated resources and models.
Evaluating Short-Term Temporal Fluctuations of Social Biases in Social Media Data and Masked Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Social biases such as gender or racial biase are reported in language models . a recent study has shown that MLMs encode discriminatory social biase .
Approach: They analyse temporal corpora of MLMs trained on chronologically ordered temporal snapshots . they find that gender and racial biases are encoded in MLM models .
Outcome: The proposed model identifies gender biases in MLMs but most remain stable over time . gender bias is associated with higher likelihood scores in some demographic groups .
A Practical Toolkit for Multilingual Question and Answer Generation (2023.acl-demo)

Copied to clipboard

Challenge: Generating questions and answers from text is a challenging task due to the expected structured output.
Approach: They propose an online service for multilingual QAG along with a python package for model fine-tuning, generation, and evaluation.
Outcome: The proposed model is available in eight languages and can be used online or locally via lmqg.
Automatic Extraction of Metaphoric Analogies from Literary Texts: Task Formulation, Dataset Construction, and Evaluation (2025.coling-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have shown to be difficult to extract metaphors from free text because they can involve some implicit concepts and link dissimilar concepts.
Approach: They compare the ability of large language models to extract metaphors from literary texts using domain experts.
Outcome: The proposed models can extract metaphors from literary texts without using domain experts.
TweetTER: A Benchmark for Target Entity Retrieval on Twitter without Knowledge Bases (2024.lrec-main)

Copied to clipboard

Challenge: Entity linking is a well-established task in NLP consisting of associating entity mentions with entries in a knowledge base.
Approach: They propose a benchmark that reframes entity linking as a binary entity retrieval task and uses a knowledge base to evaluate model performance.
Outcome: The proposed benchmark aims to bridge the challenges in entity linking in noisy domains such as social media.
Multilingual Topic Classification in X: Dataset and Analysis (2024.emnlp-main)

Copied to clipboard

Challenge: Social media platforms such as X (Twitter), Snapchat and Instagram provide an environment for content creation and information sharing.
Approach: They propose a multilingual dataset featuring tweet topic classification in four languages . they leverage X-Topic to perform cross-linguistic and multilingual analysis .
Outcome: The proposed dataset includes topics in four languages and is useful for cross-linguistic analysis and the development of robust multilingual models.
The Problem of Ambiguity in Table Question Answering (2026.findings-eacl)

Copied to clipboard

Challenge: Existing approaches to question answering on tabular data have limited capabilities due to ambiguousness inherent to tabular datasets.
Approach: They propose to use large language models to answer questions on tabular data by analyzing tabular tables and detecting ambiguity.
Outcome: The proposed model can detect ambiguity in tabular data and provide an initial ground for a deeper discussion on how to approach it in the age of LLMs.
A Predictive Factor Analysis of Social Biases and Task-Performance in Pretrained Masked Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Various types of social biases have been reported with pretrained Masked Language Models (MLMs) in prior work.
Approach: They conduct a comprehensive study on 39 pretrained MLMs to examine their model factors and their social biases.
Outcome: The proposed model factors influence social biases learned by an MLM and their downstream task performance.
Generative Language Models for Paragraph-Level Question Generation (2022.emnlp-main)

Copied to clipboard

Challenge: Powerful generative models have led to recent progress in question generation.
Approach: They propose a multilingual and multidomain benchmark for question generation that unifies existing datasets by converting them to a standard QG setting.
Outcome: The proposed benchmark unifies existing question answering datasets to a QG setting.
Don’t Neglect the Obvious: On the Role of Unambiguous Words in Word Sense Disambiguation (2020.emnlp-main)

Copied to clipboard

Challenge: Existing senseannotated corpora lack coverage of many instances in WordNet . however, unambiguous words make up a large portion of WordNet while being poorly covered in existing senseannnotated .
Approach: They propose a method to provide annotations for most unambiguous words in a large corpus by using a dataset.
Outcome: The proposed method improves on the original results on Word Sense Disambiguation (WSD) using pre-trained language models and propagation algorithms.
Efficient Multilingual Language Model Compression through Vocabulary Trimming (2023.findings-emnlp)

Copied to clipboard

Challenge: Multilingual language models (LMs) have become a powerful tool in NLP, especially for non-English languages.
Approach: They propose a method to reduce a multilingual LM vocabulary to a target language by deleting potentially irrelevant tokens from its vocabulary.
Outcome: The proposed method can retain the original performance of the multilingual LM while being considerably smaller in size than the original model.
Analysing Zero-Shot Readability-Controlled Sentence Simplification (2025.coling-main)

Copied to clipboard

Challenge: Text simplification (RCTS) models often depend on parallel corpora with readability annotations on both source and target sides.
Approach: They propose to use instruction-tuned large language models for zero-shot RCTS to reduce reliance on parallel corpora with readability annotations on both source and target sides.
Outcome: The proposed model can generate sentences with the desired readability, but the model's limitations and characteristics of the source sentences impede it.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations