Papers by Yova Kementchedjhieva
A systematic comparison of methods for low-resource dependency parsing on genuinely low-resource languages (D19-1)
Copied to clipboard
| Challenge: | Large annotated treebanks are available for only a tiny fraction of the world's languages, and there is a wealth of literature on strategies for parsing with few resources. |
| Approach: | They propose three strategies for improving low-resource parsers: data augmentation, cross-lingual training, and transliteration. |
| Outcome: | The proposed methods improve low-resource parsers by using data augmentation, cross-lingual training, and transliteration. |
Uncovering Probabilistic Implications in Typological Knowledge Bases (P19-1)
Copied to clipboard
| Challenge: | linguistic typology is concerned with mapping out the relationships between languages with structural and functional properties. |
| Approach: | They propose a computational model which identifies known and new linguistic universals and uncovers them worthy of further linguistic investigation. |
| Outcome: | The proposed model outperforms baselines and knowledge base baselines. |
The ApposCorpus: a new multilingual, multi-domain dataset for factual appositive generation (2020.coling-main)
Copied to clipboard
| Challenge: | appositives are phrases that appear next to a noun phrase and serve an explicative function. |
| Approach: | They propose a more realistic end-to-end definition of appositive generation with a dataset that spans four languages and two entity types. |
| Outcome: | The proposed model is non-trivial and leaves plenty of room for improvement. |
Adversarial Removal of Demographic Attributes Revisited (D19-1)
Copied to clipboard
| Challenge: | Several approaches have been proposed to learn classifiers that are invariant (unbiased with respect) to protected attributes. |
| Approach: | They propose to use a diagnostic classifier trained on a held-out subsample to find protected attributes for mention detection at above-chance levels. |
| Outcome: | The proposed classifier generalizes poorly to new in-domain and new domains, suggesting it relies on correlations specific to their particular data sample. |
John praised Mary because _he_? Implicit Causality Bias and Its Interaction with Explicit Cues in LMs (2021.findings-acl)
Copied to clipboard
| Challenge: | Psycholinguists have identified one such cue in the implicit causality bias of interpersonal verbs. |
| Approach: | They propose to use pre-trained language models to encode IC bias at inference time . they hypothesize that when a cause is explicitly stated, an incongruent IC biased leads to a delay in human processing. |
| Outcome: | The results suggest that pre-trained language models tend to prioritize lexical patterns over higher-order signals. |
Grammatical Error Correction through Round-Trip Machine Translation (2023.findings-eacl)
Copied to clipboard
| Challenge: | A decade ago the idea of using round-trip MT to guide grammatical error correction was not feasible due to the low quality of MT systems of the day. |
| Approach: | They propose to use round-trip machine translation to guide grammatical error correction to preserve meaning while mapping its surface form from one language into another. |
| Outcome: | The proposed system is re-examined across five languages and models of various sizes and yields consistent improvements. |
An Exploration of Encoder-Decoder Approaches to Multi-Label Classification for Legal and Biomedical Text (2023.findings-acl)
Copied to clipboard
| Challenge: | Standard methods for multi-label text classification rely on encoder-only pre-trained models . encoder decoder models have proven more effective in other classification tasks . |
| Approach: | They compare four methods for multi-label classification based on encoder-only models . they use a pre-trained model for multilabel text classification . |
| Outcome: | The proposed methods outperform encoder-only methods on complex datasets and labeling schemes. |
Why is unsupervised alignment of English embeddings from different algorithms so hard? (D18-1)
Copied to clipboard
| Challenge: | a new paper challenges word embedding algorithms to align independent English word embeds with 100% precision . authors show that when two different embeddables are used, they fail to do so . |
| Approach: | They propose to use unsupervised bilingual dictionary induction to study English-English alignments. |
| Outcome: | The proposed approach is more of a challenge than a technical contribution . it shows that the results challenge unsupervised bilingual dictionary induction algorithms . |
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation (2025.emnlp-main)
Copied to clipboard
| Challenge: | N-gram-based evaluation metrics are unreliable due to low correlation to human judgments. |
| Approach: | They propose a metric that rewards correct details and penalizes incorrect ones. |
| Outcome: | The proposed metric matches the performance of open-source LLM-based metrics in correlation to human judgments while being far more efficient. |
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for vision-language models treat compositionality and long-caption understanding in isolation. |
| Approach: | They analyze when compositional reasoning and long-caption understanding transfer across tasks and when this relationship fails. |
| Outcome: | The proposed model can generalize on poorly grounded captions and with strong visual grounding, while architectural choices can limit compositional learning. |
JEEM: Vision-Language Understanding in Four Arabic Dialects (2026.findings-eacl)
Copied to clipboard
Karima Kadaoui, Hanin Atwany, Hamdan Al-Ali, Abdelrahman Mohamed, Ali Mekky, Sergei Tilga, Natalia Fedorova, Ekaterina Artemova, Hanan Aldarmaki, Yova Kementchedjhieva
| Challenge: | Existing evaluation datasets feature Western-centric images and English text, while their non-English counterparts are often derived from the latter. |
| Approach: | They propose to evaluate Vision-Language Models (VLMs) on visual understanding across four Arabic-speaking countries: Jordan, The Emirates, Egypt, and Morocco. |
| Outcome: | The proposed model underperforms in visual understanding and dialect-specific generation across four Arabic-speaking countries. |
MuLan: A Study of Fact Mutability in Language Models (2024.naacl-short)
Copied to clipboard
| Challenge: | Pretrained and large language models encode factual knowledge, but factual information changes over time and mutates with the passage of time. |
| Approach: | They propose to use a model to evaluate the ability of English language models to anticipate time-contingency by comparing their models to a benchmark model. |
| Outcome: | The proposed model can predict the president of a country or the winner of sa championship in time, but it is difficult to update them due to their mutability. |
A Discriminative Latent-Variable Model for Bilingual Lexicon Induction (D18-1)
Copied to clipboard
| Challenge: | Existing methods for bilingual lexicon induction take advantage of word embeddings, but our model is not as efficient as previous work. |
| Approach: | They propose a discriminative latent-variable model for bilingual lexicon induction that combines the bipartite matching dictionary prior and an embedding-based approach. |
| Outcome: | The proposed model outperforms existing models on six language pairs and shows that it mitigates hubness problem. |
A Probabilistic Generative Model of Linguistic Typology (N19-1)
Copied to clipboard
| Challenge: | a generative model of languages based on principles-and-parameters posits that languages toggle on or off . linguistic typologists use a set of universal parameters to determine which languages toggle . we show that the correlation between parameters is significant, and that it is not enough to write down the set of parameters available to languages. |
| Approach: | They propose a generative model of language based on exponential-family matrix factorisation. |
| Outcome: | a linguistic model outperforms baseline models on predicting held-out features by exploiting similarities between languages and their features. |
Dynamic Forecasting of Conversation Derailment (2021.emnlp-main)
Copied to clipboard
| Challenge: | a pretrained language encoder can predict derailment in online conversations . this is a useful task for detecting and preventing abusive language . |
| Approach: | They extend a task to predict derailment in online conversations by using a pretrained language encoder. |
| Outcome: | The proposed task outperforms previous approaches in terms of performance and quality. |
LLMs Can Compensate for Deficiencies in Visual Representations (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a strong language backbone in vision-language models compensates for weak visual features by contextualizing or enriching them. |
| Approach: | They investigate whether strong language backbone compensates for weak visual features . they use CLIP-based vision encoders to perform controlled self-attention ablations . |
| Outcome: | The proposed model compensates for weak visual features by contextualizing or enriching them. |
Lost in Evaluation: Misleading Benchmarks for Bilingual Dictionary Induction (D19-1)
Copied to clipboard
| Challenge: | a quarter of the data consists of proper nouns, which can be hardly indicative of BDI performance, and there are pervasive gaps in the gold-standard targets. |
| Approach: | They examine the composition and quality of test sets for five different languages . they suggest future research avoids drawing conclusions from quantitative results . |
| Outcome: | The results show that a quarter of the data consists of proper nouns, which can be hardly indicative of BDI performance, and there are gaps in the gold-standard targets. |
PuzzLing Machines: A Challenge on Learning From Small Data (2020.acl-main)
Copied to clipboard
| Challenge: | a benchmark dataset of 81 languages is released to test deep neural models' human-like reasoning and generalization skills. |
| Approach: | They propose a challenge on learning from small data using Rosetta Stone puzzles from Linguistic Olympiads for high school students. |
| Outcome: | The proposed benchmark consists of Rosetta Stone puzzles from Linguistic Olympiads for high school students. |