Papers by Ekaterina Kochmar
Detecting Multiword Expression Type Helps Lexical Complexity Assessment (2020.lrec-1)
Copied to clipboard
| Challenge: | Multiword expressions (MWEs) represent lexemes that should be treated as single lexical units due to their idiosyncratic nature. |
| Approach: | They re-annotate a complex word identification shared task 2018 dataset . they find that a lexical complexity assessment system benefits from the information . |
| Outcome: | The proposed dataset provides valuable information for the text simplification community. |
A Fully Automated Pipeline for Conversational Discourse Annotation: Tree Scheme Generation and Labeling with Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have shown promise in automating discourse annotation for conversations. |
| Approach: | They propose a pipeline that uses large language models to construct and perform annotations using speech functions and the Switchboard-DAMSL taxonomies. |
| Outcome: | The proposed pipeline outperforms existing tree annotation schemes and can match or surpass human annotations while significantly reducing time required for annotation. |
Recursive Context-Aware Lexical Simplification (D19-1)
Copied to clipboard
| Challenge: | REC-LS is a system that can be used to perform a number of simplifications at once, but the results are sometimes ungrammatical and meaning can be changed, making the original text less clear and more complex. |
| Approach: | They propose a recursive context-aware lexical simplification architecture that takes previous simplification steps into account and makes use of the wider context when detecting the words in need of simplification. |
| Outcome: | The proposed system outperforms the current state-of-the-art systems in lexical simplification. |
Automatic Readability Assessment for Closely Related Languages (2023.findings-acl)
Copied to clipboard
| Challenge: | In recent years, the main focus of research on automatic readability assessment (ARA) has shifted towards using expensive deep learning-based methods with the primary goal of increasing models’ accuracy. |
| Approach: | They focus on how linguistic aspects such as mutual intelligibility or degree of language relatedness can improve ARA in a low-resource setting. |
| Outcome: | The inclusion of CrossNGO, a novel feature exploiting n-gram overlap, significantly improves the performance of ARA models compared to the use of off-the-shelf large multilingual language models alone. |
REFeREE: A REference-FREE Model-Based Metric for Text Simplification (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for text simplification lack a universal standard of quality and require a small number of human annotations. |
| Approach: | They propose to introduce a reference-free model-based metric with a 3-stage curriculum that can be applied to any quality standard with fewer annotations. |
| Outcome: | The proposed metric outperforms existing reference-based metrics in predicting ratings while requiring no reference simplifications at inference time. |
LLMs cannot spot math errors, even when allowed to peek into the solution (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) demonstrate impressive performance on existing reasoning benchmarks, but struggle with meta-reasoning tasks such as locating the first error step in student solutions. |
| Approach: | They propose an approach that generates an intermediate corrected student solution, aligning more closely with the original student’s solution, which helps improve performance. |
| Outcome: | The proposed approach generates an intermediate corrected student solution, aligning more closely with the original student’s solution, which helps improve performance. |
KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan (2025.acl-long)
Copied to clipboard
Mukhammed Togmanov, Nurdaulet Mukhituly, Diana Turmakhan, Jonibek Mansurov, Maiya Goloburda, Akhmed Sakip, Zhuohan Xie, Yuxia Wang, Bekassyl Syzdykov, Nurkhan Laiyk, Alham Fikri Aji, Ekaterina Kochmar, Preslav Nakov, Fajri Koto
| Challenge: | Kazakh language remains underrepresented in the field of natural language processing despite the country's population exceeding twenty million . however, there is a lack of dedicated models and benchmark evaluations specifically tailored to Kazakh languages. |
| Approach: | They propose to create a dataset specifically designed for Kazakh language with 23,000 questions sourced from authentic educational materials and manually validated by native speakers and educators. |
| Outcome: | The first MMLU-style dataset specifically designed for Kazakh language. |
SeCoDa: Sense Complexity Dataset (2020.lrec-1)
Copied to clipboard
| Challenge: | Sense Complexity Dataset (SeCoDa) provides a corpus that is annotated jointly for word senses and word tokens. |
| Approach: | They propose to use a hierarchical sense annotation scheme that draws on information available in the Cambridge Advanced Learner's Dictionary to provide more coarse-grained senses than WordNet. |
| Outcome: | The Sense Complexity Dataset (SeCoDa) provides a corpus that is annotated jointly for complexity and word senses. |
SelectLLM: Query-Aware Efficient Selection Algorithm for Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing large language models struggle with complex tasks such as factually-grounded reasoning and planning due to inherent training biases, model size constraints, and the quality or diversity of pre-training datasets. |
| Approach: | They propose a novel algorithm to select the most suitable LLMs from a large pool and use it to efficiently generalize and perform tasks. |
| Outcome: | The proposed model outperforms existing ensemble-based baselines and achieves competitive performance with similarly sized top-performing LLMs while maintaining efficiency. |
BasahaCorpus: An Expanded Linguistic Resource for Readability Assessment in Central Philippine Languages (2023.emnlp-main)
Copied to clipboard
| Challenge: | Current research on automatic readability assessment (ARA) has focused on improving the performance of models in high-resource languages such as English. |
| Approach: | They propose a hierarchical cross-lingual modeling approach that takes advantage of a language’s placement in the family tree to increase the amount of available training data. |
| Outcome: | The proposed model improves the performance of models in high-resource languages such as English and Hiligaynon, minasbate, Karay-a, and Rinconada. |
AITutor-EvalKit: Exploring the Capabilities of AI Tutors (2026.eacl-demo)
Copied to clipboard
| Challenge: | Personalized one-on-one tutoring is an effective educational approach, yet its widespread adoption is constrained by the limited availability of qualified tutors and the high costs associated with tutor training. |
| Approach: | They propose an evaluation tool that uses language technology to evaluate the pedagogical quality of AI tutors. |
| Outcome: | The proposed evaluation tool is aimed at education stakeholders as well as the *ACL community at large, as it supports learning and can also collect user feedback and annotation. |
What Makes Math Word Problems Challenging for LLMs? (2024.findings-naacl)
Copied to clipboard
| Challenge: | Experiments show that even quite powerful LLMs are still challenged by MWPs. |
| Approach: | They propose to analyze what makes math word problems (MWPs) in English challenging for large language models (LLMs). |
| Outcome: | The proposed model can handle a range of core NLP tasks, but it has emergent abilities, such as ability to solve mathematical puzzles. |
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing evaluations of large language models have been limited to subjective protocols and benchmarks. |
| Approach: | They propose a unified evaluation taxonomy with eight pedagogical dimensions based on key learning sciences principles to assess the pedagical value of LLM-powered AI tutor responses grounded in student mistakes or confusions in the mathematical domain. |
| Outcome: | The proposed taxonomy, benchmark, and human-annotated labels will streamline the evaluation process and help track the progress in AI tutors’ development. |
Word Complexity is in the Eye of the Beholder (2021.naacl-main)
Copied to clipboard
| Challenge: | Lexical complexity is a subjective notion, yet it is often neglected in lexical simplification and readability systems which use a ”one-size-fits-all” approach. |
| Approach: | They propose to use a dataset of complex words annotated by readers with different backgrounds to investigate which aspects contribute to the notion of lexical complexity. |
| Outcome: | The proposed approach can be replicated in a dataset of complex words annotated by readers with different backgrounds. |
Complex Word Identification as a Sequence Labelling Task (P19-1)
Copied to clipboard
| Challenge: | Complex Word Identification (CWI) is a crucial first step in a simplification pipeline. |
| Approach: | They propose a system that performs CWI in context without extensive feature engineering and outperforms state-of-the-art systems on this task. |
| Outcome: | The proposed system outperforms state-of-the-art systems on complex word identification. |
What Makes Cryptic Crosswords Challenging for LLMs? (2025.coling-main)
Copied to clipboard
| Challenge: | Recent research suggests that solving cryptic crosswords is challenging even for modern NLP models, including Large Language Models (LLMs). |
| Approach: | They establish benchmark results for three popular LLMs: Gemma2, LLaMA3 and ChatGPT, and investigate why these models struggle to achieve superior performance. |
| Outcome: | The proposed models perform significantly below humans on the cryptic crossword puzzle task, while human solvers achieve 99% accuracy. |
UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment (2025.emnlp-main)
Copied to clipboard
Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Joshua Reynolds, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi
| Challenge: | Language proficiency research plays a central role in education and often intersects with advances in linguistics and AI. |
| Approach: | They propose a multilingual multidimensional dataset of texts annotated according to the CEFR scale in 13 languages. |
| Outcome: | The proposed dataset supports linguistic features and pretrained models in multilingual CEFR level assessment. |
Automatic learner summary assessment for reading comprehension (N19-1)
Copied to clipboard
| Challenge: | Summarization is a well-established method of measuring reading proficiency in traditional English as a second or other language assessments. |
| Approach: | They propose three approaches to automatically assess learner summary for evaluating non-native reading comprehension using a summarization task and a long-term memory model. |
| Outcome: | The proposed models outperform traditional methods and produce quality assessments close to professional examiners. |