Papers by Amey Hengle
Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval over haystacks (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing multilingual long-context benchmarks are myopic and inherently limited, as successful recall alone does not indicate a model’s capacity to reason over extended contexts. |
| Approach: | They propose a new synthetic benchmark for multilingual long-context reasoning that includes bAbI-style tasks that test multi-hop inference, aggregation, and epistemic reasoning. |
| Outcome: | The proposed benchmarks are based on a multilingual long-context model and span seven languages. |
Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Recent large language models demonstrate remarkable abilities in responding to queries in diverse languages, but their ability to handle long multilingual contexts is unexplored. |
| Approach: | They propose a multilingual Needle-in-a-Haystack (MLNeedle) test to assess a model's ability to retrieve relevant information from a collection of multilingual distractor texts. |
| Outcome: | The proposed model performance is the lowest when the needle is in a language outside the English language family and (ii) located in the middle of the input context. |
CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs (2025.naacl-long)
Copied to clipboard
| Challenge: | Current evaluation methods do not capture complex attributes of counterspeech quality, such as contextual relevance, aggressiveness, or argumentative coherence. |
| Approach: | They propose to use a dataset and framework to evaluate counterspeech quality across four dimensions: contextual relevance, aggressiveness, argument-coherence, and suitability. |
| Outcome: | The proposed method outperforms ROUGE, METEOR, and BertScore in correlating with human judgement, indicating a significant improvement in automated counterspeech evaluation. |
Counting the Bugs in ChatGPT’s Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language Model (2023.emnlp-main)
Copied to clipboard
Leonie Weissweiler, Valentin Hofmann, Anjali Kantharuban, Anna Cai, Ritam Dutt, Amey Hengle, Anubha Kabra, Atharva Kulkarni, Abhishek Vijayakumar, Haofei Yu, Hinrich Schuetze, Kemal Oflazer, David Mortensen
| Challenge: | Existing studies on large language models (LLMs) ignore the remarkable ability of humans to generalize and focus only on English. |
| Approach: | They conduct the first rigorous analysis of the morphological capabilities of ChatGPT in four typologically varied languages. |
| Outcome: | The proposed model massively underperforms purpose-built systems, particularly in English. |
Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing systems that target hate speech with intent-conditioned counterspeech generate better results with longer contexts. |
| Approach: | They propose a framework that enables counterspeech generation by modeling the pragmatic implications underlying social biases in hateful statements. |
| Outcome: | The proposed framework outperforms existing benchmarks in intent-conditioned counterspeech generation. |
Still Not Quite There! Evaluating Large Language Models for Comorbid Mental Health Diagnosis (2024.emnlp-main)
Copied to clipboard
Amey Hengle, Atharva Kulkarni, Shantanu Patankar, Madhumitha Chandrasekaran, Sneha D’silva, Jemima Jacob, Rashmi Gupta
| Challenge: | ANGST is a benchmark for depression-anxiety comorbidity classification from social media posts. |
| Approach: | They propose a social media-based benchmark for depression-anxiety comorbidity classification . ANGST enables multi-label classification, allowing each post to be simultaneously identified as indicating depression and/or anxiety. |
| Outcome: | The proposed dataset enables multi-label classification of depression and anxiety . it outperforms existing models but none achieves an F1 score exceeding 72% . |