Un-considering Contextual Information: Assessing LLMs’ Understanding of Indexical Elements (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) excel in coreference resolution tasks, but previous studies only assessed performance with nouns and third person pronouns. |
| Approach: | They evaluate LLMs' performance on coreference resolution with indexicals like I, you, here and tomorrow which come with unique challenges due to their linguistic properties. |
| Outcome: | The proposed models perform well with some indexicals while struggling with others. |
Similar Papers
Assessing the Capabilities of Large Language Models in Coreference: An Evaluation (2024.lrec-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are a new approach to coreference resolution, but their performance is not yet fully understood. |
| Approach: | They propose that future efforts should improve scope, data, and evaluation methods of traditional coreference research to adapt to the development of LLMs. |
| Outcome: | The proposed methods improve scope, data, and evaluation methods of traditional coreference research to adapt to the development of LLMs. |
Correct-Detect: Balancing Performance and Ambiguity Through the Lens of Coreference Resolution in LLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are intended to reflect human linguistic competencies . but when context is absent or insufficient, ambiguity resolution becomes more tenuous . |
| Approach: | They propose a CORRECT-DETECT trade-off between large language models and ambiguity detection . they show that large language model models can achieve good performance with minimal prompting . |
| Outcome: | The proposed models can achieve good performance with minimal prompting in coreference disambiguation and detection of ambiguity in corefertility tasks, but they cannot do both at the same time. |
The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead? (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models have shown capabilities close to human performance in various analytical tasks. |
| Approach: | They investigate the efficiency and accuracy of Large Language Models in specialized tasks . they integrate LLMs with expert annotators to observe the impact of LLM suggestions . |
| Outcome: | The proposed model improves task completion speed but introduces anchoring bias . the proposed model is not suitable for open-ended analysis, but is capable of handling specialized tasks. |
Efficient Solutions For An Intriguing Failure of LLMs: Long Context Window Does Not Mean LLMs Can Analyze Long Sequences Flawlessly (2025.coling-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable capabilities in comprehending and analyzing lengthy sequential inputs. |
| Approach: | They propose to implement ad-hoc solutions that enhance LLMs’ performance on long input sequences by up to 50% while reducing API cost and latency by up . to address this limitation, they propose to use three datasets and two tasks to analyze news categorization and sentence analysis to evaluate their models. |
| Outcome: | The proposed solutions significantly improve LLMs’ performance on long input sequences by up to 50% while reducing API cost and latency by up . to 93% and 50%, respectively. |
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs (2025.naacl-short)
Copied to clipboard
| Challenge: | Recent evaluations of LLMs on coreference resolution have revealed that traditional output formats and evaluation metrics do not fully capture the models’ referential understanding. |
| Approach: | They propose a benchmark for mention resolution presented in a multiple-choice question format and a curated mixture of different mention types and corresponding entities. |
| Outcome: | The proposed model achieves 81.9% accuracy while the open model achieve 80%. |
CorefInst: Leveraging LLMs for Multilingual Coreference Resolution (2026.tacl-1)
Copied to clipboard
| Challenge: | Existing methods for CR are encoder-only, decoder-based and asynchronous models. |
| Approach: | They propose a multilingual CR methodology which leverages decoder-only LLMs to handle overt and zero mentions. |
| Outcome: | The proposed model outperforms the leading multilingual CR model by 2 percentage points across all languages in the CorefUD v1.2 dataset. |
CUTE: Measuring LLMs’ Understanding of Their Tokens (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) perform well on a wide variety of tasks, authors say . they lack direct access to characters, which can be difficult to generalize to new languages . |
| Approach: | They propose a benchmark to test the orthographic knowledge of Large Language Models . they find that most LLMs seem to know the spelling of their tokens - yet fail to manipulate text . |
| Outcome: | The proposed benchmark tests the orthographic knowledge of large language models . it finds that most LLMs seem to know the spelling of their tokens, but fail to manipulate text . |
Reassessing Semantic Knowledge Encoded in Large Language Models through the Word-in-Context Task (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have propelled significant progress, extending their application across various domains including dialogue systems, text generation, translation systems, and beyond. |
| Approach: | They propose to use the Word-in-Context (WiC) task to reassess the semantic knowledge encoded in large language models (LLMs) they prompt LLMs to generate natural language descriptions that contrast the meanings of the target word in two contextual sentences given in the WiC dataset. |
| Outcome: | The proposed model significantly improves the classification accuracy of the two models. |
From Fallback to Frontline: When Can LLMs be Superior Annotators of Human Perspectives? (2026.findings-acl)
Copied to clipboard
| Challenge: | large language models are often used as annotators at scale, but are not faithful estimators of human perspectives. |
| Approach: | They characterize the conditions under which large language models outperform human annotators . they find they are statistically superior frontline estimators based on low variance . |
| Outcome: | The proposed model outperforms human annotators when predicting subgroup opinions on subjective tasks. |
Large Language Models Are Partially Primed in Pronoun Interpretation (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing studies suggest large language models acquire rich linguistic representations, but little is known about whether they adapt to linguistic biases in a human-like way. |
| Approach: | They examine whether large language models display human-like referential biases using stimuli and procedures from real psycholinguistic experiments. |
| Outcome: | The proposed models display human-like referential biases when exposed to referential patterns in the local context. |