Challenge: Existing benchmarks for Large Language Models (LLMs) follow the data distribution of pre-training data.
Approach: They propose a benchmark ConvRe focusing on converse relations which contains 17 relations and 1240 triples extracted from popular knowledge graph completion datasets.
Outcome: The proposed benchmark focuses on converse relations, which contains 17 relations and 1240 triples extracted from popular knowledge graph completion datasets.

Similar Papers

Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET Benchmark (2023.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks of social language are lacking for large language models.
Approach: They propose a new benchmark that measures how well large language models understand social language by grouping 58 tasks into five categories: humor & sarcasm, offensiveness, sentiment & emotion, and trustworthiness.
Outcome: The proposed model performs well at 58 tasks that are divided into five categories: humor & sarcasm, offensiveness, sentiment & emotion, and trustworthiness.
Large Language Models: The Need for Nuance in Current Debates and a Pragmatic Perspective on Understanding (2023.emnlp-main)

Copied to clipboard

Challenge: Current Large Language Models (LLMs) are unparalleled in their ability to generate grammatically correct, fluent text.
Approach: They argue that LLMs only parrot statistical patterns in training data and that language learning in LLM cannot inform human language learning.
Outcome: The proposed model can generate grammatically correct, fluent text without requiring human intervention.
Exploring Graph Learning Tasks with Pure LLMs: A Comprehensive Benchmark and Investigation (2026.findings-acl)

Copied to clipboard

Challenge: Recent studies focus on performance benchmarks without fully comparing LLMs to graph learning models.
Approach: They evaluate off-the-shelf and instruction-tuned graph learning models across a variety of scenarios.
Outcome: The proposed models outperform traditional graph learning models in few-shot settings, the authors show . their models out perform models with instruction tuning, and they show excellent generalization and robustness.
Can Large Language Models Understand DL-Lite Ontologies? An Empirical Study (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have shown remarkable proficiency in understanding textual data and revolutionizing the field of natural language processing.
Approach: They empirically analyze LLMs' capability of understanding Description Logic (DL) ontologies covering 6 representative tasks from syntactic and semantic aspects.
Outcome: The proposed model can understand formal syntax and model-theoretic semantics of concepts and roles, but struggle with understanding TBox NI transitivity and handling ontologies with large ABoxes.
DIVKNOWQA: Assessing the Reasoning Ability of LLMs via Open-Domain Question Answering over Knowledge Base and Text (2024.findings-naacl)

Copied to clipboard

Challenge: Retrievalaugmented LLMs have been used to ground LLM in external knowledge . a gap exists in the current landscape regarding the effectiveness of grounding LLM on heterogeneous knowledge sources.
Approach: They propose a model that uses symbolic language to generate symbolic queries . they use a dataset that is generated using predefined reasoning chains and human annotation .
Outcome: The proposed model outperforms previous approaches by a significant margin in QA tasks over text.
LLM4RE: A Data-centric Feasibility Study for Relation Extraction (2025.coling-main)

Copied to clipboard

Challenge: Relation Extraction (RE) is a critical step in information extraction due to its wide-scale applicability for downstream applications such as Knowledge Base creation and Question Answering (QA).
Approach: They propose to conduct the first feasibility analysis to explore the viability of Large Language Models for RE by investigating their robustness to various RE scenarios stemming from data-specific characteristics.
Outcome: The proposed models are robust to various RE scenarios stemming from data-specific characteristics, but their performance is not yet fully understood.
LLMs meet Bloom’s Taxonomy: A Cognitive View on Large Language Model Evaluations (2025.coling-main)

Copied to clipboard

Challenge: Existing evaluation approaches for Large Language Models lack a structured approach that reflects the underlying cognitive abilities required for solving the tasks.
Approach: They propose a hierarchical approach to evaluation of Large Language Models that leverages Bloom’s Taxonomy to identify how well they cover the levels of Bloom’ s taxonomies.
Outcome: The proposed evaluation frameworks cover the Bloom’s Taxonomy, a hierarchical framework for categorizing cognitive skills, on the most widely used benchmarks.
Metaphor and Large Language Models: When Surface Features Matter More than Deep Understanding (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies on metaphor processing have focused on single datasets and specific task settings, often using artificially constructed data through lexical replacement.
Approach: They propose to evaluate the capabilities of Large Language Models (LLMs) in metaphor interpretation across multiple datasets, tasks, and prompt configurations.
Outcome: The proposed frameworks are more realistic and efficient than current models and are more efficient than existing models.
CUTE: Measuring LLMs’ Understanding of Their Tokens (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) perform well on a wide variety of tasks, authors say . they lack direct access to characters, which can be difficult to generalize to new languages .
Approach: They propose a benchmark to test the orthographic knowledge of Large Language Models . they find that most LLMs seem to know the spelling of their tokens - yet fail to manipulate text .
Outcome: The proposed benchmark tests the orthographic knowledge of large language models . it finds that most LLMs seem to know the spelling of their tokens, but fail to manipulate text .
Lexical Semantics with Large Language Models: A Case Study of English “break” (2023.findings-eacl)

Copied to clipboard

Challenge: Large neural language models (LLMs) can be powerful tools for research in lexical semantics.
Approach: They argue that large neural language models can be powerful tools for research in lexical semantics by capturing known sense distinctions and identifying informative new sense combinations.
Outcome: The proposed models capture many of the sense distinctions found in the English verb break and can be used to identify informative new sense combinations for further analysis.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations