Challenge: Using a multilingual model, we examine the ability of large language models to perform reasoning tasks.
Approach: They propose to use a multilingual model to analyze commonsense reasoning in large language models for Italian and to provide a semi-automated system to complete the annotation.
Outcome: The proposed model performs at high-level classification tasks but its easoning is inconsistent and unverifiable, since it does not capture intermediate evidence.

Similar Papers

Tiered Reasoning for Intuitive Physics: Toward Verifiable Commonsense Language Understanding (2021.findings-emnlp)

Copied to clipboard

Challenge: Large-scale, pre-trained language models (LMs) have achieved human-level performance on a breadth of language understanding tasks.
Approach: They propose a commonsense reasoning dataset with dense annotations that allows multi-tiered evaluation of machines’ reasoning process.
Outcome: The proposed model can achieve high end performance but struggle to support predictions with valid supporting evidence.
ITALIC: An Italian Culture-Aware Natural Language Benchmark (2025.naacl-long)

Copied to clipboard

Challenge: ITALIC is a large-scale benchmark dataset of 10,000 multiple-choice questions designed to evaluate the natural language understanding of the Italian language and culture.
Approach: They propose to use a large-scale benchmark dataset to evaluate the natural language understanding of the Italian language and culture.
Outcome: The ITALIC dataset spans 12 domains and uses 17 state-of-the-art LLMs to assess the natural language understanding of the italian language and culture.
On the Consistency of Commonsense in Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluations of commonsense for large language models focus on downstream knowledge tasks, failing to probe whether LLMs truly understand and utilize knowledge or merely memorize it.
Approach: They propose to automatically construct a large benchmark named CoCo which measures LLMs’ knowledge memorization, comprehension, and application and examines the consistency between these tasks.
Outcome: The proposed benchmark systematically assesses LLMs’ knowledge memorization, comprehension, and application and examines the consistency between these tasks.
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are used globally across many languages, but their English-centric pretraining raises concerns about cross-lingual disparities for cultural awareness .
Approach: They introduce an automatic multilingual framework for evaluating cultural awareness in large language models across languages, regions, and topics.
Outcome: The framework evaluates open-ended text generation, capturing how models express culturally grounded knowledge in natural language.
Commonsense Reasoning for Natural Language Processing (2020.acl-tutorials)

Copied to clipboard

Challenge: In this tutorial, we will outline the various types of commonsense knowledge and discuss techniques to gather and represent commonsence knowledge.
Approach: This tutorial will provide researchers with the critical foundations and recent advances in commonsense representation and reasoning.
Outcome: This tutorial will outline the various types of commonsense and discuss techniques to gather and represent commonsence knowledge while highlighting the challenges specific to this type of knowledge (e.g., reporting bias).
Learning the Effects of Physical Actions in a Multi-modal Environment (2023.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are trained on large corpora of disembodied texts.
Approach: They propose a multi-modal task of predicting the outcomes of actions solely from realistic sensory inputs (images and text). They extend an LLM to model latent representations of objects to better predict action outcomes in an environment.
Outcome: The proposed model can capture commonsense when augmented with visual information and generalize and learn commonsensical reasoning better.
DanteLLM: Let’s Push Italian LLM Research Forward! (2024.lrec-main)

Copied to clipboard

Challenge: Existing models for large language processing in the English language are limited in resources and evaluation tools for non-English languages.
Approach: They propose a benchmark and an open LLM Leaderboard to evaluate LLMs’ performance in Italian and propose 'DanteLLM' it is the most performant LLM in the world, with improvements of up to 6 points .
Outcome: The proposed model outperforms existing models in Italian and offers a blueprint for the development and evaluation of LLMs in other languages.
A Method for Building a Commonsense Inference Dataset based on Basic Events (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to acquire commonsense are limited by the general-purpose language models.
Approach: They propose a method for building a commonsense inference dataset using crowdsourcing and automatic extraction from a corpus.
Outcome: The proposed method can solve 104k commonsense inference problems in a Japanese corpus with high accuracy, but low bias.
ReportLogic: Evaluating Logical Quality in Deep Research Reports (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation frameworks that evaluate large language models for Deep Research largely ignore this requirement.
Approach: They propose a benchmark that quantifies report-level logical quality through a reader-centric lens of auditability.
Outcome: The proposed model quantifies logical quality through a reader-centric lens of auditability.
Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and Sundanese (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly being used to generate synthetic data for training and evaluating models.
Approach: They investigate the effectiveness of using Large Language Models to generate culturally relevant commonsense QA datasets for Indonesian and Sundanese languages using both LLMs and human annotators.
Outcome: The proposed model generates 4.5K questions per language, compared with 4.5k for Indonesian and 4.5km for Sundanese.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations