Challenge: Large language models (LLMs) perform well on table tasks, but they still make data referencing errors (DREs) prior studies have only offered limited, small-scale analyses.
Approach: They propose inference-time strategies and lightweight critics to mitigate data referencing errors.
Outcome: The proposed model achieves an average F1 score of 78.2% in detecting both in-distribution and out-of-difference DREs and assists inference for larger models.

Similar Papers

Lost in the Source Language: How Large Language Models Evaluate the Quality of Machine Translation (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies have shown that Large Language Models (LLMs) can be used as translation evaluators.
Approach: They propose to use both coarse-grained and fine-grounded prompts to discern the utility of source versus reference data in machine translation evaluation tasks.
Outcome: The proposed model can be used to evaluate translations in multiple languages.
Beyond Performance: Quantifying and Mitigating Label Bias in LLMs (2024.naacl-long)

Copied to clipboard

Challenge: Large language models exhibit undesirable preference toward predicting certain answers over others, despite their adaptability to diverse tasks.
Approach: They propose a label bias calibration method that outperforms recent calibration approaches for improving performance and mitigating label bias.
Outcome: The proposed method outperforms calibration approaches for improving performance and mitigating label bias.
Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) offer new opportunities to enhance the annotation process, particularly for detecting label errors in existing datasets.
Approach: They propose to use an ensemble of large language models to flag mislabeled examples by using an LLM-as-a-judge approach to detect label errors in existing datasets.
Outcome: The proposed method improves label accuracy and consistency in large language models.
TableLLM: Enabling Tabular Data Manipulation by LLMs in Real Office Usage Scenarios (2025.findings-acl)

Copied to clipboard

Challenge: TableLLM is a robust large language model capable of handling tabular data manipulation tasks.
Approach: They propose a distant supervision method for training which includes a reasoning process extension strategy and a cross-way validation strategy.
Outcome: The proposed model has 8 billion parameters and is capable of handling tabular data tasks.
Rethinking Tabular Data Understanding with Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are capable of various tasks, yet their capability in interpreting and reasoning over tabular data remains an underexplored area.
Approach: They propose a method for table structure normalization to improve model performance . they propose aggregation of multiple reasoning pathways to improve performance based on textual and symbolic reasoning.
Outcome: The proposed method improves performance on symbolic reasoning tasks with textual reasoning slightly outperforming symbolic reasoning on tables.
Automatic Evaluation of Attribution by Large Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Generative large language models (LLMs) incorporate external references to generate and support claims. however, evaluating the attribution remains an open problem.
Approach: They investigate automatic evaluation of attribution given by large language models . they define different types of attributed errors and then explore two approaches .
Outcome: The proposed methods highlight promising signals and challenges.
What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects (2026.findings-eacl)

Copied to clipboard

Challenge: a series of paradigm shifts have come with distinct characteristics and challenges associated with table modeling.
Approach: They propose to replicate four table LLMs by instruction-tuning three foundation models on four existing datasets.
Outcome: The results show that base model choice plays a more dominant role than training data itself.
What Did I Do Wrong? Quantifying LLMs’ Sensitivity and Consistency to Prompt Engineering (2025.naacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have significantly improved productivity in a number of routine tasks.
Approach: They propose two metrics for classification tasks, namely *sensitivity* and *consistency*, which are complementary to task performance.
Outcome: The proposed metrics are complementary to task performance and can be used to guide prompt engineering and obtain LLMs that balance robustness and performance.
Aligning What LLMs Do and Say: Towards Self-Consistent Explanations (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are often prompted to produce natural language explanations, but the features driving the answer are often different from those emphasized in their explanations.
Approach: They propose a large-scale benchmark linking model decisions with diverse explanations and attribution vectors across datasets, methods, and model families to address this gap.
Outcome: The proposed model generates an answer where the word NLP in the prompt has high feature importance.
Confronting LLMs with Traditional ML: Rethinking the Fairness of Large Language Models in Tabular Classifications (2024.naacl-long)

Copied to clipboard

Challenge: Recent studies suggest using large language models to make tabular classifications . however, LLMs have been shown to exhibit harmful social biases based on stereotypes and inequalities present in society.
Approach: They propose to use large language models to make tabular classifications . they show that LLMs inherit biases from their training data .
Outcome: The proposed models exhibit harmful biases that reflect stereotypes and inequalities in society.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations