Papers with RQ2

8 papers
Comparing Template-based and Template-free Language Model Probing (2024.eacl-long)

Copied to clipboard

Challenge: Template-based and template-based approaches rank models differently except for the top domain-specific models.
Approach: They evaluate 16 different cloze-task language model probing approaches on 10 probing English datasets to answer questions about model rankings and absolute scores.
Outcome: The results show that the template-based and template-free approaches rank models differently except for the top domain-specific models.
Analyzing Interpretability of Summarization Model with Eye-gaze Information (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have provided saliency scores for neural summarization models . eye-gaze information is often used as a proxy for human attention in reading tasks .
Approach: They propose to compare model saliency to human eye-gaze data to determine whether it conforms to human gaze during summarization.
Outcome: The proposed framework compares the model behavior to human summarization performance.
Evaluating Self-Generated Documents for Enhancing Retrieval-Augmented Generation with Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Recent research focuses on optimizing the use of Self-Docs with their inherent properties remaining underexplored.
Approach: They develop a taxonomy to compare the effectiveness of different types of Self-Docs and explore strategies for combining them with external sources.
Outcome: The proposed model can supplement retrieved content and provide a powerful way to improve knowledge-intensive question answering tasks.
Not Enough Data to Pre-train Your Language Model? MT to the Rescue! (2023.findings-acl)

Copied to clipboard

Challenge: In recent years, transformer-based language models (LMs) have become the default approach for many NLP tasks.
Approach: They compare the performance of transformer-based language models with machine-translated corpora.
Outcome: The proposed model can be improved with real data, but further research is needed.
ECON: On the Detection and Resolution of Evidence Conflicts (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that AI generated content is more likely to dominate search results, making it difficult to detect when compared to human-produced content.
Approach: They propose a method for generating diverse, validated evidence conflicts to simulate real-world misinformation scenarios.
Outcome: The proposed method enables the detection of conflicting information in real-world scenarios and shows that weaker models struggle with similar answer conflicts while stronger models show robust performance.
On the Robustness of Editing Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies have exhibited impressive success and significant potential.
Approach: They propose to modify the knowledge memory with minimum computational cost while preserving the performance on the retained knowledge.
Outcome: The proposed methods avoid retraining to update the model parameters and have demonstrated promising performance and efficiency.
On Generalization across Measurement Systems: LLMs Entail More Test-Time Compute for Underrepresented Cultures (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) should be able to provide accurate information irrespective of the measurement system at hand .
Approach: They use newly compiled datasets to test if this is true for seven open-source LLMs.
Outcome: The proposed model can provide accurate information regardless of the measurement system at hand.
PopAlign: Diversifying Contrasting Patterns for a More Comprehensive Alignment (2025.acl-long)

Copied to clipboard

Challenge: Typical approaches to training large language models rely on limited contrasting patterns . contrasting data is limited and models are susceptible to harmful response tendencies .
Approach: They propose a framework that integrates contrasting patterns across the prompt, model, and pipeline levels.
Outcome: The proposed framework outperforms existing methods in the comparison of RQ1 and RQ2 . the proposed framework significantly outperformed existing methods, leading to more comprehensive alignment.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations