Papers with RQ2
Comparing Template-based and Template-free Language Model Probing (2024.eacl-long)
Copied to clipboard
| Challenge: | Template-based and template-based approaches rank models differently except for the top domain-specific models. |
| Approach: | They evaluate 16 different cloze-task language model probing approaches on 10 probing English datasets to answer questions about model rankings and absolute scores. |
| Outcome: | The results show that the template-based and template-free approaches rank models differently except for the top domain-specific models. |
Analyzing Interpretability of Summarization Model with Eye-gaze Information (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies have provided saliency scores for neural summarization models . eye-gaze information is often used as a proxy for human attention in reading tasks . |
| Approach: | They propose to compare model saliency to human eye-gaze data to determine whether it conforms to human gaze during summarization. |
| Outcome: | The proposed framework compares the model behavior to human summarization performance. |
Evaluating Self-Generated Documents for Enhancing Retrieval-Augmented Generation with Large Language Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Recent research focuses on optimizing the use of Self-Docs with their inherent properties remaining underexplored. |
| Approach: | They develop a taxonomy to compare the effectiveness of different types of Self-Docs and explore strategies for combining them with external sources. |
| Outcome: | The proposed model can supplement retrieved content and provide a powerful way to improve knowledge-intensive question answering tasks. |
Not Enough Data to Pre-train Your Language Model? MT to the Rescue! (2023.findings-acl)
Copied to clipboard
| Challenge: | In recent years, transformer-based language models (LMs) have become the default approach for many NLP tasks. |
| Approach: | They compare the performance of transformer-based language models with machine-translated corpora. |
| Outcome: | The proposed model can be improved with real data, but further research is needed. |
ECON: On the Detection and Resolution of Evidence Conflicts (2024.emnlp-main)
Copied to clipboard
Cheng Jiayang, Chunkit Chan, Qianqian Zhuang, Lin Qiu, Tianhang Zhang, Tengxiao Liu, Yangqiu Song, Yue Zhang, Pengfei Liu, Zheng Zhang
| Challenge: | Recent studies have shown that AI generated content is more likely to dominate search results, making it difficult to detect when compared to human-produced content. |
| Approach: | They propose a method for generating diverse, validated evidence conflicts to simulate real-world misinformation scenarios. |
| Outcome: | The proposed method enables the detection of conflicting information in real-world scenarios and shows that weaker models struggle with similar answer conflicts while stronger models show robust performance. |
On the Robustness of Editing Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have exhibited impressive success and significant potential. |
| Approach: | They propose to modify the knowledge memory with minimum computational cost while preserving the performance on the retained knowledge. |
| Outcome: | The proposed methods avoid retraining to update the model parameters and have demonstrated promising performance and efficiency. |
On Generalization across Measurement Systems: LLMs Entail More Test-Time Compute for Underrepresented Cultures (2025.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) should be able to provide accurate information irrespective of the measurement system at hand . |
| Approach: | They use newly compiled datasets to test if this is true for seven open-source LLMs. |
| Outcome: | The proposed model can provide accurate information regardless of the measurement system at hand. |
PopAlign: Diversifying Contrasting Patterns for a More Comprehensive Alignment (2025.acl-long)
Copied to clipboard
Zekun Moore Wang, Shenzhi Wang, King Zhu, Jiaheng Liu, Ke Xu, Jie Fu, Wangchunshu Zhou, Wenhao Huang
| Challenge: | Typical approaches to training large language models rely on limited contrasting patterns . contrasting data is limited and models are susceptible to harmful response tendencies . |
| Approach: | They propose a framework that integrates contrasting patterns across the prompt, model, and pipeline levels. |
| Outcome: | The proposed framework outperforms existing methods in the comparison of RQ1 and RQ2 . the proposed framework significantly outperformed existing methods, leading to more comprehensive alignment. |