Challenge: Argumentation in natural language processing (NLP) is becoming an indispensable tool in many application domains such as public policy, law, medicine, and education.
Approach: They propose a reconstructed dataset of argument and counter-argument pairs . they propose integrating dynamic external knowledge from the web to improve counter-arguments .
Outcome: The proposed method shows stronger correlation with human judgments compared to reference-based metrics.

Similar Papers

Exploring the Potential of Large Language Models in Computational Argumentation (2024.acl-long)

Copied to clipboard

Challenge: Argumentation is an essential tool in various domains, including law, public policy, and artificial intelligence.
Approach: They propose to evaluate LLMs on various computational argumentation tasks . they organize existing tasks into six main categories and standardize the format of 14 datasets .
Outcome: The proposed model performs well on argument mining and argument generation tasks.
Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models have raised concerns about reliability and trustworthiness of the models.
Approach: They analyze 134 papers and introduce a taxonomy of evidence-based text generation with LLMs.
Outcome: The proposed methods highlight open challenges and outline promising directions for future work.
Can Language Models Recognize Convincing Arguments? (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have found that large language models can generate persuasive content without engaging in human experimentation.
Approach: They extend a dataset with debates, votes, and user traits to measure LLMs' ability to distinguish between strong and weak arguments, predict stances based on beliefs and demographic characteristics, and determine appeal of argument to individual based upon their traits.
Outcome: The proposed tasks outperform human predictions in detecting convincing arguments in debates, votes, and user traits.
Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey (2025.findings-emnlp)

Copied to clipboard

Challenge: specialized LLMs are often limited in domain-specific applications that require specialized knowledge.
Approach: They provide a comprehensive overview of four key methods to enhance large language models by integrating domain-specific knowledge.
Outcome: The proposed methods are categorized into four key approaches: dynamic knowledge injection, static knowledge embedding, modular adapters, and prompt optimization.
LLM DEBATE OPPONENT : Counter-argument Generation focusing on Implicit and Critical Premises (2025.naacl-srw)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) show promise in automating counter-argument generation.
Approach: They compare multi-step and one-step generation methods for counter-arguments across 100 debate topics.
Outcome: The proposed model outperforms multi-step and one-step pipelines for counter-arguments across 100 debate topics.
Natural Language Reasoning in Large Language Models: Analysis and Evaluation (2025.findings-acl)

Copied to clipboard

Challenge: Argumentative reasoning presents unique challenges due to its reliance on context, implicit assumptions, and value judgments.
Approach: They propose a large-scale evaluation of LLMs' unconstrained natural language reasoning capabilities . they formalise a new strategy designed to evaluate argumentative reasoning in LLM .
Outcome: The proposed model performs better on a range of reasoning tasks than other models.
Argument Summarization and its Evaluation in the Era of Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have revolutionized various Natural Language Generation tasks, including Argument Summarization (ArgSum).
Approach: They propose a prompt-based evaluation scheme and validate it through a human benchmark dataset.
Outcome: The proposed evaluation scheme outperforms existing methods and is validated by a human benchmark dataset.
DIVKNOWQA: Assessing the Reasoning Ability of LLMs via Open-Domain Question Answering over Knowledge Base and Text (2024.findings-naacl)

Copied to clipboard

Challenge: Retrievalaugmented LLMs have been used to ground LLM in external knowledge . a gap exists in the current landscape regarding the effectiveness of grounding LLM on heterogeneous knowledge sources.
Approach: They propose a model that uses symbolic language to generate symbolic queries . they use a dataset that is generated using predefined reasoning chains and human annotation .
Outcome: The proposed model outperforms previous approaches by a significant margin in QA tasks over text.
Can LLMs Really Judge? A Progressive Argumentation-Mining Framework for Distinguishing Understanding from Aggregation (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluations of large language models rely on dataset-based generation accuracy . however, generative correctness does not guarantee discriminative capability to verify solutions .
Approach: They propose a diagnostic framework that explicitly controls context and isolates discriminative behaviors.
Outcome: The proposed framework explicitly controls context and isolates discriminative behaviors.
1+1>2: Can Large Language Models Serve as Cross-Lingual Knowledge Aggregators? (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been recognized for their impressive capabilities in natural language processing (NLP).
Approach: They propose a method to enhance the multilingual performance of Large Language Models by aggregating knowledge from diverse languages.
Outcome: The proposed method reduces the performance disparity across languages and offers valuable insights for further exploration.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations