Reproduction and Revival of the Argument Reasoning Comprehension Task (2020.lrec-1)
Copied to clipboard
| Challenge: | Reproduction of scientific results is essential for scientific development across all disciplines. |
| Approach: | They evaluate scientific reproduction of arguments reasoning comprehension systems . they find reproducing results of previous work is a basic requirement for validating hypothesis . |
| Outcome: | The proposed systems were compared with the revised data set and scored in line with the results of the argument reasoning comprehension task. |
Similar Papers
The Argument Reasoning Comprehension Task: Identification and Reconstruction of Implicit Warrants (N18-1)
Copied to clipboard
| Challenge: | Existing methods for analyzing warrants in natural language arguments are insufficient. |
| Approach: | They propose a method for reconstructing warrants from news comments . they use a crowdsourcing process to obtain warrants for 2k authentic arguments . |
| Outcome: | The proposed method will define a substantial step towards automatic warrant reconstruction. |
Probing Neural Network Comprehension of Natural Language Arguments (P19-1)
Copied to clipboard
| Challenge: | Argument Reasoning Comprehension Task (ARCT) focuses on inferences, not just discovering warrants. |
| Approach: | They propose to build an adversarial dataset on which all models achieve random accuracy. |
| Outcome: | The proposed dataset provides a more robust assessment of argument comprehension and should be adopted as the standard in future work. |
Demonstration Retrieval-Augmented Generative Event Argument Extraction (2024.lrec-main)
Copied to clipboard
| Challenge: | Experimental results show that our method outperforms all strong baselines and can be generalized to various datasets. |
| Approach: | They propose a generative EAE that uses event knowledge-injected generator and demonstration retriever to generate event arguments from training data. |
| Outcome: | The proposed method outperforms baselines and can be generalized to various datasets. |
Enhancing Study-Level Inference from Clinical Trial Papers via Reinforcement Learning-Based Numeric Reasoning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Prior work has framed this task as a textual inference task by retrieving relevant content fragments and inferring conclusions from them. |
| Approach: | They propose to extract structured numerical evidence and apply domain knowledge informed logic to derive outcome-specific conclusions. |
| Outcome: | The proposed approach outperforms general-purpose LLMs of over 400B parameters and achieves a 21% improvement in F1 score over retrieval-based systems. |
Automatic Argument Quality Assessment - New Datasets and Methods (D19-1)
Copied to clipboard
Assaf Toledo, Shai Gretz, Edo Cohen-Karlik, Roni Friedman, Elad Venezian, Dan Lahav, Michal Jacovi, Ranit Aharonov, Noam Slonim
| Challenge: | 6.3k arguments were collected from contributors of various levels, and are released as part of this work. |
| Approach: | They propose to use a language model to annotate arguments for argument ranking and argument-pair classification. |
| Outcome: | The proposed methods outperform state-of-the-art methods in the argument ranking task and argument-pair classification task. |
Bringing replication and reproduction together with generalisability in NLP: Three reproduction studies for Target Dependent Sentiment Analysis (C18-1)
Copied to clipboard
| Challenge: | a lack of reproducibility and generalisability is a major threat to scientific development in Natural Language Processing. |
| Approach: | They propose to use a model zoo to document and release language models and published code . they recommend that future replication experiments should consider a variety of datasets . |
| Outcome: | The proposed methods are compared on six English datasets and are based on the results. |
Reproducing Neural Ensemble Classifier for Semantic Relation Extraction inScientific Papers (2020.lrec-1)
Copied to clipboard
| Challenge: | Replicability and reproducibility are core ideas of modern scientific methods. |
| Approach: | They describe challenges encountered in reproducing the results of a top performing system in computational linguistics. |
| Outcome: | The proposed system was able to reproduce the results of a task 7 in the domain of natural language processing and computational linguistics. |
Predicting Desirable Revisions of Evidence and Reasoning in Argumentative Writing (2023.findings-eacl)
Copied to clipboard
| Challenge: | Using the essay context of the revision and feedback from students prior to the revision, we identify desirable and undesirable revisions. |
| Approach: | They propose to use the essay context of the revision and the feedback students received before the revision to improve classifier performance. |
| Outcome: | The proposed models improve over baseline models, while models utilizing context improve over the baseline models. |
ConQRet: A New Benchmark for Fine-Grained Automatic Evaluation of Retrieval Augmented Computational Argumentation (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing methods for evaluating RAArg are costly and lack long, complex arguments and real-world evidence. |
| Approach: | They propose to use multiple fine-grained LLM judges to evaluate RAArg using a new benchmark that features long and complex human-authored arguments on debated topics. |
| Outcome: | The proposed methods provide better and more interpretable assessments than traditional single-score metrics and even previously reported human crowdsourcing. |
ASU at TextGraphs 2019 Shared Task: Explanation ReGeneration using Language Models and Iterative Re-Ranking (D19-53)
Copied to clipboard
| Challenge: | Explanation Regeneration task is an intermediate step towards general multi-hop inference on large graphs. |
| Approach: | They propose a system that performs multi-hop inference and ranks a set of explanatory facts for a given elementary science question and correct answer pair. |
| Outcome: | The proposed system secured 2nd rank in the text graphs 2019 shared task with a mean average precision (MAP) of 41.3% on the test set. |