Challenge: Existing approaches to explain the correct answer in multiple-choice QA are low in F-scores and lack of performance.
Approach: They introduce a lightweight focus feature in a transformer-based NLP model and examine performance improvements.
Outcome: The proposed model achieves the highest scores, second only to a computationally intensive system.

Similar Papers

Reading Comprehension as Natural Language Inference:A Semantic Analysis (2020.starsem-1)

Copied to clipboard

Challenge: In recent past, Natural language Inference (NLI) has gained significant attention, but its true impact has not been well studied.
Approach: They propose to transform a large RACE dataset into an NLI model and compare it to a state-of-the-art model.
Outcome: The proposed model outperforms the previous model on a question-answer concatenation form and a coherent entailment form.
Paraphrase Identification via Textual Inference (2024.starsem-1)

Copied to clipboard

Challenge: Paraphrase identification (PI) and natural language inference (NLI) are important tasks in natural language processing.
Approach: They propose a method for paraphrase identification and natural language inference using an NLI system to solve these tasks.
Outcome: The proposed method outperforms dedicated PI models on PI datasets and provides insights into limitations of current benchmarks.
Dyna-bAbI: unlocking bAbI’s potential with dynamic synthetic benchmarking (2022.starsem-1)

Copied to clipboard

Challenge: Controlled synthetic tasks are an important resource for diagnosing model behavior.
Approach: They propose a framework that provides fine-grained control over task generation in bAbI.
Outcome: The proposed framework provides fine-grained control over task generation in the bAbI benchmark.
Guiding Zero-Shot Paraphrase Generation with Fine-Grained Control Tokens (2023.starsem-1)

Copied to clipboard

Challenge: Sequence-to-sequence paraphrase generation models struggle with the generation of diverse paraphrases.
Approach: They propose a translation-based guided paraphrase generation model that learns useful features for promoting surface form variation in generated paraphrases from cross-lingual parallel data.
Outcome: The proposed model learns useful features for promoting surface form variation in generated paraphrases from cross-lingual parallel data.
On the Systematicity of Probing Contextualized Word Representations: The Case of Hypernymy in BERT (2020.starsem-1)

Copied to clipboard

Challenge: Existing studies have found that BERT can correctly retrieve noun hypernyms in cloze tasks, but this does not correspond to systematic knowledge in BERT.
Approach: They propose to use BERT to probe for hypernymy knowledge encoded in representations for cloze tasks to find out whether it is systematic or not .
Outcome: The proposed model can retrieve hypernyms in cloze tasks, but not systematic knowledge in BERT.
Find or Classify? Dual Strategy for Slot-Value Predictions on Multi-Domain Dialog State Tracking (2020.starsem-1)

Copied to clipboard

Challenge: Existing methods for dialog state tracking are ontology-based and ontologie-free . however, it is not clear enough which slots are better handled by either of the two methods .
Approach: They propose a dual-strategy model that integrates both ontology-based and ontological-free methods.
Outcome: The proposed model outperforms the existing model on noisy and cleaner datasets.
Compositional Structured Explanation Generation with Dynamic Modularized Reasoning (2024.starsem-1)

Copied to clipboard

Challenge: Large-scale language models have shown remarkable performance on reasoning tasks such as reading comprehension, natural language inference, story generation, etc.
Approach: They propose a compositional structured explanation generation task to test a model's ability to generalize from generating entailment trees to more steps, focusing on the length and shapes of engorgement trees.
Outcome: The proposed model shows competitive compositional generalization abilities in a generation setting.
What do Large Language Models Learn about Scripts? (2022.starsem-1)

Copied to clipboard

Challenge: Script Knowledge is important for language understanding but expensive to produce manually and difficult to induce from text due to reporting bias.
Approach: They propose a pipeline-based script induction framework which can generate good quality ESDs for unseen scenarios.
Outcome: The proposed framework produces good quality ESDs for unseen scenarios, but manual evaluation shows there is room for improvement.
Did the Cat Drink the Coffee? Challenging Transformers with Generalized Event Knowledge (2021.starsem-1)

Copied to clipboard

Challenge: Prior work has explored the ability of computational models to predict word semantic fit with a given predicate.
Approach: They compare Transformers Language Models to SDM to assess their performance . they found that TLMs do not capture important aspects of event knowledge . people can discriminate between typical and atypical events, they say .
Outcome: The proposed models can achieve comparable performance to SDM, but they lack important aspects of event knowledge.
BiQuAD: Towards QA based on deeper text understanding (2021.starsem-1)

Copied to clipboard

Challenge: Recent question answering and machine reading benchmarks require systems to pinpoint the span of the answer to a given text.
Approach: They propose a dataset that requires deeper comprehension to answer questions extractively and deductively.
Outcome: The proposed dataset outperforms existing benchmarks on extractive and deductive questions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations