Papers by Boi Faltings

17 papers
Exploring Defeasibility in Causal Reasoning (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies ignore defeasibility in causal reasoning and fail to evaluate existing causal strength metrics in defensible settings.
Approach: They propose a metric that measures causal strength based on token-level causal relationships.
Outcome: The proposed metric improves on existing metrics by 69.7% . supporters and defeaters are more effective than opponents, the authors show .
Rationalization through Concepts (2021.findings-acl)

Copied to clipboard

Challenge: Existing models that explain complex decisions are limited because of their lack of interpretability.
Approach: They propose a model that extracts text snippets as concepts and infers which ones are described in the document.
Outcome: The proposed model outperforms state-of-the-art methods trained on each aspect label independently.
Conditional Dichotomy Quantification via Geometric Embedding (2025.acl-long)

Copied to clipboard

Challenge: Existing methods that rely on semantic similarity fail to capture the nuanced oppositional dynamics essential for these applications.
Approach: They propose a task that formalizes the measurement of conditional dichotomy by using a dichotomian framework.
Outcome: The proposed framework provides carefully constructed datasets covering debate, defeasible inference, and causal reasoning scenarios.
Self-training Improves Pre-training for Few-shot Learning in Task-oriented Dialog Systems (2021.emnlp-main)

Copied to clipboard

Challenge: Large-scale pre-trained language models have shown promising results for few-shot learning in task-oriented dialog (ToD) systems.
Approach: They propose a self-training approach that iteratively labels the most confident unlabeled data to train a stronger Student model.
Outcome: The proposed approach improves state-of-the-art pre-trained models in few-shot learning scenarios for task-oriented dialog (ToD) systems when only a small number of labeled data are available.
Unraveling Misinformation Propagation in LLM Reasoning (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive capabilities in reasoning, but how they propagate within their reasoning process remains underexplored.
Approach: They propose a practical approach to mitigating misinformation propagation in LLMs by applying factual corrections early in the reasoning process and fine-tuning on synthesized data with early-stage corrections significantly improves reasoning factuality.
Outcome: The proposed model can correct misinformation when explicitly instructed, but fails to correct misinformation less than half the time even with explicit instructions.
GameWikiSum: a Novel Large Multi-Document Summarization Dataset (2020.lrec-1)

Copied to clipboard

Challenge: Existing datasets contain only hundreds of samples, resulting in heavy reliance on hand-crafted features or manually annotated data.
Approach: They propose a new domain-specific dataset for multi-document summarization that is 100 times larger than commonly used datasets.
Outcome: The proposed dataset is 100 times larger than commonly used datasets and in another domain than news.
Continual Learning for Natural Language Generation in Task-oriented Dialog Systems (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing neural approaches for natural language generation are typically developed offline for specific domains.
Approach: They propose a method to expand NLG knowledge incrementally to new domains . major challenge is catastrophic forgetting, meaning a model forgets the knowledge it has learned before .
Outcome: The proposed method outperforms other methods by effectively mitigating catastrophic forgetting issue.
Assistive Recipe Editing through Critiquing (2023.eacl-main)

Copied to clipboard

Challenge: Existing methods for generating recipes that satisfy dietary restrictions are inconsistent or incoherent and paired datasets are not available at scale.
Approach: They propose to build a hierarchical denoising auto-encoder that edits recipes given ingredient-level critiques by interacting with the predicted ingredients.
Outcome: The proposed model can more effectively edit recipes compared to strong language models and iteratively rewrites recipes to satisfy user feedback.
The Odyssey of Commonsense Causality: From Foundational Benchmarks to Cutting-Edge Reasoning (2024.emnlp-main)

Copied to clipboard

Challenge: Despite its significance, a systematic exploration of commonsense causality is lacking.
Approach: They focus on taxonomies, benchmarks, acquisition methods, qualitative reasoning, and quantitative measurements in commonsense causality.
Outcome: The proposed method synthesizes insights from over 200 representative articles and provides a practical guide for beginners.
Language Model Decoding as Likelihood–Utility Alignment (2023.findings-eacl)

Copied to clipboard

Challenge: Existing studies only compare decoding algorithms in narrow scenarios, and their findings do not generalize across tasks.
Approach: They propose a taxonomy of misalignment mitigation strategies to provide a unifying view of decoding as a tool for alignment.
Outcome: The proposed taxonomy combines likelihood and utility assumptions to provide general statements about decoding as a tool for alignment across tasks.
A Logical Fallacy-Informed Framework for Argument Generation (2025.naacl-long)

Copied to clipboard

Challenge: Argument generation is crucial in daily life and has numerous online and offline applications.
Approach: They propose a fallacy-informed preference optimization that includes a classification loss to capture the fine-grained information on fallacy types to help LLMs generate logically sound arguments.
Outcome: The proposed method reduces fallacy errors by 17.5% on argument generation tasks and outperforms fine-tuned baselines and other preference optimization methods, such as DPO.
Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are shown to perform better when asked to reason step-by-step before generating a final answer.
Approach: They propose a framework to tailor small-sized LMs to generate correct reasoning steps and robustly reason over these steps.
Outcome: The proposed framework outperforms four competitive baselines and improves the robustness and generalization ability of the reasoning LM, yielding higher performance on out-of-distribution test sets.
Learning to Create Sentence Semantic Relation Graphs for Multi-Document Summarization (D19-54)

Copied to clipboard

Challenge: Existing methods for summarizing documents rely on hand-crafted features or additional annotated data.
Approach: They propose a method that makes use of two types of sentence embeddings . the method uses universal embeddable and domain-specific embeddible features .
Outcome: The proposed method achieves competitive results on two types of summary, consisting of 665 bytes and 100 words.
HotelRec: a Novel Very Large-Scale Hotel Recommendation Dataset (2020.lrec-1)

Copied to clipboard

Challenge: State-of-the-art deep learning-based recommender systems require large datasets to achieve their best performance.
Approach: They propose to use TripAdvisor to build a large-scale hotel recommendation dataset with 50 million reviews.
Outcome: The proposed dataset is the largest publicly available hotel recommendation dataset, based on TripAdvisor, with 50 million reviews.
Uncertainty in Causality: A New Frontier (2025.acl-long)

Copied to clipboard

Challenge: Existing literature on uncertainty in causality is lacking a comprehensive review of this area.
Approach: They propose a trichotomy categorizing causal uncertainty into aleatoric, epistemic, ontological and ontological categories . they propose key traits for an optimal causal LLM to handle uncertainty .
Outcome: The proposed method categorizes causal uncertainty into aleatoric, epistemic, and ontological uncertainty.
Unveiling the Art of Heading Design: A Harmonious Blend of Summarization, Neology, and Algorithm (2024.findings-acl)

Copied to clipboard

Challenge: Creating an appealing heading is crucial for attracting readers and marketing work or products.
Approach: They propose a benchmark to measure the quality of heading generation using summarization, neology, and algorithm metrics.
Outcome: The proposed benchmark compared 6,653 abstracts with corresponding descriptions and acronyms and found that it excels across summarization, neology, and algorithm aspects.
REFINER: Reasoning Feedback on Intermediate Representations (2024.eacl-long)

Copied to clipboard

Challenge: Language models (LLMs) have shown remarkable performance by explicitly generating intermediate inferences,e.g., chain-of-thought prompting.
Approach: They propose a framework for finetuning LMs to generate intermediate reasoning steps while interacting with a critic model that provides automated feedback on the reasoning.
Outcome: Empirical evaluations of REFINER on three diverse reasoning tasks show that it significantly improves over baseline models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations