Papers by Boi Faltings
Exploring Defeasibility in Causal Reasoning (2024.findings-acl)
Copied to clipboard
Shaobo Cui, Lazar Milikic, Yiyang Feng, Mete Ismayilzada, Debjit Paul, Antoine Bosselut, Boi Faltings
| Challenge: | Existing studies ignore defeasibility in causal reasoning and fail to evaluate existing causal strength metrics in defensible settings. |
| Approach: | They propose a metric that measures causal strength based on token-level causal relationships. |
| Outcome: | The proposed metric improves on existing metrics by 69.7% . supporters and defeaters are more effective than opponents, the authors show . |
Rationalization through Concepts (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing models that explain complex decisions are limited because of their lack of interpretability. |
| Approach: | They propose a model that extracts text snippets as concepts and infers which ones are described in the document. |
| Outcome: | The proposed model outperforms state-of-the-art methods trained on each aspect label independently. |
Conditional Dichotomy Quantification via Geometric Embedding (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods that rely on semantic similarity fail to capture the nuanced oppositional dynamics essential for these applications. |
| Approach: | They propose a task that formalizes the measurement of conditional dichotomy by using a dichotomian framework. |
| Outcome: | The proposed framework provides carefully constructed datasets covering debate, defeasible inference, and causal reasoning scenarios. |
Self-training Improves Pre-training for Few-shot Learning in Task-oriented Dialog Systems (2021.emnlp-main)
Copied to clipboard
| Challenge: | Large-scale pre-trained language models have shown promising results for few-shot learning in task-oriented dialog (ToD) systems. |
| Approach: | They propose a self-training approach that iteratively labels the most confident unlabeled data to train a stronger Student model. |
| Outcome: | The proposed approach improves state-of-the-art pre-trained models in few-shot learning scenarios for task-oriented dialog (ToD) systems when only a small number of labeled data are available. |
Unraveling Misinformation Propagation in LLM Reasoning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated impressive capabilities in reasoning, but how they propagate within their reasoning process remains underexplored. |
| Approach: | They propose a practical approach to mitigating misinformation propagation in LLMs by applying factual corrections early in the reasoning process and fine-tuning on synthesized data with early-stage corrections significantly improves reasoning factuality. |
| Outcome: | The proposed model can correct misinformation when explicitly instructed, but fails to correct misinformation less than half the time even with explicit instructions. |
GameWikiSum: a Novel Large Multi-Document Summarization Dataset (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing datasets contain only hundreds of samples, resulting in heavy reliance on hand-crafted features or manually annotated data. |
| Approach: | They propose a new domain-specific dataset for multi-document summarization that is 100 times larger than commonly used datasets. |
| Outcome: | The proposed dataset is 100 times larger than commonly used datasets and in another domain than news. |
Continual Learning for Natural Language Generation in Task-oriented Dialog Systems (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing neural approaches for natural language generation are typically developed offline for specific domains. |
| Approach: | They propose a method to expand NLG knowledge incrementally to new domains . major challenge is catastrophic forgetting, meaning a model forgets the knowledge it has learned before . |
| Outcome: | The proposed method outperforms other methods by effectively mitigating catastrophic forgetting issue. |
Assistive Recipe Editing through Critiquing (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing methods for generating recipes that satisfy dietary restrictions are inconsistent or incoherent and paired datasets are not available at scale. |
| Approach: | They propose to build a hierarchical denoising auto-encoder that edits recipes given ingredient-level critiques by interacting with the predicted ingredients. |
| Outcome: | The proposed model can more effectively edit recipes compared to strong language models and iteratively rewrites recipes to satisfy user feedback. |
The Odyssey of Commonsense Causality: From Foundational Benchmarks to Cutting-Edge Reasoning (2024.emnlp-main)
Copied to clipboard
| Challenge: | Despite its significance, a systematic exploration of commonsense causality is lacking. |
| Approach: | They focus on taxonomies, benchmarks, acquisition methods, qualitative reasoning, and quantitative measurements in commonsense causality. |
| Outcome: | The proposed method synthesizes insights from over 200 representative articles and provides a practical guide for beginners. |
Language Model Decoding as Likelihood–Utility Alignment (2023.findings-eacl)
Copied to clipboard
Martin Josifoski, Maxime Peyrard, Frano Rajič, Jiheng Wei, Debjit Paul, Valentin Hartmann, Barun Patra, Vishrav Chaudhary, Emre Kiciman, Boi Faltings
| Challenge: | Existing studies only compare decoding algorithms in narrow scenarios, and their findings do not generalize across tasks. |
| Approach: | They propose a taxonomy of misalignment mitigation strategies to provide a unifying view of decoding as a tool for alignment. |
| Outcome: | The proposed taxonomy combines likelihood and utility assumptions to provide general statements about decoding as a tool for alignment across tasks. |
A Logical Fallacy-Informed Framework for Argument Generation (2025.naacl-long)
Copied to clipboard
| Challenge: | Argument generation is crucial in daily life and has numerous online and offline applications. |
| Approach: | They propose a fallacy-informed preference optimization that includes a classification loss to capture the fine-grained information on fallacy types to help LLMs generate logically sound arguments. |
| Outcome: | The proposed method reduces fallacy errors by 17.5% on argument generation tasks and outperforms fine-tuned baselines and other preference optimization methods, such as DPO. |
Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) are shown to perform better when asked to reason step-by-step before generating a final answer. |
| Approach: | They propose a framework to tailor small-sized LMs to generate correct reasoning steps and robustly reason over these steps. |
| Outcome: | The proposed framework outperforms four competitive baselines and improves the robustness and generalization ability of the reasoning LM, yielding higher performance on out-of-distribution test sets. |
Learning to Create Sentence Semantic Relation Graphs for Multi-Document Summarization (D19-54)
Copied to clipboard
| Challenge: | Existing methods for summarizing documents rely on hand-crafted features or additional annotated data. |
| Approach: | They propose a method that makes use of two types of sentence embeddings . the method uses universal embeddable and domain-specific embeddible features . |
| Outcome: | The proposed method achieves competitive results on two types of summary, consisting of 665 bytes and 100 words. |
HotelRec: a Novel Very Large-Scale Hotel Recommendation Dataset (2020.lrec-1)
Copied to clipboard
| Challenge: | State-of-the-art deep learning-based recommender systems require large datasets to achieve their best performance. |
| Approach: | They propose to use TripAdvisor to build a large-scale hotel recommendation dataset with 50 million reviews. |
| Outcome: | The proposed dataset is the largest publicly available hotel recommendation dataset, based on TripAdvisor, with 50 million reviews. |
Uncertainty in Causality: A New Frontier (2025.acl-long)
Copied to clipboard
| Challenge: | Existing literature on uncertainty in causality is lacking a comprehensive review of this area. |
| Approach: | They propose a trichotomy categorizing causal uncertainty into aleatoric, epistemic, ontological and ontological categories . they propose key traits for an optimal causal LLM to handle uncertainty . |
| Outcome: | The proposed method categorizes causal uncertainty into aleatoric, epistemic, and ontological uncertainty. |
Unveiling the Art of Heading Design: A Harmonious Blend of Summarization, Neology, and Algorithm (2024.findings-acl)
Copied to clipboard
| Challenge: | Creating an appealing heading is crucial for attracting readers and marketing work or products. |
| Approach: | They propose a benchmark to measure the quality of heading generation using summarization, neology, and algorithm metrics. |
| Outcome: | The proposed benchmark compared 6,653 abstracts with corresponding descriptions and acronyms and found that it excels across summarization, neology, and algorithm aspects. |
REFINER: Reasoning Feedback on Intermediate Representations (2024.eacl-long)
Copied to clipboard
Debjit Paul, Mete Ismayilzada, Maxime Peyrard, Beatriz Borges, Antoine Bosselut, Robert West, Boi Faltings
| Challenge: | Language models (LLMs) have shown remarkable performance by explicitly generating intermediate inferences,e.g., chain-of-thought prompting. |
| Approach: | They propose a framework for finetuning LMs to generate intermediate reasoning steps while interacting with a critic model that provides automated feedback on the reasoning. |
| Outcome: | Empirical evaluations of REFINER on three diverse reasoning tasks show that it significantly improves over baseline models. |