Papers by Muhammad Khalifa
BOLT: Fast Energy-based Controlled Text Generation with Tunable Biases (2023.acl-short)
Copied to clipboard
| Challenge: | Energy-based models (EBMs) have gained popularity for controlled text generation due to their high applicability to a wide range of constraints. |
| Approach: | They propose a language model with tunable biases to adjust the language model’s output logits. |
| Outcome: | The proposed model maintains the generator’s autoregressive nature to assert a strong control on token-wise conditional dependencies and overall fluency, and converges faster. |
Self-Training Pre-Trained Language Models for Zero- and Few-Shot Multi-Dialectal Arabic Sequence Labeling (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing approaches to fine-tune pre-trained language models for downstream tasks require labeled data. |
| Approach: | They propose to self-train pre-trained language models to improve performance on data-scarce varieties by as large as 10% F1 and 2% accuracy. |
| Outcome: | The proposed model improves zero-shot MSA-to-DA transfer by as large as 10% F1 (NER) and 2% accuracy (POS tagging). |
GRACE: Discriminator-Guided Chain-of-Thought Reasoning (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing language models (LMs) can assign a high likelihood to incorrect steps . Existing models (LLMs), however, struggle with complex multi-step reasoning. |
| Approach: | They propose a stepwise decoding approach that steers the decoding process towards producing correct reasoning steps. |
| Outcome: | The proposed approach outperforms existing methods on math and symbolic reasoning tasks. |
A Bag of Tricks for Dialogue Summarization (2021.emnlp-main)
Copied to clipboard
| Challenge: | Using a pretrained sequence-to-sequence language model, we explore speaker name substitution, negation scope highlighting, multi-task learning with relevant tasks, and pretraining on in-domain data. |
| Approach: | They propose a pretrained sequence-to-sequence language model that can handle different parts of dialogue belonging to multiple speakers and combine them to produce a coherent monologue summary. |
| Outcome: | The proposed techniques outperform baseline models on a dialogue summarization dataset. |
FACTCHECKMATE: Preemptively Detecting and Mitigating Hallucinations in LMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Language models (LMs) hallucinate. |
| Approach: | They introduce a classifier that predicts whether LMs hallucinate based on model’s hidden states before decoding begins. |
| Outcome: | The proposed model preemptively detects hallucinations by learning a classifier that predicts whether the LM will hallucinate . if a hallucinomy is detected, FactCheckmate intervenes by adjusting the model’s hidden states to produce more factual outputs. |
Few-shot Reranking for Multi-hop QA via Language Model Prompting (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods for multi-hop QA with open-domain questions require a large number of labeled question-document pairs for retrieval. |
| Approach: | They propose a language-based prompt for multi-hop path reranking that relies on language model prompting to generate a relevance score between a question and the path. |
| Outcome: | The proposed method yields strong retrieval performance on HotpotQA with only 128 training examples compared to state-of-the-art methods trained on thousands of examples. |
Contrastive Training Improves Zero-Shot Classification of Semi-structured Documents (2023.findings-acl)
Copied to clipboard
| Challenge: | Xu et al., 2020 focus on semi-structured document classification in a zero-shot setting . positional, layout, and style information play a vital role in interpreting such documents . |
| Approach: | They propose a matching-based approach that relies on a pairwise contrastive objective for pretraining and fine-tuning. |
| Outcome: | The proposed method significantly improves Macro F1 in the zero-shot learning setting. |
Logit Arithmetic Elicits Long Reasoning Capabilities Without Training (2026.findings-acl)
Copied to clipboard
Yunxiang Zhang, Muhammad Khalifa, Lechen Zhang, Xin Liu, Ayoung Lee, Xinliang Frederick Zhang, Farima Fatahi Bayat, Lu Wang
| Challenge: | Large reasoning models exhibit long chain-of-thought reasoning with complex strategies such as backtracking and self-verification, yet, these capabilities typically require resource-intensive post-training. |
| Approach: | They propose a decoding-time approach which transfers long chain-of-thought reasoning capabilities from a substantially smaller reasoning guider to a large non-reasoning target. |
| Outcome: | The proposed method improves performance over a model 21x smaller than the target model by 21.5% and 24.2% over the model. |
Merging Generated and Retrieved Knowledge for Open-Domain QA (2023.emnlp-main)
Copied to clipboard
| Challenge: | Open-domain question answering systems often have retrieval modules but retrieving passages from external knowledge sources is known to suffer from insufficient knowledge coverage. |
| Approach: | They propose a Compatibility-Oriented knowledge Merging framework to leverage both sources of information by matching LLM-generated passages with retrieved counterparts into compatible pairs. |
| Outcome: | The proposed framework outperforms baselines on three out of four tested open-domain QA benchmarks. |
On Many-Shot In-Context Learning for Long-Context Evaluation (2025.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks primarily evaluate long-context language models' retrieval capabilities. |
| Approach: | They propose a benchmark to evaluate long-context language models' retrieval capabilities by using MANYICLBENCH. |
| Outcome: | The proposed model performs better with additional demonstrations than translation and reasoning tasks. |
Has It All Been Solved? Open NLP Research Questions Not Solved by Large Language Models (2024.lrec-main)
Copied to clipboard
Oana Ignat, Zhijing Jin, Artem Abzaliev, Laura Biester, Santiago Castro, Naihao Deng, Xinyi Gao, Aylin Ece Gunal, Jacky He, Ashkan Kazemi, Muhammad Khalifa, Namho Koh, Andrew Lee, Siyang Liu, Do June Min, Shinka Mori, Joan C. Nwatu, Veronica Perez-Rosas, Siqi Shen, Zekun Wang, Winston Wu, Rada Mihalcea
| Challenge: | Recent advances in large language models have led to misleading public discourse that “it’s all been solved.” |
| Approach: | They identify 14 research areas encompassing 45 research directions that require new research and are not directly solvable by LLMs. |
| Outcome: | The research areas identified are 45 research directions that require new research and are not directly solvable by LLMs. |
Small Language Models Need Strong Verifiers to Self-Correct Reasoning (2024.findings-acl)
Copied to clipboard
Yunxiang Zhang, Muhammad Khalifa, Lajanugen Logeswaran, Jaekyeom Kim, Moontae Lee, Honglak Lee, Lu Wang
| Challenge: | Existing studies show that large language models can self-correct their outputs by generating a critique and revising it based on the critique. |
| Approach: | They propose a pipeline that prompts small language models to collect self-correction data that supports the training of self-refinement abilities. |
| Outcome: | The proposed pipeline improves the self-correction abilities of two models on five datasets spanning math and commonsense reasoning. |