Papers by Muhammad Khalifa

12 papers
BOLT: Fast Energy-based Controlled Text Generation with Tunable Biases (2023.acl-short)

Copied to clipboard

Challenge: Energy-based models (EBMs) have gained popularity for controlled text generation due to their high applicability to a wide range of constraints.
Approach: They propose a language model with tunable biases to adjust the language model’s output logits.
Outcome: The proposed model maintains the generator’s autoregressive nature to assert a strong control on token-wise conditional dependencies and overall fluency, and converges faster.
Self-Training Pre-Trained Language Models for Zero- and Few-Shot Multi-Dialectal Arabic Sequence Labeling (2021.eacl-main)

Copied to clipboard

Challenge: Existing approaches to fine-tune pre-trained language models for downstream tasks require labeled data.
Approach: They propose to self-train pre-trained language models to improve performance on data-scarce varieties by as large as 10% F1 and 2% accuracy.
Outcome: The proposed model improves zero-shot MSA-to-DA transfer by as large as 10% F1 (NER) and 2% accuracy (POS tagging).
GRACE: Discriminator-Guided Chain-of-Thought Reasoning (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing language models (LMs) can assign a high likelihood to incorrect steps . Existing models (LLMs), however, struggle with complex multi-step reasoning.
Approach: They propose a stepwise decoding approach that steers the decoding process towards producing correct reasoning steps.
Outcome: The proposed approach outperforms existing methods on math and symbolic reasoning tasks.
A Bag of Tricks for Dialogue Summarization (2021.emnlp-main)

Copied to clipboard

Challenge: Using a pretrained sequence-to-sequence language model, we explore speaker name substitution, negation scope highlighting, multi-task learning with relevant tasks, and pretraining on in-domain data.
Approach: They propose a pretrained sequence-to-sequence language model that can handle different parts of dialogue belonging to multiple speakers and combine them to produce a coherent monologue summary.
Outcome: The proposed techniques outperform baseline models on a dialogue summarization dataset.
FACTCHECKMATE: Preemptively Detecting and Mitigating Hallucinations in LMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Language models (LMs) hallucinate.
Approach: They introduce a classifier that predicts whether LMs hallucinate based on model’s hidden states before decoding begins.
Outcome: The proposed model preemptively detects hallucinations by learning a classifier that predicts whether the LM will hallucinate . if a hallucinomy is detected, FactCheckmate intervenes by adjusting the model’s hidden states to produce more factual outputs.
Few-shot Reranking for Multi-hop QA via Language Model Prompting (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for multi-hop QA with open-domain questions require a large number of labeled question-document pairs for retrieval.
Approach: They propose a language-based prompt for multi-hop path reranking that relies on language model prompting to generate a relevance score between a question and the path.
Outcome: The proposed method yields strong retrieval performance on HotpotQA with only 128 training examples compared to state-of-the-art methods trained on thousands of examples.
Contrastive Training Improves Zero-Shot Classification of Semi-structured Documents (2023.findings-acl)

Copied to clipboard

Challenge: Xu et al., 2020 focus on semi-structured document classification in a zero-shot setting . positional, layout, and style information play a vital role in interpreting such documents .
Approach: They propose a matching-based approach that relies on a pairwise contrastive objective for pretraining and fine-tuning.
Outcome: The proposed method significantly improves Macro F1 in the zero-shot learning setting.
Logit Arithmetic Elicits Long Reasoning Capabilities Without Training (2026.findings-acl)

Copied to clipboard

Challenge: Large reasoning models exhibit long chain-of-thought reasoning with complex strategies such as backtracking and self-verification, yet, these capabilities typically require resource-intensive post-training.
Approach: They propose a decoding-time approach which transfers long chain-of-thought reasoning capabilities from a substantially smaller reasoning guider to a large non-reasoning target.
Outcome: The proposed method improves performance over a model 21x smaller than the target model by 21.5% and 24.2% over the model.
Merging Generated and Retrieved Knowledge for Open-Domain QA (2023.emnlp-main)

Copied to clipboard

Challenge: Open-domain question answering systems often have retrieval modules but retrieving passages from external knowledge sources is known to suffer from insufficient knowledge coverage.
Approach: They propose a Compatibility-Oriented knowledge Merging framework to leverage both sources of information by matching LLM-generated passages with retrieved counterparts into compatible pairs.
Outcome: The proposed framework outperforms baselines on three out of four tested open-domain QA benchmarks.
On Many-Shot In-Context Learning for Long-Context Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks primarily evaluate long-context language models' retrieval capabilities.
Approach: They propose a benchmark to evaluate long-context language models' retrieval capabilities by using MANYICLBENCH.
Outcome: The proposed model performs better with additional demonstrations than translation and reasoning tasks.
Has It All Been Solved? Open NLP Research Questions Not Solved by Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in large language models have led to misleading public discourse that “it’s all been solved.”
Approach: They identify 14 research areas encompassing 45 research directions that require new research and are not directly solvable by LLMs.
Outcome: The research areas identified are 45 research directions that require new research and are not directly solvable by LLMs.
Small Language Models Need Strong Verifiers to Self-Correct Reasoning (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies show that large language models can self-correct their outputs by generating a critique and revising it based on the critique.
Approach: They propose a pipeline that prompts small language models to collect self-correction data that supports the training of self-refinement abilities.
Outcome: The proposed pipeline improves the self-correction abilities of two models on five datasets spanning math and commonsense reasoning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations