Papers by Prasanna Parthasarathi
Extending Neural Generative Conversational Model using External Knowledge Sources (D18-1)
Copied to clipboard
| Challenge: | Existing generative dialogue models lack coherence and are content poor . however, current models lack the capacity to handle large unstructured knowledge sources. |
| Approach: | They propose an architecture to incorporate unstructured knowledge sources to enhance the next utterance prediction in chit-chat type of generative dialogue models. |
| Outcome: | The proposed architecture improves the next utterance prediction in chit-chat type of generative dialogue models by incorporating external knowledge from Wikipedia summaries and the NELL knowledge base. |
Detecting Languages Unintelligible to Multilingual Models through Local Structure Probes (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in multilingual pretrained models have proven effective at zero-shot transfer to a wide variety of languages, but this transfer is not universal, with many languages not currently understood by multilingual approaches. |
| Approach: | They propose a general approach that requires only unlabelled text to detect which languages are not well understood by a cross-lingual model. |
| Outcome: | The proposed model can detect which languages are not well understood by a multilingual model on 350 low-resource languages. |
Do Large Language Models Know How Much They Know? (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models are highly capable systems, but their capabilities and limitations are unclear. |
| Approach: | They develop a benchmark that challenges LLMs to recall all information they possess on specific topics. |
| Outcome: | The proposed model can recall excessive, insufficient, or the precise amount of information they possess on a given topic, indicating their awareness of how much they know about the given topic. |
CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems (2024.findings-acl)
Copied to clipboard
Abbas Ghaddar, David Alfonso-Hermelo, Philippe Langlais, Mehdi Rezagholizadeh, Boxing Chen, Prasanna Parthasarathi
| Challenge: | CHARP is a testbed for knowledge-grounded dialogue evaluation of models trained on FaithDial data. |
| Approach: | They propose a testbed for evaluating models trained on FaithDial with annotation artifacts that may bias models towards completely ignoring the conversation history. |
| Outcome: | The proposed model fails to accurately evaluate the conversational history and lacks hallucination detection. |
Local Structure Matters Most: Perturbation Study in NLU (2022.findings-acl)
Copied to clipboard
| Challenge: | Recent research shows that neural models are insensitive to word-order perturbations, but other studies suggest that models learn some abstract notion of syntax. |
| Approach: | They develop order-altering perturbations on the order of words, subwords, and characters to analyze their effect on neural models’ performance on language understanding tasks. |
| Outcome: | The proposed models are insensitive to word-order perturbations while the local ordering remains relatively unperturbed. |
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination (2025.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) have a tendency to hallucinate false or misleading information, limiting their reliability. |
| Approach: | They examine how architecture-based inductive biases affect the propensity to hallucinate . they find that the models are more reliable and more reliable than traditional models . |
| Outcome: | The proposed models can be used to train and train large language models that are factual or able to explain themselves through their knowledge. |
Measuring the Knowledge Acquisition-Utilization Gap in Pretrained Language Models (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent research has demonstrated that pre-trained language models acquire a broad range of knowledge about linguistic structures, encyclopedic relations, levels of commonsense, and even coding and reasoning rules. |
| Approach: | They propose a systematic framework to measure parametric knowledge utilization in pre-trained language models by extracting parametric information from a PLM and constructing a downstream task around this extracted knowledge. |
| Outcome: | The proposed framework extracts parametric knowledge from a PLM and constructs a downstream task around this extracted knowledge. |
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are prohibitive to use under resource constraints due to their high latency and high latex. |
| Approach: | They propose to use a contextual bandit to help choose a model based on a context to improve performance. |
| Outcome: | The proposed model can be used to improve performance on multiple domains even without prior knowledge of the model. |
Local Structure Matters Most in Most Languages (2022.aacl-short)
Copied to clipboard
| Challenge: | Recent perturbation studies have found unintuitive results on what does and does not matter when performing Natural Language Understanding (NLU) tasks in English. |
| Approach: | They replicate a study on the importance of local structure and relative unimportance of global structure in a multilingual setting. |
| Outcome: | The proposed model replicates a study on the importance of local structure and relative unimportance of global structure in a multilingual setting. |
Thinking Long, but Short: Stable Sequential Test-Time Scaling for Large Reasoning Models (2026.findings-eacl)
Copied to clipboard
| Challenge: | Inducing models to think for longer can increase accuracy, but as the length of reasoning is further extended, it has also been shown to result in accuracy degradation and model instability. |
| Approach: | They propose a sequential test-time scaling method which induces models to think for longer, but which also generates an increasingly long output. |
| Outcome: | The proposed method improves model accuracy significantly over a wide range of induced thoughts, stabilizing the accuracy of sequential scaling, and eliminating the need for reasoning length fine-tuning. |
Learning an Unreferenced Metric for Online Dialogue Evaluation (2020.acl-main)
Copied to clipboard
| Challenge: | Existing tools for dialogue evaluation do not generalize to unseen datasets and/or need a human-generated reference response during inference. |
| Approach: | They propose an unreferenced automated dialogue evaluation metric that uses large pre-trained language models to extract latent representations of utterances and leverages the temporal transitions that exist between them. |
| Outcome: | The proposed model achieves higher correlation with human annotations in an online setting, while not requiring true responses for comparison during inference. |
EpiK-Eval: Evaluation for Language Models as Epistemic Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Developing systems that can reason through language understanding has been a cornerstone in natural language processing research. |
| Approach: | They propose a question-answering benchmark to evaluate LLMs' ability to combine knowledge from different training documents within their parameter space. |
| Outcome: | The proposed benchmark aims to evaluate LLMs' ability to combine knowledge from different training documents within their parameter space. |
Sometimes We Want Ungrammatical Translations (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in Neural Machine Translation (NMT) systems focus on improving translation quality and improving robustness to perturbations. |
| Approach: | They propose a way to quantify faithfulness to the original text by focusing on word-order perturbations. |
| Outcome: | The proposed method aims to measure faithfulness and robustness in word-order perturbations without deleting or injecting tokens. |
UnNatural Language Inference (2021.acl-long)
Copied to clipboard
| Challenge: | Recent studies show that pre-trained NLU models understand human-like syntax . however, these models are word order invariant, causing them to assign gold labels to permutations . |
| Approach: | They propose to measure the severity of this issue by examining the properties of particular permutations that lead models to be word order invariant. |
| Outcome: | The proposed model is word order invariant, but it's not human-like syntax. |
CHARPEVAL: Benchmarking Large Language Models’ Contextual Reasoning in Knowledge-Grounded Dialogue (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks that evaluate the ability of Large Language Models (LLMs) to perform contextualized reasoning in knowledge-grounded dialogue scenarios are lacking. |
| Approach: | They propose a benchmark to evaluate the ability of Large Language Models to perform contextualized reasoning in knowledge-grounded dialogue scenarios. |
| Outcome: | The proposed benchmark shows that open-weight LLMs are ineffective at reasoning over discontinuous chunks of text across the input. |
EWEK-QA : Enhanced Web and Efficient Knowledge Graph Retrieval for Citation-based Question Answering Systems (2024.acl-long)
Copied to clipboard
Mohammad Dehghan, Mohammad Alomrani, Sunyam Bagga, David Alfonso-Hermelo, Khalil Bibi, Abbas Ghaddar, Yingxue Zhang, Xiaoguang Li, Jianye Hao, Qun Liu, Jimmy Lin, Boxing Chen, Prasanna Parthasarathi, Mahdi Biparva, Mehdi Rezagholizadeh
| Challenge: | citation-based QA systems are suffering from two shortcomings . they usually rely only on web as a source of extracted knowledge and external knowledge sources can hamper the efficiency of the system. |
| Approach: | They propose to use a web-based knowledge graph retrieval solution to enrich extracted knowledge fed to a citation-based QA system. |
| Outcome: | The proposed model outperforms open-source state-of-the-art models in 7 quantitative and human evaluation tasks. |
Practical Takes on Federated Learning with Pretrained Language Models (2023.findings-eacl)
Copied to clipboard
| Challenge: | federated learning with pretrained language models for language tasks entails data privacy constraints when learning from diverse data domains. |
| Approach: | They propose to use pretrained language models to learn from diverse data domains . they elaborate hypotheses over the components in federated NLP architectures based on three tasks . |
| Outcome: | The proposed model can generalize by adapting to the different domains. |