Papers by Sina Semnani

15 papers
Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in language models (LMs) and retrieval-augmented generation (RAG) have led to more capable chatbots and generative search engines.
Approach: They propose to emulate the educational scenario where children/students learn by listening to and participating in conversations of their parents/teachers by watching and steering the discourse among several LM agents.
Outcome: The proposed system outperforms baseline methods on discourse trace and report quality and is preferred by 70% of participants over a search engine and 78% over sabota.
A Few-Shot Semantic Parser for Wizard-of-Oz Dialogues with the Precise ThingTalk Representation (2022.findings-acl)

Copied to clipboard

Challenge: Existing approaches to build effective semantic parsers for Wizard-of-Oz are insufficient.
Approach: They propose a new dialogue representation and a sample-efficient methodology that can predict precise dialogue states in WOZ conversations.
Outcome: The proposed model can predict precise dialogue states in WOZ conversations.
SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources with Retrieval and Semantic Parsing (2024.findings-acl)

Copied to clipboard

Challenge: SPAGHETTI: Semantic Parsing Augmented Generation for Hybrid English information from Text Tables and Infoboxes is a hybrid question-answering pipeline .
Approach: They propose a hybrid question-answering pipeline that leverages knowledge from multiple knowledge sources.
Outcome: The proposed approach achieves state-of-the-art on the Compmix dataset with 56.5% exact match rate.
AutoQA: From Databases To QA Semantic Parsers With Only Synthetic Training Data (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods to generate semantic parsers that answer questions on databases require large amounts of annotated data.
Approach: They propose a method to generate semantic parsers that answer questions on databases . they use automatic paraphrasing and template-based parsing to find alternative expressions .
Outcome: The proposed method achieves 69.8% answer accuracy on natural questions, 16.4% higher than state-of-the-art models and 5.2% lower than the same model trained with human data.
Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: a new study examines the accuracy of Wikipedia's factual inconsistencies . a corpus-level inconsistent detection system can help editors identify inconsistances .
Approach: They propose a corpus-level inconsistency detection system that combines LLM reasoning with retrieval to detect and contextualize potential contradictions for human review.
Outcome: The proposed system can detect inconsistencies in Wikipedia and human review.
X-RiSAWOZ: High-Quality End-to-End Multilingual Dialogue Datasets and Few-shot Agents (2023.findings-acl)

Copied to clipboard

Challenge: X-RiSAWOZ dataset has more than 18,000 human-verified dialogue utterances for each language . Xiaoping and Xinhui are the main challenges for task-oriented dialogue research .
Approach: They develop a toolkit to accelerate the post-editing of a new language dataset after translation . their dataset, code, and toolkit are released open-source .
Outcome: The proposed toolkit accelerates the post-editing of a new language dataset after translation.
CHURRO: Making History Readable with an Open-Weight Large Vision-Language Model for High-Accuracy, Low-Cost Historical Text Recognition (2025.emnlp-main)

Copied to clipboard

Challenge: Existing vision-language models are not equipped to read diverse languages and scripts found in historical materials.
Approach: They propose to train an open-weight vision-language model for historical text recognition on CHURRO-DS, the largest historical text-recognition dataset to date.
Outcome: The proposed model outperforms existing vision-language models on CHURRO-DS, the largest historical text recognition dataset to date.
SPINACH: SPARQL-Based Information Navigation for Challenging Real-World Questions (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have led to significant improvements in the Knowledge Base Question Answering task.
Approach: They introduce an expert-annotated KBQA dataset from Wikidata’s “Request a Query” forum with 320 decontextualized question-SPARQL pairs.
Outcome: The SPINACH dataset outperforms baselines on the QALD-7, QADL-9 Plus and QAL-10 datasets by 31.0%, 27.0% and 10.0% in F1 respectively.
WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia (2023.findings-emnlp)

Copied to clipboard

Challenge: a new few-shot LLM-based chatbot is able to provide factual and engaging responses . a novel hybrid human-and-LLM evaluation methodology is used to evaluate the system .
Approach: They propose a few-shot LLM-based chatbot that almost never hallucinates . they distill WikiChat into a 7B-parameter LLaMA model with minimal loss of quality .
Outcome: The proposed system outperforms retrieval-based and LLM-based systems on the Wikipedia corpus.
Localizing Open-Ontology QA Semantic Parsers in a Day Using Machine Translation (2020.emnlp-main)

Copied to clipboard

Challenge: a new toolkit for localizing a semantic parser for a language is proposed . the proposed approach is based on a method for question answering systems .
Approach: They propose a toolkit that leverages Neural Machine Translation systems to localize a semantic parser for a new language.
Outcome: The proposed approach outperforms state-of-the-art methods in 10 new languages . it can be deployed in restaurants and hotels in less than 24 hours .
LEMONADE: A Large Multilingual Expert-Annotated Abstractive Event Dataset for the Real World (2025.findings-acl)

Copied to clipboard

Challenge: Using a partially reannotated subset of the Armed Conflict Location & Event Data, we analyze 39,786 conflict events across 20 languages and 171 countries.
Approach: They propose a large-scale conflict event dataset with extensive coverage of region-specific entities.
Outcome: The proposed method detects event arguments and entities through holistic document understanding and normalizes them across the multilingual dataset.
Fine-tuned LLMs Know More, Hallucinate Less with Few-Shot Sequence-to-Sequence Semantic Parsing over Wikidata (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models can answer many questions correctly, but can also hallucinate and give wrong answers.
Approach: They propose a question-answering benchmark for Wikidata that uses SPARQL to ground large language models.
Outcome: The proposed method outperforms the state-of-the-art for QALD-7 by 3.6% in F1 score.
Zero-shot Persuasive Chatbots with LLM-Generated Strategies and Information Retrieval (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to improve persuasive chatbots use only a handful of predefined strategies.
Approach: They propose a persuasive chatbot based on large language models that is factual and more persuasive by leveraging many more nuanced strategies.
Outcome: The proposed chatbot is factual and more persuasive by leveraging many more nuanced strategies.
SUQL: Conversational Search over Structured and Unstructured Data with Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: SUQL is a conversational language that supports the generality of hybrid data access for large knowledge corpora.
Approach: They propose a conversational agent that supports the full generality of hybrid data access for large knowledge corpora using SUQL.
Outcome: The proposed language can handle hybrid data sources.
Zero and Few-Shot Localization of Task-Oriented Dialogue Agents with a Distilled Representation (2023.eacl-main)

Copied to clipboard

Challenge: Existing low-cost approaches to build a high-quality functioning dialogue agent are limited to a few widely-spoken languages.
Approach: They propose automatic methods that use ToD training data to build a functioning agent in another language . they compare the method to existing methods that only use a small training set .
Outcome: The proposed method improves the state-of-the-art in Chinese to English transfer using zero-shot data compared to existing full-shot methods . the proposed method achieves 46.7% and 22.0% in task success rate and dialogue success rate, respectively.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations