Papers by Patrick Zhao

5 papers
Importance of Synthesizing High-quality Data for Text-to-SQL Parsing (2023.findings-acl)

Copied to clipboard

Challenge: Existing text-to-SQL parsers lack the data to perform well with augmented synthetic data.
Approach: They propose a framework that imposes strong typing constraints and incorporates key relationships from schema.
Outcome: The proposed framework improves on the high-quality synthesized SQL and natural language question (NLQ) models have significant accuracy boosts and achieve new state-of-the-art performance on spider.
FEED PETs: Further Experimentation and Expansion on the Disambiguation of Potentially Euphemistic Terms (2023.starsem-1)

Copied to clipboard

Challenge: Existing work on euphemism disambiguation tasks has focused on transformers . euphorias are expressions that soften the message they convey, therefore dictionary-based approaches are ineffective .
Approach: They propose to annotate PETs for vagueness and use transformers to classify PETs . they perform euphemism disambiguation experiments in three different languages .
Outcome: The proposed models perform well in English euphemism disambiguation task . preliminary results will be used to launch future work .
Scaling Parameter-Constrained Language Models with Quality Data (2024.emnlp-industry)

Copied to clipboard

Challenge: Scaling laws in language modeling quantify training loss as a function of dataset size and model parameters, but neglect the critical role of data quality in model generalization.
Approach: They propose to use effective training tokens as a combination of text diversity and syntheticity as measured by a teacher model to calculate scaling laws.
Outcome: The proposed term effective training tokens is a combination of two readily-computed indicators of text diversity and syntheticity as measured by a teacher model.
Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences (2025.findings-emnlp)

Copied to clipboard

Challenge: Current LLMs are trained to refuse potentially harmful input queries regardless of intent . a study of 480 participants evaluating 3,840 query-response pairs reveals that response strategy largely shapes user experience .
Approach: They examine how different refusal strategies affect user perceptions across varying motivations . they find partial compliance reduces negative user perception by over 50% to flat-out refusals a 480 participants study .
Outcome: The study examines the perceptions of LLMs on user intents and their response strategies . it shows that partial compliance reduces negative user perceptions by over 50% to flat refusals .
MEDs for PETs: Multilingual Euphemism Disambiguation for Potentially Euphemistic Terms (2024.findings-eacl)

Copied to clipboard

Challenge: Euphemisms are a linguistic device used to soften or neutralize language that may otherwise be harsh or awkward to state directly.
Approach: They train a multilingual transformer model to disambiguate potentially euphemistic terms in multilingual and cross-lingual settings.
Outcome: The proposed model performs better than monolingual models on the disambiguation task compared to monolingual ones in multilingual and cross-lingual settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations