Papers by Yao Dou
Simplified Graph Learning for Inductive Short Text Classification (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for short text classification are limited and lack of labeled data is not enough. |
| Approach: | They propose a novel short text classification algorithm which leverages words to handle the lack of labeled data. |
| Outcome: | The proposed model performs better with lower memory consumption and faster inference speed than previous models. |
Dancing Between Success and Failure: Edit-level Simplification Evaluation using SALSA (2023.emnlp-main)
Copied to clipboard
| Challenge: | Traditional human evaluation methods for text simplification often relies on individual, shallow sentence-level ratings, easily affected by the annotator's preference or bias. |
| Approach: | They propose an edit-based human annotation framework that enables holistic and fine-grained text simplification evaluation. |
| Outcome: | The proposed framework is able to predict sentence- and word-level quality simultaneously and report promising results. |
Improving Large-scale Paraphrase Acquisition and Generation (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing Twitter-based paraphrase datasets lack quality definitions for identification and generation tasks. |
| Approach: | They propose to use two separate definitions of paraphrase for identification and generation tasks in existing Twitter-based paraphrase datasets. |
| Outcome: | The proposed model achieves state-of-the-art performance of 84.2 F1 for automatic paraphrase identification compared to other models fine-tuned on other corpora such as Quora, MSCOCO, and ParaNMT. |
Value Compass Benchmarks: A Comprehensive, Generative and Self-Evolving Platform for LLMs’ Value Evaluation (2025.acl-demo)
Copied to clipboard
Jing Yao, Xiaoyuan Yi, Shitong Duan, Jindong Wang, Yuzhuo Bai, Muhua Huang, Yang Ou, Scarlett Li, Peng Zhang, Tun Lu, Zhicheng Dou, Maosong Sun, James Evans, Xing Xie
| Challenge: | Current evaluation methods for large language models face two key challenges: 1. evaluation validity and 2. Result interpretation reduce the pluralistic and incommensurable values to one-dimensional scores. |
| Approach: | They propose a platform for comprehensive value diagnosis of large language models (LLMs) that provides a generative evaluation paradigm that automatically creates real-world test items co-evolving with ever-advancing LLMs. |
| Outcome: | The proposed platform provides a framework for comprehensive value diagnosis of large language models (LLMs) with fine-grained scores and case studies across 27 value dimensions for 33 leading LLMs, customized comparisons, and visualized analysis of LLM’s alignment with cultural values. |
SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants? (2025.emnlp-main)
Copied to clipboard
Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie, Weixin Cai, Alan Ritter, Chris Quirk, Wei Xu, Jianfeng Gao
| Challenge: | Large language models (LLMs) are increasingly used in interactive applications, and human evaluation remains the gold standard for assessing their performance in multi-turn conversations. |
| Approach: | They propose to use large language models to simulate users for automatic assistant evaluation. |
| Outcome: | The proposed model outperforms human evaluations on two interactive tasks and achieves Spearman’s of 0.7 on both tasks. |
Data2Text Studio: Automated Text Generation from Structured Data (D18-2)
Copied to clipboard
| Challenge: | Data2Text Studio is a platform for automated text generation from structured data. |
| Approach: | They conduct experiments on RotoWire datasets for template extraction and text generation . they find that the Semi-HMMs model improves interactivity and interpretability . |
| Outcome: | The proposed model improves on template extraction and text generation tasks on RotoWire datasets. |
Hybrid Inverted Index Is a Robust Accelerator for Dense Retrieval (2023.emnlp-main)
Copied to clipboard
| Challenge: | Inverted file structure is a common technique for accelerating dense retrieval, but its lossy nature degrades it. |
| Approach: | They propose a hybrid index where embedding clusters and salient terms work collaboratively to accelerate dense retrieval. |
| Outcome: | The proposed method achieves lossless retrieval quality with competitive efficiency across index settings. |
LENS: A Learnable Evaluation Metric for Text Simplification (2023.acl-long)
Copied to clipboard
| Challenge: | Existing metrics for text simplification are based on unitary or outdated models, making them unsuitable for this approach. |
| Approach: | They present a learnable evaluation metric for text simplification using language models . they also introduce a human evaluation framework that rates simplifications from several models a list-wise manner . |
| Outcome: | The proposed model correlates much better with human judgment than existing metrics. |
Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text (2022.acl-long)
Copied to clipboard
| Challenge: | a recent study has reported that crowdsourcing cannot distinguish between machine-authored and human-authored text. |
| Approach: | They propose a framework called Scarecrow for scrutinizing machine text via crowd annotation . they use crowd annotation to identify redundancy, commonsense errors, and incoherence . |
| Outcome: | The proposed method quantifies gaps between human-authored and machine-generated text . it can detect redundancy, commonsense errors, and incoherence . |
Thresh: A Unified, Customizable and Deployable Platform for Fine-Grained Text Evaluation (2023.emnlp-demo)
Copied to clipboard
| Challenge: | Existing tools for fine-grained human evaluation lack adaptability to different domains or languages, or modify annotation settings according to user needs. |
| Approach: | They propose a unified platform for fine-grained evaluation that is customizable and deployable with a single YAML configuration file. |
| Outcome: | The proposed frameworks are based on a single YAML configuration file and can be easily extended to different domains or languages. |
Hierarchical Heterogeneous Graph Representation Learning for Short Text Classification (2021.emnlp-main)
Copied to clipboard
| Challenge: | Short text classification is a fundamental task in natural language processing. |
| Approach: | They propose a new method called SHINE which is based on graph neural network for short text classification. |
| Outcome: | The proposed method outperforms state-of-the-art methods on benchmark short text datasets. |
Automatic and Human-AI Interactive Text Generation (with a focus on Text Simplification and Revision) (2024.acl-tutorials)
Copied to clipboard
| Challenge: | In this tutorial, we focus on text-to-text generation, a class of natural language generation tasks, that takes a piece of text as input and then generates a revision that is improved according to some specific criteria. |
| Approach: | This tutorial focuses on text-to-text generation, a class of natural language generation tasks that takes a piece of text as input and generates a revision that is improved according to some specific criteria. |
| Outcome: | This tutorial focuses on text-to-text generation, a class of natural language generation tasks, that takes a piece of text as input and generates a revision that is improved according to some specificcriteria. |
Improving Minimum Bayes Risk Decoding with Multi-Prompt (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to generate LLMs with a single ‘best’ prompt are unstable and sub-optimal in practice. |
| Approach: | They propose to decode multiple candidate generations from a prompt bank at inference-time and use Minimum Bayes Risk (MBR) to select a final output. |
| Outcome: | The proposed method improves MBR across a set of conditional generation tasks and models. |
Reducing Privacy Risks in Online Self-Disclosures with Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Disclosure is a social media activity that can be rewarding but also poses privacy risks. |
| Approach: | They propose to detect and abstract online self-disclosures using a large corpus of 4.8K annotated disclosure spans and a language model to fine-tune for detection. |
| Outcome: | The proposed model can detect and abstract self-disclosures with 80% accuracy, on-par with GPT-3.5. |
GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Using large language models to jailbreak is important for testing safety and security issues. |
| Approach: | They propose an approach that leverages the reflective capabilities of large language models for jailbreaking with only black-box access. |
| Outcome: | The proposed method achieves jailbreak success rates of 98% on GPT-4, 92% on GTP-4 Turbo, and 94% on Llama-3.1-70B in under 7 queries. |