Papers by Yao Dou

15 papers
Simplified Graph Learning for Inductive Short Text Classification (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for short text classification are limited and lack of labeled data is not enough.
Approach: They propose a novel short text classification algorithm which leverages words to handle the lack of labeled data.
Outcome: The proposed model performs better with lower memory consumption and faster inference speed than previous models.
Dancing Between Success and Failure: Edit-level Simplification Evaluation using SALSA (2023.emnlp-main)

Copied to clipboard

Challenge: Traditional human evaluation methods for text simplification often relies on individual, shallow sentence-level ratings, easily affected by the annotator's preference or bias.
Approach: They propose an edit-based human annotation framework that enables holistic and fine-grained text simplification evaluation.
Outcome: The proposed framework is able to predict sentence- and word-level quality simultaneously and report promising results.
Improving Large-scale Paraphrase Acquisition and Generation (2022.emnlp-main)

Copied to clipboard

Challenge: Existing Twitter-based paraphrase datasets lack quality definitions for identification and generation tasks.
Approach: They propose to use two separate definitions of paraphrase for identification and generation tasks in existing Twitter-based paraphrase datasets.
Outcome: The proposed model achieves state-of-the-art performance of 84.2 F1 for automatic paraphrase identification compared to other models fine-tuned on other corpora such as Quora, MSCOCO, and ParaNMT.
Value Compass Benchmarks: A Comprehensive, Generative and Self-Evolving Platform for LLMs’ Value Evaluation (2025.acl-demo)

Copied to clipboard

Challenge: Current evaluation methods for large language models face two key challenges: 1. evaluation validity and 2. Result interpretation reduce the pluralistic and incommensurable values to one-dimensional scores.
Approach: They propose a platform for comprehensive value diagnosis of large language models (LLMs) that provides a generative evaluation paradigm that automatically creates real-world test items co-evolving with ever-advancing LLMs.
Outcome: The proposed platform provides a framework for comprehensive value diagnosis of large language models (LLMs) with fine-grained scores and case studies across 27 value dimensions for 33 leading LLMs, customized comparisons, and visualized analysis of LLM’s alignment with cultural values.
SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants? (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used in interactive applications, and human evaluation remains the gold standard for assessing their performance in multi-turn conversations.
Approach: They propose to use large language models to simulate users for automatic assistant evaluation.
Outcome: The proposed model outperforms human evaluations on two interactive tasks and achieves Spearman’s of 0.7 on both tasks.
Data2Text Studio: Automated Text Generation from Structured Data (D18-2)

Copied to clipboard

Challenge: Data2Text Studio is a platform for automated text generation from structured data.
Approach: They conduct experiments on RotoWire datasets for template extraction and text generation . they find that the Semi-HMMs model improves interactivity and interpretability .
Outcome: The proposed model improves on template extraction and text generation tasks on RotoWire datasets.
Hybrid Inverted Index Is a Robust Accelerator for Dense Retrieval (2023.emnlp-main)

Copied to clipboard

Challenge: Inverted file structure is a common technique for accelerating dense retrieval, but its lossy nature degrades it.
Approach: They propose a hybrid index where embedding clusters and salient terms work collaboratively to accelerate dense retrieval.
Outcome: The proposed method achieves lossless retrieval quality with competitive efficiency across index settings.
LENS: A Learnable Evaluation Metric for Text Simplification (2023.acl-long)

Copied to clipboard

Challenge: Existing metrics for text simplification are based on unitary or outdated models, making them unsuitable for this approach.
Approach: They present a learnable evaluation metric for text simplification using language models . they also introduce a human evaluation framework that rates simplifications from several models a list-wise manner .
Outcome: The proposed model correlates much better with human judgment than existing metrics.
Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text (2022.acl-long)

Copied to clipboard

Challenge: a recent study has reported that crowdsourcing cannot distinguish between machine-authored and human-authored text.
Approach: They propose a framework called Scarecrow for scrutinizing machine text via crowd annotation . they use crowd annotation to identify redundancy, commonsense errors, and incoherence .
Outcome: The proposed method quantifies gaps between human-authored and machine-generated text . it can detect redundancy, commonsense errors, and incoherence .
Thresh: A Unified, Customizable and Deployable Platform for Fine-Grained Text Evaluation (2023.emnlp-demo)

Copied to clipboard

Challenge: Existing tools for fine-grained human evaluation lack adaptability to different domains or languages, or modify annotation settings according to user needs.
Approach: They propose a unified platform for fine-grained evaluation that is customizable and deployable with a single YAML configuration file.
Outcome: The proposed frameworks are based on a single YAML configuration file and can be easily extended to different domains or languages.
Hierarchical Heterogeneous Graph Representation Learning for Short Text Classification (2021.emnlp-main)

Copied to clipboard

Challenge: Short text classification is a fundamental task in natural language processing.
Approach: They propose a new method called SHINE which is based on graph neural network for short text classification.
Outcome: The proposed method outperforms state-of-the-art methods on benchmark short text datasets.
Automatic and Human-AI Interactive Text Generation (with a focus on Text Simplification and Revision) (2024.acl-tutorials)

Copied to clipboard

Challenge: In this tutorial, we focus on text-to-text generation, a class of natural language generation tasks, that takes a piece of text as input and then generates a revision that is improved according to some specific criteria.
Approach: This tutorial focuses on text-to-text generation, a class of natural language generation tasks that takes a piece of text as input and generates a revision that is improved according to some specific criteria.
Outcome: This tutorial focuses on text-to-text generation, a class of natural language generation tasks, that takes a piece of text as input and generates a revision that is improved according to some specificcriteria.
Improving Minimum Bayes Risk Decoding with Multi-Prompt (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods to generate LLMs with a single ‘best’ prompt are unstable and sub-optimal in practice.
Approach: They propose to decode multiple candidate generations from a prompt bank at inference-time and use Minimum Bayes Risk (MBR) to select a final output.
Outcome: The proposed method improves MBR across a set of conditional generation tasks and models.
Reducing Privacy Risks in Online Self-Disclosures with Language Models (2024.acl-long)

Copied to clipboard

Challenge: Disclosure is a social media activity that can be rewarding but also poses privacy risks.
Approach: They propose to detect and abstract online self-disclosures using a large corpus of 4.8K annotated disclosure spans and a language model to fine-tune for detection.
Outcome: The proposed model can detect and abstract self-disclosures with 80% accuracy, on-par with GPT-3.5.
GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation (2024.emnlp-main)

Copied to clipboard

Challenge: Using large language models to jailbreak is important for testing safety and security issues.
Approach: They propose an approach that leverages the reflective capabilities of large language models for jailbreaking with only black-box access.
Outcome: The proposed method achieves jailbreak success rates of 98% on GPT-4, 92% on GTP-4 Turbo, and 94% on Llama-3.1-70B in under 7 queries.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations