Papers by Ying Su

19 papers
CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error Correction (2025.acl-long)

Copied to clipboard

Challenge: Existing studies have focused on the interpretability of Grammatical Error Correction (GEC) evaluation metrics, but the interpretabilty of these metrics has been neglected.
Approach: They propose a reference-based metric that describes four aspects of GEC systems: hit-correction, wrong-corrections, under-correcties, and over-corrects.
Outcome: The proposed metric reveals critical qualities and locates drawbacks of GEC systems.
LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing (2024.emnlp-main)

Copied to clipboard

Challenge: a comparative analysis of paper (meta-)reviews by large language models (LLMs) aims to identify and distinguish LLMs from human activities .
Approach: They present a comparative analysis to identify and distinguish LLM activities from human activities.
Outcome: The proposed analysis aims to improve recognition of instances when someone implicitly uses LLMs for reviewing activities.
ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have been adopted to process textual task description and accomplish procedural planning in embodied AI tasks because of their powerful reasoning ability.
Approach: They propose to evaluate the planning ability of large language models and multi-modal counterfactual vision language models (VLMs) using a multi-factual household activity simulator and a chatGPT task description to evaluate their reasoning ability.
Outcome: The proposed benchmark evaluates the planning ability of multi-modal and counterfactual vision language models on a household activity simulator and a chatGPT task description.
SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are being used in urban planning but there is concern that they reproduce or amplify such biases.
Approach: They propose a framework to evaluate spatial gender bias in large language models . they use a taxonomy of 62 urban micro-spaces, a prompt library and three diagnostic layers .
Outcome: The proposed framework identifies structured gender-space associations that go beyond the public-private divide, forming nuanced micro-level mappings.
MovieCORE: COgnitive REasoning in Movies (2025.emnlp-main)

Copied to clipboard

Challenge: MovieCORE is a video question answering dataset that focuses on surface-level comprehension.
Approach: They propose a video question-answer dataset that uses large language models as thought agents to generate and refine high-quality question-anchor pairs.
Outcome: The proposed model improves model reasoning capabilities post-training by 25% . the proposed model is based on a large language model and is scalable to a wide range of tasks .
Refine, Align, and Aggregate: Multi-view Linguistic Features Enhancement for Aspect Sentiment Triplet Extraction (2024.findings-acl)

Copied to clipboard

Challenge: Aspect Sentiment Triplet Extraction (ASTE) aims to extract the triplets of aspect terms, their associated sentiment and opinion terms.
Approach: They propose to use multi-view linguistic features enhancement to explore the prior indication effect in the “Refine, Align, and Aggregate” learning process to enhance aspect-opinion relations.
Outcome: The proposed model achieves state-of-the-art on several benchmark datasets and is robust to state- of-the art constraints.
PipeNet: Question Answering with Semantic Pruning over Knowledge Graphs (2024.starsem-1)

Copied to clipboard

Challenge: Existing approaches to utilizing explicit knowledge graphs (KGs) are limited by the number of nodes in the subgraph.
Approach: They propose a grounding-pruning-reasoning pipeline to prune noisy nodes in subgraphs to improve the efficiency of graph reasoning with KG.
Outcome: The proposed method reduces computation cost and memory usage while obtaining decent representation of pruned subgraphs.
Unified Grid Tagging Scheme for Aspect Sentiment Quad Prediction (2025.coling-main)

Copied to clipboard

Challenge: Existing table-filling methods decompose the ASQP task into subtasks without considering the association between sentiment elements.
Approach: They propose a simple yet effective Unified Grid Tagging Scheme to extract sentiment quadruplets in one shot . they leverage syntactic dependency tree and AMR graph to enrich association between sentiment elements .
Outcome: The proposed model extracts all sentiment elements in quads for a given review to explain the reason for the sentiment.
Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key? (2024.acl-long)

Copied to clipboard

Challenge: Recent progress in LLMs discussion suggests that multi-agent discussion improves the reasoning abilities of LLM.
Approach: They propose a group discussion framework to enrich the set of discussion mechanisms.
Outcome: The proposed framework performs better on a wide range of reasoning tasks and backbone LLMs.
Memory-enhanced Large Language Model for Cross-lingual Dependency Parsing via Deep Hierarchical Syntax Understanding (2025.findings-emnlp)

Copied to clipboard

Challenge: Experimental results show that our approach can significantly improve the parsing accuracy of all baseline models, leading to new state-of-the-art results.
Approach: They propose a deep hierarchical syntax understanding approach to improve the cross-lingual semantic memory capability of large language models by implicitly aligning linguistic knowledge between source and target languages.
Outcome: The proposed approach improves the cross-lingual semantic memory capability of large language models by combining implicit multi-task fine-tuning and explicit label bank guiding.
Rare and Zero-shot Word Sense Disambiguation using Z-Reweighting (2022.acl-long)

Copied to clipboard

Challenge: Word sense disambiguation (WSD) is a problem in the natural language processing community.
Approach: They propose a method to adjust training on imbalanced word sense dataset . they propose to achieve performance gain on standard English all words benchmark .
Outcome: The proposed method achieves performance gain on the standard English all words benchmark.
A Unified Generative Framework for Bilingual Euphemism Detection and Identification (2024.findings-acl)

Copied to clipboard

Challenge: Existing euphemism datasets are only domain-specific or language-specific.
Approach: They propose a unified model to jointly conduct bilingual euphemism detection and identification tasks.
Outcome: The proposed model is effective and provides a new reference standard for euphemism detection and identification.
Refining Idioms Semantics Comprehension via Contrastive Learning and Cross-Attention (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods based on deep learning struggle to grasp idiom semantics due to the figurative meanings of many idiomas deviating from their literal interpretations.
Approach: They propose a Chinese idiom cloze test to capture comprehensive idiomatics and a semantic sense contrastive learning module to enhance the representation of idiomics.
Outcome: The proposed model outperforms state-of-the-art models on the Chinese idiom cloze test and on other benchmark datasets.
MICO: A Multi-alternative Contrastive Learning Framework for Commonsense Knowledge Representation (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to commonsense reasoning include fine-tuning large pre-trained language models or injecting the entire knowledge base for CKGC.
Approach: They propose to learn commonsense knowledge representation by using a multi-alternative contrastive learning framework on COmmonsense Knowledge graphs.
Outcome: Extensive experiments show that the proposed framework is effective in commonsense reasoning tasks.
TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for visual storytelling ignore latent topic information.
Approach: They propose a topic-aware reinforcement network for VIsual StoryTelling that takes topic information into account to generate a coherent story.
Outcome: The proposed method outperforms most of the competing models across multiple evaluation metrics.
Multilingual Word Sense Disambiguation with Unified Sense Representation (2022.coling-1)

Copied to clipboard

Challenge: Existing researches on word sense disambiguation focus on English only.
Approach: They propose to build knowledge and supervised based multilingual word sense disambiguation systems on a multilingual lexicon describing the same set of concepts across languages.
Outcome: The proposed model can understand the fine-grained semantics of words under specific contexts.
SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science (2025.acl-long)

Copied to clipboard

Challenge: Seed science is essential for modern agriculture, but its application in seed science remains limited due to a shortage of experts and limited availability of online resources.
Approach: They evaluate 26 leading large language models and compare them against a set of benchmarks . they find that there is a gap between the power of LLMs and real-world seed science problems .
Outcome: The new seed benchmark highlights the gap between the power of large language models and real-world seed science problems.
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset (2025.acl-long)

Copied to clipboard

Challenge: Recent Common Crawl datasets remove 90% of data, limiting their suitability for long token horizon training.
Approach: They propose to combine classifier ensembling, synthetic data rephrasing and heuristic filters to achieve better trade-offs between accuracy and data quantity.
Outcome: The proposed model-based filtering improves MMLU by 5.6 over DCLM for 15T tokens . the full 6.3T token dataset matches DCLM on MMLO, but contains four times more unique real tokens than DCLM .
Jailbreaking Large Language Models with Morality Attacks (2026.findings-acl)

Copied to clipboard

Challenge: Pluralism alignment is the goal of creating AI that can coexist with and serve morally multifaceted humanity.
Approach: They propose to use jailbreak attacks to manipulate LLMs’ judgment over pluralistic values by using a morality dataset with 10.4K instances.
Outcome: The proposed method exploits the persuasion abilities of LLMs to produce moral content over pluralistic values.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations