Papers by Linhao Yu

11 papers
Optimizing Question Semantic Space for Dynamic Retrieval-Augmented Multi-hop Question Answering (2025.acl-long)

Copied to clipboard

Challenge: Existing retrieval-augmented generation methods are insufficient for multi-hop question answering . however, they tend to generate hallucinations due to semantic mismatching .
Approach: They propose to optimize question semantic space for dynamic retrieval-augmented multi-hop question answering by optimizing the semantic embeddings.
Outcome: The proposed method outperforms existing RAG methods in both in- and out-of-domain settings.
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Recent years have witnessed remarkable progress achieved by large language models in both natural language understanding and generation.
Approach: They propose a large benchmark CMoralEval for moral evaluation of Chinese LLMs . they use a Chinese TV program discussing Chinese moral norms and Chinese moral anomies based on various sources .
Outcome: The proposed dataset is characterized by diversity and authenticity.
Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluations of Large Language Models (LLMs) focus on item-level behavioral metrics without capturing how models prioritize competing values as a whole.
Approach: They propose a symmetric human-LLM evaluation framework to measure value-structure alignment . they evaluate 12 LLMs across four model families via 240 replicated Q-sorts .
Outcome: The proposed framework measures value-structure alignment across four model families.
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety (2024.acl-demos)

Copied to clipboard

Challenge: a rapid development of Chinese large language models poses big challenges for efficient LLM evaluation.
Approach: They propose an evaluation testbed that benchmarks Chinese LLMs across capability, alignment and safety.
Outcome: The evaluation platform OpenEval benchmarks Chinese LLMs across capability, alignment and safety.
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks and evaluation protocols suffer from inadequate or homogeneous creation of key points, exorbitant cost of data creation, and limited evaluation scopes.
Approach: They propose an automatic framework which leverages Monte Carlo Tree Search to construct numerous and diverse descriptive sentences that thoroughly represent video content in an iterative way.
Outcome: The proposed framework improves MCTS-VCB and DREAM-1K on video captioning tasks by 25.0% and 16.3% respectively.
AttnPO: Attention-Guided Process Supervision for Efficient Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing trajectory-level length penalties fail to effectively shorten reasoning length and degrade accuracy, as they treat all reasoning steps uniformly and lack fine-grained signals to distinguish redundancy from necessity.
Approach: They propose a low-overhead process-supervised RL framework that leverages the model’s intrinsic attention signals for step-level credit assignment.
Outcome: The proposed framework reduces reasoning length while improving performance across 9 benchmarks.
Self-Pluralising Culture Alignment for Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Existing approaches to align large language models don't take cultural diversity into account.
Approach: They propose a framework that generates questions on various culture topics and outputs to LLMs under both culture-aware and culture-unaware settings.
Outcome: The proposed framework improves the alignment of large language models to diverse cultures without compromising general abilities.
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for video understanding often focus on specific aspects, overlooking the holistic nature of video content.
Approach: They propose a temporal-oriented benchmark for fine-grained understanding on dense dynamic videos with two complementary tasks: captioning and QA.
Outcome: The proposed model performs well on diverse video scenarios and dynamic videos, with interpretable and robust evaluation criteria.
Rethinking the Reversal Curse of LLMs: a Prescription from Human Knowledge Reversal (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for large language models (LLMs) are limited by their aggressive sample permutation and lack a detailed understanding of the underlying reasons for the reversal curse.
Approach: They propose a method which enhances bidirectional entity correlation modeling and pairwise relationship reasoning to overcome the reversal curse.
Outcome: The proposed method overcomes the reversal curse by augmenting the samples with entity order-reversals and semantically preserved question-answer pairs.
CS2W: A Chinese Spoken-to-Written Style Conversion Dataset with Multiple Conversion Types (2023.emnlp-main)

Copied to clipboard

Challenge: Existing datasets focus on a single type of spoken style, such as disfluencies.
Approach: They propose a Chinese Spoken-to-Written style conversion dataset with 7,237 spoken sentences extracted from transcribed conversational texts.
Outcome: The proposed dataset covers four major conversion problems corresponding to the majority of spoken styles.
LFED: A Literary Fiction Evaluation Dataset for Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: LFED is a literary fiction evaluation dataset for large language models that evaluate the capability of LLMs on the long fiction comprehension and reasoning.
Approach: They propose a Literary Fiction Evaluation Dataset to evaluate LLMs' comprehension and reasoning on long fictions.
Outcome: The proposed dataset evaluates the capability of large language models on the long fiction comprehension and reasoning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations