Papers by Yang Hou

49 papers
A Probabilistic Framework for LLM Hallucination Detection via Belief Tree Propagation (2025.naacl-long)

Copied to clipboard

Challenge: Current large language models (LLMs) produce factually incorrect statements .
Approach: They propose a probabilistic framework for LLM hallucination detection that generates a belief tree by expanding a statement into logically related claims and reasoning globally about the relationships between these claims.
Outcome: The proposed method improves on multiple hallucination detection benchmarks by 3%-9% over state-of-the-art models.
Dynamic Head Selection for Neural Lexicalized Constituency Parsing (2025.acl-long)

Copied to clipboard

Challenge: Lexicalized parsing has traditionally been neglected in favor of unlexicalized, span-based methods.
Approach: They propose a latent lexicalization framework that dynamically infers lexicals from data without relying on predefined head-finding rules.
Outcome: The proposed model learns lexical dependencies directly from data, offering greater adaptability across languages and datasets.
GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing models for general intelligence fail to model how mental states interact and crystallize into group-level outcomes.
Approach: They propose a multimodal benchmark for group-level Theory of Mind (ToM) to probe nonlinear collective behavior.
Outcome: The proposed model performs significantly below human levels, exposing blind spots in modeling social structures and nonlinear collective behavior.
CLHA: A Simple Yet Effective Contrastive Learning Framework for Human Alignment (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) have attracted considerable attention from academic and industrial communities due to their outstanding performance in various natural language processing tasks.
Approach: They propose a Contrastive Learning Framework for Human Alignment to evaluate the noise within the data and dynamically adjust the training process.
Outcome: The proposed framework surpasses other algorithms in terms of reward model scores, automatic evaluations, and human assessments on the widely used dataset "Helpful and Harmless"
CLERC: A Dataset for U. S. Legal Case Retrieval and Retrieval-Augmented Analysis Generation (2025.findings-naacl)

Copied to clipboard

Challenge: a dataset of case law is used to train and evaluate models for writing legal analyses . current approaches struggle to find relevant cases and generate legal analyses, authors say .
Approach: They build a dataset of case law to support information retrieval and retrieval-augmented generation.
Outcome: The proposed dataset supports two important backbone tasks: retrieval (IR) and retrieval-augmented generation (RAG).
WebDART: Dynamic Decomposition and Re-planning for Complex Web Tasks (2026.findings-acl)

Copied to clipboard

Challenge: Large-language-model (LLM) agents are competent at straightforward web tasks, but struggle with complex tasks.
Approach: They propose a general framework that decomposes web tasks into three subtasks . they show that WebDART lifts end-to-end success rates by 13.7 percentage points .
Outcome: Evaluated on WebChoreArena, WebDART lifts success rates by 13.7 percentage points over previous state-of-the-art agents.
Learn Like Humans: Use Meta-cognitive Reflection for Efficient Self-Improvement (2026.acl-long)

Copied to clipboard

Challenge: Existing self-improving frameworks rely on inefficient, multi-turn recursive loops that incur high computational costs.
Approach: They propose a framework that achieves efficient self-evolution within a single recurrence cycle.
Outcome: The proposed framework outperforms state-of-the-art self-evolving systems while significantly reducing computational overhead.
DALK: Dynamic Co-Augmentation of LLMs and KG to answer Alzheimer’s Disease Questions with Scientific Literature (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in large language models have achieved promising performances across various applications, but the challenge of integrating long-tail knowledge continues to impede the seamless adoption of LLMs in specialized domains.
Approach: They propose a dynamic co-augmentation framework for the refinement of large language models and knowledge graphs in the context of Alzheimer's Disease.
Outcome: The proposed framework can be used to study Alzheimer's Disease (AD) using LLMs and KGs.
A Multi-persona Framework for Argument Quality Assessment (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for argument quality assessment do not consider multi-perspective evaluation due to subjective nature of arguments.
Approach: They propose a multi-persona framework for argument quality assessment that simulates diverse evaluator perspectives through large language models.
Outcome: The proposed framework outperforms baselines while providing comprehensive multi-perspective rationales on IBM-Rank-30k and IBM-ArgQ-5.3kArgs datasets.
CR-GIS: Improving Conversational Recommendation via Goal-aware Interest Sequence Modeling (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to determine a goal item by sequentially tracking users’ interests ignore the rich goal-aware implicit interest sequence patterns in a dialog.
Approach: They propose to model goal-aware implicit user interest sequence patterns in a dialog and a hierarchical Star Transformer to guide multi-turn utterances generation.
Outcome: The proposed framework achieves more accurate recommendations with more fluent and coherent dialog utterances.
Fantastic Questions and Where to Find Them: FairytaleQA – An Authentic Dataset for Narrative Comprehension (2022.acl-long)

Copied to clipboard

Challenge: Existing QA datasets rarely distinguish fine-grained reading skills, such as the understanding of varying narrative elements.
Approach: They propose to use FairytaleQA to generate 10,580 questions based on 278 children-friendly stories to assess model's fine-grained learning skills.
Outcome: The proposed dataset consists of 10,580 questions derived from 278 children-friendly stories, covering seven types of narrative elements or relations.
EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for story evaluation lack reasoning capabilities for open-source models . evolvR framework provides high-fidelity evaluators for story generation tasks .
Approach: They propose a framework that self-synthesizes chain-of-thought data via a multi-persona strategy . they propose evolvR to provide a reward model for story generation .
Outcome: The proposed framework achieves state-of-the-art performance on three evaluation benchmarks . it also enhances the quality of generated stories, validating the superiority of the framework .
Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have focused on test-time scaling to improve reasoning quality but at the cost of efficiency.
Approach: They propose a training-free framework that enhances reasoning accuracy and stability with minimal overhead.
Outcome: The proposed framework yields consistent gains across general, coding, and STEM tasks while remaining highly efficient.
Mining Word Boundaries from Speech-Text Parallel Data for Cross-domain Chinese Word Segmentation (2025.coling-main)

Copied to clipboard

Challenge: Recent studies on Chinese Word Segmentation (CWS) have focused on the cross-domain scenarios, but there is a high cost of manually annotating high-quality data.
Approach: They propose to explicitly mine word boundaries from parallel speech-text data by using the Montreal Forced Aligner toolkit to perform character-level alignment on speech- text data.
Outcome: The proposed approach is based on character-level alignment on speech-text data and a robust complete-then-train (CTT) strategy.
Beyond Static Evaluation: A Dynamic Approach to Assessing AI Assistants’ API Invocation Capabilities (2024.lrec-main)

Copied to clipboard

Challenge: Existing evaluation methods for human-machine interactions are static and can be misleading.
Approach: They propose to use a LLM-based user agent to assess an assistant's API call capability without human involvement.
Outcome: The proposed method mirrors real human conversation patterns in human-machine interactions, and shows that it aligns more closely with human assessment.
Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms (2026.findings-acl)

Copied to clipboard

Challenge: Neologisms and emerging slang are central to daily conversation, yet challenging for non-native speakers (NNS) to interpret and use appropriately in cross-cultural communication with native speakers (NS).
Approach: They use AI to learn English neologisms and write messages using the learned word to an NS friend.
Outcome: The proposed model shows that AI Explanation yields the largest gains over no support in NS-rated competence, while contextual appropriateness judgments show indifference across support.
TopKG: Target-oriented Dialog via Global Planning on Knowledge Graph (2022.coling-1)

Copied to clipboard

Challenge: Existing target-oriented dialogs take a local and greedy strategy for response generation, where global planning is absent.
Approach: They propose a global planning method for target-oriented dialog on a commonsense knowledge graph to adjust local response generation towards the global target.
Outcome: The proposed method can reach the target with a higher success rate, fewer turns, and more coherent responses.
Can Large Language Models Translate Unseen Languages in Underrepresented Scripts? (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated impressive performance in machine translation, but struggle with unseen low-resource languages.
Approach: They propose a benchmark to evaluate translation for Mongolian and Yi using linguistic resources.
Outcome: The proposed model can translate Mongolian (in traditional script) and Yi with the help of linguistic resources, but is limited in its ability to handle these languages effectively.
Can Large Language Models Effectively Support Decision-Making in Sudden Emergencies? (2026.findings-acl)

Copied to clipboard

Challenge: Existing research has focused on the earlier stages of emergency response . lack of suitable datasets for reliable and compliance-aware decision-oriented modeling and evaluation is limiting current research .
Approach: They propose a first real-world emergency decision-making dataset EDM-Bench . they propose 'rule-enhanced reasoning framework' that integrates external regulatory knowledge with constrained inference mechanisms to improve both decision safety and interpretability.
Outcome: The proposed framework improves decision safety and interpretability by integrating regulatory knowledge with constrained inference mechanisms.
VarMAE: Pre-training of Variational Masked Autoencoder for Domain-adaptive Language Understanding (2022.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained language models have been widely applied to standard benchmarks due to the limited resources available in a domain.
Approach: They propose a Transformer-based language model called VarMAE for domain-adaptive language understanding that encodes the context of a token into a smooth latent distribution.
Outcome: Experiments on science- and finance-domain NLU tasks show that the proposed model can be efficiently adapted to new domains with limited resources.
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios (2026.findings-acl)

Copied to clipboard

Challenge: Existing large language models (LLMs) are prone to misuse and misinformation, posing serious compliance risks.
Approach: They propose a bilingual red-teaming benchmark to test an LLM’s refusal of requests that violate financial compliance.
Outcome: The proposed benchmark is based on real-world financial crime cases and ethical violations and includes 14 subcategories covering financial crimes and ethical breaches.
Concise and Precise Context Compression for Tool-Using Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods suffer from key information loss and difficulty in adjusting the length of compressed sequences based on documentation lengths.
Approach: They propose two strategies for compressing tool documentation into concise and precise summary sequences for tool-using language models.
Outcome: The proposed approach achieves comparable performance to the upper-bound baseline under 16x compression ratio.
Triviality Corrected Endogenous Reward (2026.acl-long)

Copied to clipboard

Challenge: Recent work on unsupervised reinforcement learning for mathematical reasoning using confidence-based endogenous rewards focuses on open-ended text generation, requiring either annotated data or powerful closed-source models.
Approach: They propose a method that rewards the relative information gain between a specialist and a generalist reference policy, modulated by a probability-dependent correction mechanism.
Outcome: The proposed model improves on multiple writing benchmarks and model architectures without external supervision and validates generality across different generation tasks.
REG: Retrieval via Emotion Similarity for Guiding Empathetic Dialogue Generation (2026.acl-long)

Copied to clipboard

Challenge: Empathy relies on the cognitive capacity to relate to similar past experiences. Existing methods prioritize semantic similarity over emotion characteristics, leading to unempathetic responses.
Approach: They propose a framework that integrates four Emotion Attributes into the retrieval process to ensure explicit emotional alignment.
Outcome: Empirical results show that REG significantly outperforms baselines, offering a robust solution for empathetic generation.
Advancing the Robustness of Large Language Models through Self-Denoised Smoothing (2024.naacl-short)

Copied to clipboard

Challenge: Existing adversarial attacks can cause LLMs to make wrong predictions on downstream tasks or generate harmful content misaligned with human values.
Approach: They propose to use randomized smoothing to add noise to the input and then make predictions based on these denoised versions.
Outcome: The proposed method surpasses existing methods in both empirical and certified robustness in defending against adversarial perturbations for both downstream tasks and human alignments (i.e., jailbreak attacks).
Multi-perspective Preference Alignment of LLMs for Programming-Community Question Answering (2025.coling-main)

Copied to clipboard

Challenge: Extensive experiments on a high-quality, real-world PCQA dataset validate its accuracy and preference.
Approach: They propose a multi-perspective preference alignment for programming-community question answering to generate user-centric responses.
Outcome: Experiments on a high-quality, real-world PCQA dataset validate the proposed model's accuracy and preference.
Self-Correction Makes LLMs Better Parsers (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have achieved remarkable success across various natural language processing tasks, but they still face challenges in performing fundamental NLP tasks, such as syntactic parsing.
Approach: They propose a method that leverages grammar rules from existing treebanks to guide LLMs in correcting previous errors.
Outcome: The proposed method significantly improves performance on in-domain and cross-domain datasets.
Can MLLMs Understand the Deep Implication Behind Chinese Images? (2025.acl-long)

Copied to clipboard

Challenge: MLLMs perform poorly on traditional culture images, indicating limitations in understanding high-level semantics and lacking a deep knowledge base of Chinese traditional culture.
Approach: They propose to use Chinese images to assess MLLMs' higher-order perception and understanding of Chinese visual content.
Outcome: The proposed model incorporates images that represent Chinese traditional culture, such as famous Chinese traditional paintings, to ensure the authenticity of the Chinese context.
Good Reasoning Makes Good Demonstrations: Implicit Reasoning Quality Supervision via In-Context Reinforcement Learning (2026.findings-acl)

Copied to clipboard

Challenge: Reinforcement Learning with Verifiable Rewards (RLVR) improves reasoning in large language models but treats all correct solutions equally, potentially reinforcing flawed traces that arrive at correct answers by chance.
Approach: They propose a method that reweights rewards by a factor approximately proportional to Evidence Gain and assigns higher weights to high-quality traces without requiring costly computation.
Outcome: Experiments on mathematical reasoning benchmarks show that Reinforcement Learning with Verifiable Rewards (RLVR) improves reasoning in large language models but treats all correct solutions equally.
PASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual Learning (2026.acl-long)

Copied to clipboard

Challenge: Existing LoRA-based Mixture-of-Experts (MoE) methods often jointly update the router and experts in an indiscriminate way, causing the router’s preferences to co-drift with experts’ adaptation pathways and exacerbate forgetting.
Approach: They propose a LoRA-induced subspace that reflects which low-rank pathway directions an input activates in each expert, providing a capability-aligned coordinate system for routing and preservation.
Outcome: The proposed method outperforms conventional continual learning baselines and MoE–LoRA variants in accuracy and resistance to forgetting, without increasing model parameters.
Alignment with Fill-In-the-Middle for Enhancing Code Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for generating test cases with limited training data are not reliable and may be counterproductive.
Approach: They propose a method that splits code snippets into smaller, granular blocks, creating more diverse DPO pairs from the same test cases.
Outcome: The proposed approach shows significant improvements in code generation tasks on benchmark datasets such as HumanEval (+), MBPP (+), and APPS.
E-EVAL: A Comprehensive Chinese K-12 Education Evaluation Benchmark for Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: despite the rapid development of Large Language Models, there is no dedicated benchmark for evaluating LLMs in Chinese K-12 education.
Approach: They propose to develop a benchmark specifically tailored for Chinese K-12 education.
Outcome: EVAL is the first evaluation benchmark specifically tailored for Chinese K-12 education.
Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Spreadsheets are among the most widely used data formats in real-world applications . existing large language models treat tables as plain text, overlooking layout cues and visual semantics.
Approach: They propose a two-stage multi-agent framework for spreadsheet understanding that adopts a step-by-step reading and reasoning paradigm.
Outcome: Extensive experiments on two spreadsheet datasets show the proposed framework outperforms existing methods on Spreadsheet Bench.
FinMRAGBench: A Realistic and Complex Benchmark for Multi-Modal RAG in Financial Document Analysis (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for realistic financial analysis fail to capture realistic financial situations involving cross-document retrieval, multi-page evidence integration, and diverse analytical tasks.
Approach: They propose a multi-modal financial RAG benchmark that evaluates large language models in realistic financial analysis settings.
Outcome: The proposed framework achieves the strongest overall performance across all models.
High-order Joint Constituency and Dependency Parsing (2024.lrec-main)

Copied to clipboard

Challenge: Syntactic parsing aims to reveal how sentences are syntactically structured.
Approach: They propose to produce compatible constituency and dependency trees simultaneously for input sentences . they adopt a much more efficient decoding algorithm and explore joint modeling at training phase .
Outcome: The proposed model significantly improves matching ratio of whole trees compared to separate models . the proposed model adopts a much more efficient decoding algorithm .
METNet: A Mutual Enhanced Transformation Network for Aspect-based Sentiment Analysis (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for learning complex sentences with multiple aspects are ill-equipped to learn complex sentences .
Approach: They propose a mutual enhanced transformation network for the ABSA task . it improves representation learning of the aspect with contextual semantic features .
Outcome: The proposed model improves representation learning of the aspect with contextual semantic features, giving the aspect more abundant information.
GenomeQA: Benchmarking General Large Language Models for Genome Sequence Understanding (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on specialized DNA models trained for sequence prediction or evaluate biological knowledge using text-only questions.
Approach: They propose a benchmark to evaluate general-purpose LLMs on sequence-based genome inference tasks.
Outcome: The proposed benchmark outperforms baseline models on sequence-based genome inference tasks.
Evaluating the Expressive Appropriateness of Speech in Rich Contexts (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for evaluating expressive speech focus on word accuracy, naturalness, signal quality, or emotional intensity at the utterance level.
Approach: They propose a framework for Evaluating Expressive Appropriateness in speech that assesses whether a speech sample aligns with the underlying communicative intent implied by its discourse-level narrative context.
Outcome: The proposed framework outperforms existing speech evaluation and analysis systems on a human-annotated test set.
Data Augmentation for Cross-domain Parsing via Lightweight LLM Generation and Tree Hybridization (2025.coling-main)

Copied to clipboard

Challenge: Existing approaches for constituency parsing are expensive and lack high-quality labeled data.
Approach: They propose a data augmentation method via lightweight large language model (LLM) generation and tree hybridization to generate a large number of structurally diverse instances.
Outcome: The proposed method achieves significant improvements on five target domains with a lightweight LLM generation cost.
Character-Level Chinese Dependency Parsing via Modeling Latent Intra-Word Structure (2024.findings-acl)

Copied to clipboard

Challenge: Existing word-level dependency parsing methods in Chinese lack explicit word boundaries due to the lack of word boundaries.
Approach: They propose to model latent internal structures within Chinese words by constrained Eisner algorithm . they propose to guarantee a single root for intra-word structures and establish inter-word dependencies .
Outcome: The proposed model outperforms existing models on Chinese treebanks and shows that it can predict plausible intra-word structures.
A Probabilistic Toolkit for Multi-grained Word Segmentation in Chinese (2025.coling-demos)

Copied to clipboard

Challenge: Existing tools for word segmentation are based on different linguistic theories or target different scenarios.
Approach: They propose a probabilistic toolkit for multi-grained word segmentation in Chinese . they adopt semi-Markov CRF for single-grain word segmenting (SWS) .
Outcome: The proposed approach can produce marginal probabilities of words during inference and significantly improve performance in the cross-domain scenario.
From Knowing to Teaching: Scaffolding Pedagogical Decisions for LLM Agent (2026.acl-long)

Copied to clipboard

Challenge: Large language models produce content lacking pedagogical depth when asked to generate lessons .
Approach: They propose a framework that allows teachers to select content according to pedagogical intent and sequence topics so foundations precede applications.
Outcome: The framework achieves 67.8% win rate in human evaluation and 79.6% in LLM-based evaluation against eight baselines.
CharacterGLM: Customizing Social Characters with Large Language Models (2024.emnlp-industry)

Copied to clipboard

Challenge: Character-based dialogue systems (CharacterDial) allow users to customize social characters for social interactions.
Approach: They will collect a large-scale Chinese corpus of characters with diverse categories and behaviors and develop CharacterGLM models to address these challenges.
Outcome: Experiments show that CharacterGLM outperforms most popular open- and closed-source LLMs and performs comparable to GPT-4.
Language Anisotropic Cross-Lingual Model Editing (2023.findings-acl)

Copied to clipboard

Challenge: Existing work studies monolingual model editing, which lacks cross-lingual transferability to perform editing simultaneously across languages.
Approach: They propose a framework to naturally adapt monolingual model editing approaches to the cross-lingual scenario using parallel corpus.
Outcome: The proposed framework adapts monolingual model editing approaches to the cross-lingual scenario using parallel corpus and amplifies different subsets of parameters for each language.
Model-based Large Language Model Customization as Service (2025.emnlp-main)

Copied to clipboard

Challenge: Existing large language model services require users to upload data for fine-tuning . current methods for customization are noisy and require sensitive domain data .
Approach: *Llamdex is a framework that facilitates LLM customization as a service . client uploads pre-trained domain-specific *models* rather than data .
Outcome: *Llamdex* framework improves domain-specific accuracy by up to 26% over state-of-the-art private data synthesis methods .
Weight-Inherited Distillation for Task-Agnostic BERT Compression (2024.findings-naacl)

Copied to clipboard

Challenge: Knowledge Distillation (KD) is a predominant approach for BERT compression.
Approach: They propose a weight-inherited distillation method which directly transfers knowledge from the teacher to a compact student model by inheriting the weights.
Outcome: The proposed method outperforms state-of-the-art KD-based methods on GLUE and SQUAD benchmarks.
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval (2025.acl-long)

Copied to clipboard

Challenge: Document retrieval in real-world scenarios faces significant challenges due to diverse document formats and modalities.
Approach: They propose a visual-textual embedding framework that integrates textual and visual features for robust document representation.
Outcome: The proposed visual-textual embedding framework surpasses existing methods while preserving semantic fidelity.
Mining Effective Features Using Quantum Entropy for Humor Recognition (2023.findings-eacl)

Copied to clipboard

Challenge: Existing studies on humor recognition do not understand the mechanisms that generate humor.
Approach: They propose to use quantum entropy to represent the semantic uncertainty of the setup and punchline as features for humor recognition.
Outcome: The proposed features are more effective than baselines for recognizing humorous and non-humorous texts on the SemEval2021 task 7 dataset.
Span-based Semantic Role Labeling as Lexicalized Constituency Tree Parsing (2025.findings-acl)

Copied to clipboard

Challenge: Existing models for semantic role labeling fail to capture the relationship between syntax and semantics.
Approach: They propose a lexicalized tree representation for span-based SRL that integrates constituency and dependency parsing to explicitly model predicate-argument structures.
Outcome: The proposed model achieves competitive performance on standard English benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations