Papers by Yifan Hou

14 papers
Explore the Reasoning Capability of LLMs in the Chess Testbed (2025.naacl-short)

Copied to clipboard

Challenge: a recent study shows that large language models struggle with long-term, complex reasoning tasks.
Approach: They propose to integrate annotated strategy and tactic into large language models to improve reasoning capability.
Outcome: The proposed model performs better than GPT, Claude, and Gemini models . it integrates annotated strategy and tactic into the model .
A Survey of Post-Training Scaling in Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated proficiency in understanding and generating human natural languages.
Approach: They propose a framework for scaling large language models using supervised fine-tuning, RLxF and test-time compute methodologies.
Outcome: The proposed model can be used to understand and generate human natural languages.
Adapters for Enhanced Modeling of Multilingual Knowledge and Text (2022.findings-emnlp)

Copied to clipboard

Challenge: Large language models learn facts from text corpora, but knowledge graphs contain facts in an explicit triple format, restricting their research and application.
Approach: They propose to enhance multilingual language models with knowledge from multilingual knowledge graphs . they propose to use cross-lingual entity alignment and facts from MLKGs to improve performance .
Outcome: The proposed model improves MLLMs with cross-lingual entity alignment and facts from multilingual knowledge graphs for many languages while maintaining performance on other general language tasks.
ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have shown excellent mastering of human language but struggle in real-world applications that require mathematical problem-solving.
Approach: They propose a pipeline to train a general Math-Critique model from the LLM itself to provide feedback signals and employ rejective fine-tuning and direct preference optimization over the Llm's own generations for data collection.
Outcome: The proposed pipeline outperforms existing LLMs that could be two times larger.
Bird’s Eye: Probing for Linguistic Graph Structures with a Simple Information-Theoretic Approach (2021.acl-long)

Copied to clipboard

Challenge: Recent work on analyzing contextualized text representations has focused on hand-designed probe models to understand how and to what extent do these representations encode a particular linguistic phenomenon.
Approach: They propose a new information-theoretic probe, Bird’s Eye, which detects if and how representations encode the information in contextualized text representations.
Outcome: The proposed method estimates the mutual information between the linguistic graph embedded in a continuous space and the contextualized word representations.
What Has Been Enhanced in my Knowledge-Enhanced Language Model? (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing knowledge integration methods such as linear probes and prompts have key limitations in answering these questions.
Approach: They propose a new probe model which integrates external knowledge from knowledge graphs into pretrained language models (LMs) ERNIE and K-Adapter are proposed as KI methods .
Outcome: The proposed model interprets two well-known KELMs using graph attention on the corresponding knowledge graph for interpretation.
Evaluating the Expressive Appropriateness of Speech in Rich Contexts (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for evaluating expressive speech focus on word accuracy, naturalness, signal quality, or emotional intensity at the utterance level.
Approach: They propose a framework for Evaluating Expressive Appropriateness in speech that assesses whether a speech sample aligns with the underlying communicative intent implied by its discourse-level narrative context.
Outcome: The proposed framework outperforms existing speech evaluation and analysis systems on a human-annotated test set.
What Do Language Models Learn in Context? The Structured Task Hypothesis. (2024.acl-long)

Copied to clipboard

Challenge: Pre-trained large language models have exhibited an impressive ability to learn in context across various domains, e.g., code generation, education, medicine and even medicine.
Approach: They taxonomize existing candidate theories into three competing hypotheses that explain LLMs’ ability to learn in context.
Outcome: The proposed model can learn a task from in-context examples presented in a demonstration and generalize it to the prompt.
Can Vision-Language Models Solve Visual Math Equations? (2025.emnlp-main)

Copied to clipboard

Challenge: Vision-Language Models (VLMs) perform well on textual equations, but fail on visually grounded counterparts.
Approach: They propose to decompose visual equation solving into symbolic equation solving and visual recognition into two core components to understand this gap.
Outcome: The proposed models perform well on textual equations, but fail on visual grounded ones.
HealthCards: Exploring Text-to-Image Generation as Visual Aids for Healthcare Knowledge Democratizing and Education (2025.emnlp-main)

Copied to clipboard

Challenge: Text-to-image (T2I) generation has the potential to advance knowledge democratization and education.
Approach: They explore ways to harness T2I models for generating health knowledge flashcards . they curated a high-quality healthcare knowledge flash card dataset .
Outcome: The proposed models can generate health knowledge flashcards with appealing images . the results show that the open-source models can be fine tuned to generate health content .
Mitigating Label Biases for In-context Learning (2023.acl-long)

Copied to clipboard

Challenge: Existing methods to categorize label biases in in-context learning (ICL) have not addressed all three types of label bias.
Approach: They propose a method that estimates a language model’s label bias using random in-domain words from the task corpus to categorize and detect label biases in ICL.
Outcome: The proposed method significantly improves the performance of GPT-J and GPT-3 on a wide range of tasks.
Unveiling the Art of Heading Design: A Harmonious Blend of Summarization, Neology, and Algorithm (2024.findings-acl)

Copied to clipboard

Challenge: Creating an appealing heading is crucial for attracting readers and marketing work or products.
Approach: They propose a benchmark to measure the quality of heading generation using summarization, neology, and algorithm metrics.
Outcome: The proposed benchmark compared 6,653 abstracts with corresponding descriptions and acronyms and found that it excels across summarization, neology, and algorithm aspects.
Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Recent work has shown that language models (LMs) have strong multi-step (i.e., procedural) reasoning capabilities.
Approach: They propose a mechanistic interpretation of language models for multi-step reasoning tasks by introducing a new probing approach that recovers the reasoning tree from the model’s attention patterns.
Outcome: The proposed model implicitly embeds a reasoning tree resembling the correct reasoning process within it, and detects the information from the model’s attention patterns for most examples.
A Cause-Effect Look at Alleviating Hallucination of Knowledge-grounded Dialogue Generation (2024.lrec-main)

Copied to clipboard

Challenge: Existing dialogue systems have demonstrated impressive performance conducting fluent and natural-sounding conversations, but they are plagued by the Knowledge Hallucination problem.
Approach: They propose a method that exploits the dialogue-knowledge interaction to reduce hallucination by using external knowledge resources to generate more informative responses.
Outcome: The proposed method reduces hallucination without disrupting other dialogue performance while keeping adaptive to different generation models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations