Papers by Yansong Feng

67 papers
Exploring Question-Specific Rewards for Generating Deep Questions (2020.coling-main)

Copied to clipboard

Challenge: Recent question generation approaches use the sequence-to-sequence framework to optimize the log likelihood of ground-truth questions using teacher forcing.
Approach: They propose to optimize for QG-specific objectives via reinforcement learning to improve question quality.
Outcome: The proposed model improves the fluency, relevance, and answerability of generated questions.
Cross-Lingual Question Answering over Knowledge Base as Reading Comprehension (2023.findings-eacl)

Copied to clipboard

Challenge: Existing high-quality xMRC datasets can be further utilized to fine-tune our model.
Approach: They propose a cross-lingual question answering over knowledge base approach that converts KB subgraphs into passages to narrow the gap between KB schemas and questions.
Outcome: The proposed approach outperforms baselines and achieves strong few-shot and zero-shot performance on two xKBQA datasets in 12 languages.
Dual-Channel Evidence Fusion for Fact Verification over Texts and Tables (2022.naacl-main)

Copied to clipboard

Challenge: Existing fact extraction and verification tasks only consider evidence of a single format . Existing models convert evidence into either sentences or tables, thus losing context information .
Approach: They propose a Dual Channel Unified Format fact verification model which unifies various evidence into parallel streams, i.e., natural language sentences and a global evidence table, simultaneously.
Outcome: The proposed model outperforms existing models in two formats by a large margin . it makes the most of existing tables and tables to absorb evidence of two formats .
Natural Answer Generation with Heterogeneous Memory (N18-1)

Copied to clipboard

Challenge: Recent work on memory augmented encoder-decoder frameworks has shown promising progress for natural language generation tasks.
Approach: They propose a memory-augmented encoder-decoder framework that takes care of memory contents from different sources to explicitly avoid repetition.
Outcome: The proposed approach can produce readable and meaningful answer sentences while maintaining high coverage for given answer information.
Chain-of-Discussion: A Multi-Model Framework for Complex Evidence-Based Question Answering (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable language generation capabilities, propelling advancements in various understanding/generation tasks, including opendomain question answering (QA).
Approach: They propose a chain-of- Discussion framework to leverage synergy among multiple open-source Large Language Models (LLMs) aiming to provide more correct and more comprehensive answers for open-ended QA, although they are not strong enough individually.
Outcome: The proposed framework leverages the synergy among multiple open-source Large Language Models (LLMs) to provide more correct and comprehensive answers for open-ended QA, although they are not strong enough individually.
Towards Context-Aware Code Comment Generation (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for code comments generate comments manually, but they suffer from poor scalability and high maintenance cost due to the expensive overhead of writing comment templates.
Approach: They propose a method to automatically generate code comments at a function level by targeting object-oriented programming languages.
Outcome: The proposed approach outperforms the state-of-the-art methods and is comparable with existing methods.
Probing Multimodal Large Language Models for Global and Local Semantic Representations (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have focused on the ability of MLLMs to generate single tokens one by one, while lacking studies about how their representation vectors can encode global multimodal information.
Approach: They propose to use image-caption corpus to train Multimodal Large Language Models (MLLMs) . they find that the topmost layers encode more global semantic information .
Outcome: The proposed models can encode more global semantic information, rather than the topmost layers, and perform better on visual-language entailment tasks.
Understanding Procedural Text using Interactive Entity Networks (2020.emnlp-main)

Copied to clipboard

Challenge: Recent efforts to track multiple entities in a procedural text treat each entity separately . e.g., scientific articles, instruction books, recipes, often contain multiple entities involved .
Approach: They propose a recurrent network with memory equipped cells for state tracking . they maintain different attention matrices through specific memories to model different types of entity interactions .
Outcome: The proposed model outperforms state-of-the-art models on a benchmark dataset.
EpiCoDe: Boosting Model Performance Beyond Training with Extrapolation and Contrastive Decoding (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods to enhance performance of Large language models are limited due to the cost of training data and privacy concerns.
Approach: They propose a method that enhances a finetuned model with its inferior version and adopts contrastive decoding to reduce predicted errors.
Outcome: The proposed method outperforms existing methods in data-scarcity scenarios across three domains and shows that it is more robust and robust.
Align-then-Enhance: Multilingual Entailment Graph Enhancement with Soft Predicate Alignment (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to learn typed entailment graphs with predicates as nodes and enttailment relations as edges are incomplete.
Approach: They propose a task to utilize entailment information from one EG to enhance another in a different language.
Outcome: The proposed framework outperforms existing graphs in multilingual entailment graph enhancement tasks.
Unlocking the Potential of Model Merging for Low-Resource Languages (2024.findings-emnlp)

Copied to clipboard

Challenge: Adapting large language models (LLMs) to new languages requires continual pre-training followed by supervised fine-tuning.
Approach: They propose a model merging solution that integrates LLMs with distinct capabilities into a single model without additional training.
Outcome: The proposed model merging outperforms CT-then-SFT in low-resource languages with scarce data.
More than Classification: A Unified Framework for Event Temporal Relation Extraction (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for event temporal relation extraction ignore meaning of relations and wipe out their intrinsic dependency.
Approach: They propose a unified event temporal relation extraction framework that transforms temporal relations into logical expressions of time points and completes the ETRE by predicting the relations between certain time points.
Outcome: The proposed framework outperforms the state-of-the-art model on TB-Dense and MATRES by 0.3% on both datasets.
ELLA: Empowering LLMs for Interpretable, Accurate and Informative Legal Advice (2024.acl-demos)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown impressive performance in various tasks, showing great potential for specific domains, such as law (Lai et al., 2023), finance (Zeng e e al. 2023) and law (Lam elms, 2024).
Approach: They propose to use large language models to provide interpretable, accurate, and informative legal advice by visually presenting the correlation between legal articles and LLM's response by calculating their similarities.
Outcome: The proposed model provides users with an intuitive legal basis for the responses and retrieves relevant legal cases for user reference.
SQL-to-Text Generation with Graph-to-Sequence Model (D18-1)

Copied to clipboard

Challenge: Existing approaches to generate SQL-to-text using seq2seq models do not capture graph-structured information in SQL query.
Approach: They propose a graph-to-sequence model to encode global structure information into node embeddings.
Outcome: The proposed model outperforms the Seq2Seq and Tree2Sq baselines on the WikiSQL and Stackoverflow datasets.
How Many Answers Should I Give? An Empirical Study of Multi-Answer Reading Comprehension (2023.findings-acl)

Copied to clipboard

Challenge: Despite recent progress in multi-answer MRC, there is no systematic analysis of how this phenomenon arises and how to better address it.
Approach: They develop a taxonomy to categorize commonly-seen multi-answer MRC instances and examine how well different paradigms deal with different types of multi-announced questions.
Outcome: The proposed taxonomy categorizes commonly-seen multi-answer instances and analyzes how well different paradigms deal with different types of multi-announced instances.
UnifEE: Unified Evidence Extraction for Fact Verification (2023.eacl-main)

Copied to clipboard

Challenge: Existing models extract evidence in both sentences and table cells from Wikipedia dumps, ignoring potential connections between them.
Approach: They propose a model which uses a mixed evidence graph to extract the evidence in both formats without manually designed conversion rules.
Outcome: The proposed model outperforms existing models and improves the verification step.
Automating Legal Interpretation with LLMs: Retrieval, Generation, and Evaluation (2025.acl-long)

Copied to clipboard

Challenge: a novel framework for automated legal interpretation is proposed to alleviate the burden on legal experts.
Approach: They propose a framework for automated legal interpretation that uses large language models to extract concept-related information and interpret legal concepts.
Outcome: The proposed framework eliminates the need for legal experts to interpret legal concepts . it uses large language models to extract concept-related information and interpret legal concept interpretations .
Learning a Matching Model with Co-teaching for Multi-turn Response Selection in Retrieval-based Dialogue Systems (P19-1)

Copied to clipboard

Challenge: Existing methods for learning a robust matching model from noisy training data are retrieval-based or generation-based.
Approach: They propose a general co-teaching framework that learns matching models from noisy training data.
Outcome: The proposed learning framework can improve existing models on two public data sets.
Learning to Update Knowledge Graphs by Reading News (D19-1)

Copied to clipboard

Challenge: Existing methods to update knowledge graphs rely on elaborately designed IE systems and domain-specific rules.
Approach: They propose a novel neural network method to update knowledge graphs (KGs) they use a text-based attention mechanism to guide updating messages through KGs .
Outcome: The proposed method can effectively broadcast news information to KG structures and perform necessary link-adding or link-deleting operations to ensure the KG up-to-date according to news snippets.
Three Sentences Are All You Need: Local Path Enhanced Document Relation Extraction (2021.acl-short)

Copied to clipboard

Challenge: Document-level relation extraction (RE) is more challenging than sentence RE as it often requires reasoning over multiple sentences.
Approach: They propose a method to heuristically select evidence sentences for document-level relation extraction.
Outcome: The proposed method can be easily combined with BiLSTM to achieve good performance on benchmark datasets even better than fancy graph neural network based methods.
CASA: Causality-driven Argument Sufficiency Assessment (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods to assess the sufficiency of arguments are laborious and inconsistent due to subjective criteria.
Approach: They propose a causality-driven argument sufficiency assessment framework that uses the probability of sufficience to estimate the probability that a premise event would lead to a conclusion when both premise and conclusion events are absent.
Outcome: The proposed framework identifies insufficient arguments and improves them in a writing aid application.
Multi-grained Attention Network for Aspect-Level Sentiment Classification (D18-1)

Copied to clipboard

Challenge: Existing approaches to aspect sentiment classification use coarse-grained attention mechanisms . a novel approach captures word-level interaction between aspect and context .
Approach: They propose a novel multi-grained attention network model for aspect level sentiment classification . they use a fine-grounded attention mechanism to capture word-level interaction between aspect and context .
Outcome: The proposed model outperforms the state-of-the-art methods on three datasets . it shows that aspect-level interactions can bring extra useful information and improve performance .
MC2: Towards Transparent and Culturally-Aware NLP for Minority Languages in China (2024.acl-long)

Copied to clipboard

Challenge: MC2 is the largest open-source corpus of minority languages in china . MC2, however, includes four underrepresented languages: Tibetan, Uyghur, Kazakh, and Mongolian .
Approach: They propose a multilingual corpus of minority languages in China that includes four underrepresented languages . they prioritize accuracy while enhancing diversity by using a quality-centric approach .
Outcome: The proposed model prioritizes accuracy while enhancing diversity, the authors say . MC2 includes four underrepresented languages: Tibetan, Uyghur, Kazakh, and Mongolian .
Learning to Organize a Bag of Words into Sentences with Neural Networks: An Empirical Study (2021.naacl-main)

Copied to clipboard

Challenge: Existing approaches to encode natural languages without orders are lacking.
Approach: They conduct a comprehensive analysis of the ability of neural models to organize sentences from a bag of words under three typical scenarios.
Outcome: The proposed models can reorder or reconstruct sentences from a bag of words under three typical scenarios.
Marrying Up Regular Expressions with Neural Networks: A Case Study for Spoken Language Understanding (P18-1)

Copied to clipboard

Challenge: Experimental results show that the combination of regular expressions and NNs improves learning effectiveness when a small number of training examples are available.
Approach: They propose to combine a neural network (NN) with regular expressions (RE) to improve supervised learning for NLP by exploiting the rich expressiveness of REs at different levels within a NN.
Outcome: The proposed approach significantly improves learning effectiveness when a small number of training examples are available.
Sampling Matters! An Empirical Study of Negative Sampling Strategies for Learning of Matching Models in Retrieval-based Dialogue Systems (D19-1)

Copied to clipboard

Challenge: Existing studies focus on constructing a matching model with sophisticated neural architectures, but do little to how to effectively learn such architectures from data.
Approach: They propose to sample negative examples to automatically construct a training set for effective model learning in retrieval-based dialogue systems by using four sampling strategies.
Outcome: The proposed learning method improves the performance of matching models on two benchmarks with three matching models.
D2Plan: Dual-Agent Dynamic Global Planning for Complex Retrieval-Augmented Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in reinforcement learning (RL) have empowered Large Language Models (LLMs) with the capability to perform autonomous retrieval during reasoning tasks.
Approach: They propose a "D2Plan" paradigm for retrieval-augmented reasoning that integrates a 'Reasoner' and a'Purifier'
Outcome: Experiments show that the proposed paradigm improves on QA benchmarks.
Chain of Condition: Construct, Verify and Solve Conditions for Conditional Question Answering (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for conditional question answering struggle with finding probable answers and identifying missing conditions.
Approach: They propose a conditional question answering prompting approach that first identifies all conditions and constructs their logical relationships explicitly according to the document, then verifyes whether these conditions are satisfied and finally solves the logical expression to indicate any missing conditions.
Outcome: The proposed method outperforms existing prompting baselines on two CQA benchmark datasets and can facilitate GPT-3.5-Turbo or GPT-4 to outperFORM all existing supervised models.
Structure-Discourse Hierarchical Graph for Conditional Question Answering on Long Documents (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to conditional question answering on long documents ignore document structure and discourse relations between sentences in document sections.
Approach: They construct a Structure-Discourse Hierarchical Graph and conduct bottom-up information propagation to address this issue.
Outcome: The proposed approach outperforms the existing methods on the conditional question answering on long documents by 3.0 EM score and 2.4 F1 score on answer measuring, and 2.2 EM and 1.9 F1 scores on jointly answer and condition measuring.
Motion Generation from Fine-grained Textual Descriptions (2024.lrec-main)

Copied to clipboard

Challenge: Existing models for motion generation from textual descriptions are limited to coarse-grained descriptions.
Approach: They build a large-scale language-motion dataset specializing in fine-grained textual descriptions . they feed it with step-by-step instructions with pseudo-code compulsory checks . quantitative evaluation shows that the model outperforms MotionDiffuse in generating spatially or chronologically composite motions .
Outcome: The proposed model outperforms existing models in generating human motion sequences from textual descriptions by a large margin.
Enhancing Structured Evidence Extraction for Fact Verification (2023.emnlp-main)

Copied to clipboard

Challenge: Open-domain fact verification requires extracting and integrating both structured and unstructured evidence to verify a claim.
Approach: They propose a method to enhance the extraction of structured evidence by leveraging the row and column semantics of tables.
Outcome: The proposed method achieves evidence recall of 60.01% on the test set, higher than the previous state-of-the-art method.
Everything Has a Cause: Leveraging Causal Inference in Legal Text Analysis (2021.naacl-main)

Copied to clipboard

Challenge: Existing studies focus on analyzing structured data, while mining causal relationship among factors from unstructured data is of great importance.
Approach: They propose a graph-based causal inference framework which builds causal graphs from fact descriptions without much human involvement.
Outcome: The proposed framework can capture nuance from fact descriptions among confusing charges and provide explainable discrimination in few-shot settings.
Counterfactual Recipe Generation: Exploring Compositional Generalization in a Realistic Scenario (2022.emnlp-main)

Copied to clipboard

Challenge: Existing models fail to learn and use culinary knowledge in a compositional way, argues a new study.
Approach: They propose a task that asks models to modify a base recipe according to the change of an ingredient.
Outcome: The proposed model can perform compositional generalization in a realistic setting . existing models have difficulties in modifying ingredients while preserving original style .
The Magic of IF: Investigating Causal Reasoning Abilities in Large Language Models of Code (2023.findings-acl)

Copied to clipboard

Challenge: entailment a)
Approach: entailment : We want to explore whether Code-LLMs with code prompts are better . encoding a code prompt is better than text-only LLMs, they say .
Outcome: entailment : Our results show that Code-LLMs with code prompts are better compared to text-only LLMs.
Harder Task Needs More Experts: Dynamic Routing in MoE Models (2024.acl-long)

Copied to clipboard

Challenge: Unlike existing MoE approaches that rely on fixed TopK Routing, our dynamic expert selection framework dynamically allocates experts based on the confidence level in expert selection for each input.
Approach: They propose a dynamic expert selection framework that dynamically allocates experts based on the confidence level in expert selection for each input.
Outcome: The proposed method achieves an average improvement of 0.7% with less than 90% activated parameters and outperforms dense models in QA and machine translation tasks.
Recipe2Plan: Evaluating Planning Abilities of LLMs for Efficient and Feasible Multitasking with Time Constraints Between Actions (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation benchmarks focus on single task performance, ignoring multitask planning and execution efficiency.
Approach: They propose a benchmark framework based on real-world cooking scenarios . recipe2plan challenges agents to optimize cooking time through parallel task execution .
Outcome: The proposed benchmarks highlight the need for improved temporal awareness and global multitasking capabilities in large language models.
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts (2026.acl-long)

Copied to clipboard

Challenge: Existing solutions to supervise the reasoning process are prohibitively expensive.
Approach: They propose a cost-effective reinforcement learning framework that enhances reasoning quality using a small, general-purpose LLM only.
Outcome: Experiments show that CLARity improves reasoning quality by 16.5% over standard outcome-based reinforcement learning (RL) human evaluations confirm substantial gains in factual correctness and reasoning coherence, leading to more trustworthy model outputs.
An Intra-Class Relation Guided Approach for Code Comment Generation (2023.findings-eacl)

Copied to clipboard

Challenge: Recent work in code comment generation assumes that all information required to generate comments is encoded in the target function itself, yet in most realistic situations, it is hard to understand a function in isolation from the surrounding context.
Approach: They propose a graph-based learning framework to capture various relations among functions in a class file.
Outcome: The proposed method outperforms baseline models on automatic and human evaluation metrics on a Java dataset collected from real-world projects.
Read it in Two Steps: Translating Extremely Low-Resource Languages with Code-Augmented Grammar Books (2025.acl-long)

Copied to clipboard

Challenge: Using code rules improves rule retrieval and application of grammar books in low-resource languages.
Approach: They propose to decompose a grammar rule retrieval and application step into two steps . they propose to represent grammar rules as code functions to facilitate LLM reasoning .
Outcome: The proposed model significantly boosts rule retrieval and application, resulting in 13.1% BLEU improvement.
MiLiC-Eval: Benchmarking Multilingual LLMs for China’s Minority Languages (2025.findings-acl)

Copied to clipboard

Challenge: Large language models excel in high-resource languages but struggle with low-resourced languages . minority languages such as Tibetan, Uyghur, Kazakh, and Mongolian are marginalized in NLP research due to limited digital representation and the scarcity of training data.
Approach: They propose a benchmark for minority languages in China that tracks the progress of large language models on low-resource languages.
Outcome: The proposed benchmark focuses on underrepresented writing systems and syntax-intensive tasks.
Modeling discourse cohesion for discourse parsing via memory network (P18-2)

Copied to clipboard

Challenge: Existing approaches to discourse parsing focus on studying the semantic and syntactic aspects of EDU pairs, but they do not address long span dependencies.
Approach: They propose a new transition-based discourse parser that takes discourse cohesion into account by using memory networks.
Outcome: The proposed method outperforms traditional features and improves performance on the RST discourse treebank.
Do Charge Prediction Models Learn Legal Theory? (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing models for charge prediction are sensitive, selective, and presumption of innocence . a recent study has shown that deep learning models can predict the charges accurately, but their reliability and interpretability are still underexplored.
Approach: They propose that trustworthy charge prediction models should take legal theories into consideration . they propose three principles for trustworthy models to follow in this task .
Outcome: The proposed framework evaluates whether existing models learn legal theories . it shows that models meet selective and presumption of innocence principles .
Why Machine Reading Comprehension Models Learn Shortcuts? (2021.findings-acl)

Copied to clipboard

Challenge: Existing studies show that many MRC models learn shortcuts to outwit benchmarks, but the performance is unsatisfactory in real-world applications.
Approach: They propose to use shortcut questions to analyze learning difficulty of MRC models . they propose to analyze the learning difficulty regarding shortcut and challenging questions .
Outcome: The proposed methods show that a large proportion of shortcut questions in training data make models rely on shortcut tricks excessively.
Neighborhood Matching Network for Entity Alignment (2020.acl-main)

Copied to clipboard

Challenge: Structural heterogeneity between knowledge graphs is an outstanding challenge for entity alignment.
Approach: They propose a framework for entity alignment that uses a neighborhood matching module to combine neighborhood differences.
Outcome: The proposed framework outperforms existing methods on three datasets.
Improve Discourse Dependency Parsing with Contextualized Representations (2022.findings-naacl)

Copied to clipboard

Challenge: Existing studies show that discourse dependency analysis is easier when describing text units in a context-dependent way.
Approach: They propose to use transformers to encode contextualized representations of units of different levels to capture information needed for discourse dependency analysis.
Outcome: The proposed model outperforms traditional direct classification methods on English and Chinese datasets.
ProTrix: Building Models for Planning and Reasoning over Tables with Sentence Context (2024.findings-emnlp)

Copied to clipboard

Challenge: Tables are a crucial tool for organizing and presenting information in various domains.
Approach: They propose a Plan-then-Reason framework to answer different types of user queries over tables with sentence context.
Outcome: The proposed framework outperforms existing frameworks without self-consistency while using less API calls and in-context demonstrations.
Easy First Relation Extraction with Information Redundancy (D19-1)

Copied to clipboard

Challenge: Existing relation extraction models make decisions globally using integer linear programming . Existing approaches require time and memory to encode redundant information for ILP .
Approach: They propose an easy first approach for relation extraction with information redundancies embedded in local sentence extractors to resolve conflict decisions with domain and uniqueness constraints.
Outcome: The proposed approach outperforms both ILP and neural network-based methods in relation extraction (RE) studies have shown that the proposed approach improves the efficiency and accuracy of RE models.
Generating Classical Chinese Poems from Vernacular Chinese (D19-1)

Copied to clipboard

Challenge: Existing models for classical Chinese poetry generation only allow users to use keywords to interfere with the meaning of generated poems.
Approach: They propose a model to generate classical Chinese poems from vernacular . their model uses unsupervised machine translation to generate Chinese poems . human evaluation shows it can generate high-quality poems comparable to amateur poems - authors .
Outcome: The proposed model improves the perplexity and BLEU of the proposed model compared with typical models and human evaluation shows it generates high-quality poems comparable to amateur poems.
How Much Do Large Language Models Know about Human Motion? A Case Study in 3D Avatar Control (2025.findings-emnlp)

Copied to clipboard

Challenge: a new study explores the human motion knowledge of Large Language Models (LLMs) using 3D avatar control.
Approach: They use 20 representative motion instructions to interpolate LLMs into avatar animations . they find they are strong at interpreting high-level body movements but struggle with precise body part positioning .
Outcome: The proposed model is strong at interpreting high-level body movements but struggles with precise body part positioning.
Extract, Integrate, Compete: Towards Verification Style Reading Comprehension (2021.findings-emnlp)

Copied to clipboard

Challenge: VGaokao is a verification style reading comprehension dataset for Chinese language tests requiring advanced language understanding skills.
Approach: They propose a new extract-integration-compete approach to extract complementary evidence from Chinese Language tests of Gaokao and a pairwise competition to push models to learn the subtle difference between similar text pieces.
Outcome: The proposed approach outperforms baselines on VGaokao with retrieved complementary evidence while having the merits of efficiency and explainability.
Exploring Distantly-Labeled Rationales in Neural Network Models (2021.acl-long)

Copied to clipboard

Challenge: Existing methods focus on distantly-labeled rationales, ignoring the potential important non-rationale words and not distinguishing the importance of different rationale words.
Approach: They propose two novel auxiliary loss functions to make better use of distantly-labeled rationales, which encourage models to maintain their focus on important words beyond labeled rationals (PINs) and alleviate redundant training on non-helpful rationale (NoIRs).
Outcome: The proposed methods outperform existing methods on two representative classification tasks while maintaining the ability to spread focus to other unlabeled important words.
Enhancing Key-Value Memory Neural Networks for Knowledge Based Question Answering (N19-1)

Copied to clipboard

Challenge: Existing Key-value Memory Neural Networks are effective for shallow reasoning over documents . but extending them to Knowledge Based Question Answering is not trivial .
Approach: They propose a mechanism to enable conventional KV-MemNNs models to perform interpretable reasoning for complex questions.
Outcome: The proposed solution provides better reasoning abilities on complex questions and achieves state-of-the-art performance.
From the One, Judge of the Whole: Typed Entailment Graph Construction with Predicate Generation (2023.acl-long)

Copied to clipboard

Challenge: Existing methods to construct entailment graphs suffer from severe sparsity issues due to limited corpora and the long-tail phenomenon of predicate distributions.
Approach: They propose a multi-stage method to generate entailment graphs by generating new predicates and detecting enanglement relations among seed predicats.
Outcome: The proposed method can generate high-quality graphs with high precision over state-of-the-art methods and boost the performance of down-stream inference tasks.
Entailment Graph Learning with Textual Entailment and Soft Transitivity (2022.acl-long)

Copied to clipboard

Challenge: Typed entailment graphs suffer from severe sparsity and unreliability of distributional similarity . enlargement relation is critical to semantic understanding and natural language inference .
Approach: They propose a method to learn local entailment relations by recognizing textual enanglement between template sentences formed by typed CCG-parsed predicates.
Outcome: The proposed method can model transitivity in entailment graphs to alleviate sparsity and improve performance over current methods.
Efficient Low-Resource Language Adaptation via Multi-Source Dynamic Logit Fusion (2026.acl-long)

Copied to clipboard

Challenge: Proxy Tuning offers a logit-level strategy for introducing scaling effects, but it often fails in LRL settings because the large model’s weak LRL competence might overwhelm the knowledge of specialized smaller models.
Approach: They propose a logit-based framework that balances LRL competence from a continually pretrained small model, task competence from high-resource language instruction tuning, and the scaling benefits of large models.
Outcome: Experiments across four model families and eight LRLs show that TriMix outperforms single-model baselines and Proxy Tuning.
Semantic Graphs for Generating Deep Questions (2020.acl-main)

Copied to clipboard

Challenge: Existing research has focused on generating factoid questions relevant to one fact obtainable from a single sentence.
Approach: They propose a framework that first constructs a semantic-level graph and then encodes it by introducing an attention-based GGNN.
Outcome: The proposed framework captures the global structure of the document and facilitates reasoning over multiple facts.
Learning Dynamic Representations for Discourse Dependency Parsing (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing models characterize transition states by examining a certain number of elementary discourse units (EDUs) Existing work neglects the arcs obtained from the transition history.
Approach: They propose to employ GAT-based encoder to learn dynamic representations for sub-trees constructed in previous transition steps.
Outcome: The proposed model retains access to parsed EDUs through the obtained arcs, especially when handling lengthy text spans with complex structures.
Teaching Large Language Models an Unseen Language on the Fly (2024.findings-acl)

Copied to clipboard

Challenge: Existing large language models struggle to support numerous low-resource languages . Existing models lack sufficient training data for effective parameter updating .
Approach: They propose a framework for adapting LLMs to unseen languages by in-context learning.
Outcome: The proposed framework improves Chinese-to-Zhuang translation performance and Zhuan-to Chinese translation performance.
Jointly Learning Entity and Relation Representations for Entity Alignment (D19-1)

Copied to clipboard

Challenge: Entity alignment is a viable method for integrating heterogeneous knowledge among different knowledge graphs (KGs).
Approach: They propose a Graph Convolutional Network-based framework for learning relation representations by embedding relation seeds into entities and incorporating relation approximation into entities to iteratively improve alignment.
Outcome: The proposed approach outperforms state-of-the-art methods on three real-world cross-lingual datasets.
Things not Written in Text: Exploring Spatial Commonsense from Visual Signals (2022.acl-long)

Copied to clipboard

Challenge: Pretrained language models fail in many NLP tasks, but are ineffective in spatial commonsense reasoning.
Approach: They propose a spatial commonsense benchmark that focuses on relative scales of objects and the positional relationship between people and objects under different actions.
Outcome: The proposed framework outperforms pretrained models in answering spatial questions.
JUREX-4E: Juridical Expert-Annotated Four-Element Knowledge Base for Legal Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies have introduced legal theories into LLM workflows to improve their understanding of legal texts and reasoning accuracy.
Approach: They evaluate an expert-annotated four-element knowledge base covering 155 criminal charges.
Outcome: The proposed model can be used to analyze criminal charges and retrieve them in legal cases.
Does Recommend-Revise Produce Reliable Annotations? An Analysis on Missing Instances in DocRED (2022.acl-long)

Copied to clipboard

Challenge: Document-level relation extraction is a challenging task as it requires reasoning across multiple sentences.
Approach: They propose to use a recommend-revise scheme to reduce the workload of annotators by providing them with candidate relation instances from distant supervision to supplement and remove relational facts.
Outcome: The proposed dataset is the first large-scale and human-annotated dataset for relation extraction.
Cross-lingual Knowledge Graph Alignment via Graph Matching Neural Network (P19-1)

Copied to clipboard

Challenge: Existing approaches to cross-lingual knowledge graph (KG) alignment rely on entity embeddings derived from monolingual KG structural information.
Approach: They propose a topic entity graph to represent entities with contextual information in KGs.
Outcome: The proposed model outperforms state-of-the-art methods by a large margin.
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data (2024.findings-acl)

Copied to clipboard

Challenge: Quantitative reasoning with data is a critical skill to analyze data, yet the assessment of such ability remains limited.
Approach: They propose a quantitative reasoning with data benchmark to evaluate Large Language Models' ability in statistical and causal reasoning with real-world data.
Outcome: The proposed model GPT-4 achieves an accuracy of 58%, while open-source model Deepseek-coder-instruct gets the highest accuracy of 37%.
Cross-Lingual Transfer of Cultural Knowledge: An Asymmetric Phenomenon (2025.acl-short)

Copied to clipboard

Challenge: Existing studies evaluate whether large language models handle global cultural diversity . however, mechanisms behind cultural knowledge acquisition remain unexplored .
Approach: They propose an interpretable framework to study cultural knowledge transfer in large language models . they observe bidirectional cultural transfer between English and other high-resource languages .
Outcome: The proposed framework ensures training data transparency and controls transfer effects.
Lattice-BERT: Leveraging Multi-Granularity Representations in Chinese Pre-trained Language Models (2021.naacl-main)

Copied to clipboard

Challenge: Pre-trained language models process text as a sequence of characters, ignoring more coarse granularity, e.g., words.
Approach: They propose a new pre-training paradigm for Chinese that incorporates word representations along with characters and can model a sentence in a multi-granular manner.
Outcome: The proposed model can bring an average increase of 1.5% under the 12-layer setting, which achieves new state-of-the-art among base-size models on the CLUE benchmarks.
DiNeR: A Large Realistic Dataset for Evaluating Compositional Generalization (2023.emnlp-main)

Copied to clipboard

Challenge: Existing compositional generalization datasets lack natural language variation due to limited data scale or lack of diversity.
Approach: They propose a compositional generalization task to evaluate natural language understanding ability under compositional settings.
Outcome: The proposed method outperforms the plain seq2seq trained version by a large margin . it uses two strong baseline methods and large language models to tackle the task .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations