Papers by Chenghua Lin

68 papers
Enhancing Dialogue Generation via Dynamic Graph Knowledge Aggregation (2023.acl-long)

Copied to clipboard

Challenge: Existing graph neural networks (GNNs) teach message passing on a graph from text, resulting in a semantic gap between graph knowledge and text.
Approach: They propose a framework to integrate external graph knowledge into chatbots by coagulating representations of both text and graph knowledge.
Outcome: The proposed framework outperforms state-of-the-art (SOTA) baselines on dialogue generation.
DGST: a Dual-Generator Network for Text Style Transfer (2020.emnlp-main)

Copied to clipboard

Challenge: Existing studies on text style transfer focus on altering sentiment words to preserve attribute-independent information.
Approach: They propose a Dual-Generator network architecture for text Style Transfer using two generators.
Outcome: The proposed model performs better than existing models on Yelp and IMDb datasets.
ATLAS: Improving Lay Summarisation with Attribute-based Control (2024.acl-short)

Copied to clipboard

Challenge: Lay summarisation aims to produce scientific summaries that are comprehensible to non-experts.
Approach: They propose an abstractive summarisation approach that can control properties contributing to overall "layness" they evaluate ATLAS on a combination of biomedical lay summarization datasets.
Outcome: The proposed approach outperforms state-of-the-art summarisation metrics on biomedical datasets and shows that it can be discriminatory and emergently influenced.
Summarising Historical Text in Modern Languages (2021.eacl-main)

Copied to clipboard

Challenge: Historical text summarisation is a routine for historians and digital humanities researchers but has never been automated.
Approach: They propose a model that can be trained even with no cross-lingual data and further benchmark it against state-of-the-art algorithms.
Outcome: The proposed model outperforms standard cross-lingual benchmarks on historical text summarisation task and identifies distinctness and value of the dataset.
With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models and Vision Language Model (VLMs) have demonstrated aptitude as potential substitutes for human participants in psycholinguistic experiments.
Approach: They examine whether large language models and vision language models implicitly understand sound-based phenomena via orthography and imagery alone.
Outcome: The proposed models demonstrate sound symbolism and ability to "hear" using language and vision modules.
Evaluating Open-Domain Dialogues in Latent Space with Next Sentence Prediction and Mutual Information (2023.acl-long)

Copied to clipboard

Challenge: Existing evaluation methods for open-domain dialogues are difficult due to the one-to-many issue of the open- domain dialogues.
Approach: They propose a learning-based automatic evaluation metric which can robustly evaluate open-domain dialogues by augmenting CVAEs with a Next Sentence Prediction objective and employing Mutual Information to model the semantic similarity of text in the latent space.
Outcome: The proposed method can evaluate open-domain dialogues on two open- domain dialogue datasets.
HERB: Measuring Hierarchical Regional Bias in Pre-trained Language Models (2022.findings-aacl)

Copied to clipboard

Challenge: Existing methods do not examine social groups categorised by geographical information, leaving the region-related biases in pre-trained LMs unexplored.
Approach: They propose a hierarchical regional bias evaluation method to quantify regional bias in pre-trained language models.
Outcome: The proposed method evaluates regional bias with regard to comprehensive topics and measures potential regional bias that can be propagated to downstream tasks.
From Facts to Insights: A Study on the Generation and Evaluation of Analytical Reports for Deciphering Earnings Calls (2025.coling-main)

Copied to clipboard

Challenge: Existing studies have focused on the generation and evaluation of analytical reports derived from Earnings Calls (ECs).
Approach: They propose to use Large Language Models to generate and evaluate analytical reports derived from Earnings Calls (ECs) they propose to introduce specialized agents that introduce diverse viewpoints and desirable topics into the report generation process.
Outcome: The proposed model improves the quality of reports in different settings, while human-written reports remain preferred in the majority of cases.
Effective Distillation of Table-based Reasoning Ability from LLMs (2024.lrec-main)

Copied to clipboard

Challenge: Existing work on table-based reasoning distillation has focused on smaller models with limited performance.
Approach: They propose a table-based reasoning distillation approach to distill LLMs into smaller models . their results show that a 220 million parameter model fine-tuned using distilled data improves performance .
Outcome: The proposed model improves on a scientific table-to-text generation dataset and surpasses specific LLMs.
On the Rigour of Scientific Writing: Criteria, Analysis, and Insights (2024.findings-emnlp)

Copied to clipboard

Challenge: despite its importance, little work exists on modelling rigour in scientific writing . despite widespread use of term, scientific literature lacks definition of rigor .
Approach: They propose a framework to automatically identify and define rigour criteria and assess their relevance in scientific writing.
Outcome: The proposed framework can be tailored to the evaluation of scientific rigour for different areas.
DocMMIR: A Framework for Document Multi-modal Information Retrieval (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing multi-modal information retrieval models lack a comprehensive exploration of document-level retrieval . existing models suffer from the absence of cross-domain datasets at this granularity.
Approach: They propose a multi-modal document retrieval framework to unify diverse document formats and domains with a comprehensive retrieval scenario.
Outcome: The proposed framework improves document retrieval performance on a large multimodal dataset.
Emphasising Structured Information: Integrating Abstract Meaning Representation into LLMs for Enhanced Open-Domain Dialogue Evaluation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation metrics struggle to evaluate adversarial negative examples . existing metrics struggle in handling adversarials, resulting in low correlations with human judgments.
Approach: They propose a framework that integrates AMR and domain-specific language models for automatic open-domain dialogue evaluation.
Outcome: The proposed evaluation framework achieves strong correlations with human judgments across multiple datasets.
Guiding the Growth: Difficulty-Controllable Question Generation through Step-by-Step Rewriting (2021.acl-long)

Copied to clipboard

Challenge: Existing QG systems perform substantially worse in answering multi-hop questions than single-hop ones.
Approach: They propose a framework that progressively increases question difficulty through step-by-step rewriting under the guidance of an extracted reasoning chain.
Outcome: The proposed framework increases question difficulty through step-by-step rewriting under the guidance of an extracted reasoning chain.
Selecting Demonstrations for Many-Shot In-Context Learning via Gradient Matching (2025.findings-acl)

Copied to clipboard

Challenge: In-Context Learning (ICL) empowers Large Language Models for rapid task adaptation without fine-tuning.
Approach: They propose a method that aligns fine-tuning gradients between entire training set and selected examples to enable in-context learning and fine-uning.
Outcome: The proposed method outperforms random selection on large LLMs from 4-shot to 128-shot scenarios across 9 datasets.
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: Scientific rigour tends to be sidelined in favour of bold statements, leading authors to overstate claims beyond what their results support.
Approach: They propose a multimodal framework that retrieves supporting evidence from a paper and assigns each claim an overstatement score.
Outcome: The proposed framework retrieves supporting evidence from ICLR and NeurIPS papers and assigns each claim an overstatement score.
LATENTLOGIC: Learning Logic Rules in Latent Space over Knowledge Graphs (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for learning logic rules for knowledge graph reasoning face limitations such as searching in vast search space and inefficient optimization.
Approach: They propose a framework to efficiently mine logic rules by controllable generation in the latent space by a pre-trained VAE and a discriminator.
Outcome: The proposed framework efficiently mines logic rules by controllable generation in the latent space.
Compressing Context to Enhance Inference Efficiency of Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable power and impressive generalisation abilities across various tasks.
Approach: They propose a method that prunes redundancies in the input context to make the input more compact.
Outcome: The proposed method reduces memory and inference time while maintaining comparable performance compared to full context.
Fast and Scalable Dialogue State Tracking with Explicit Modular Decomposition (2021.naacl-main)

Copied to clipboard

Challenge: Existing approaches for dialogue state tracking are mainly based on classification-based and extraction-based methods.
Approach: They propose a model which incorporates both classification-based and extraction-based methods and integrates four modules to jointly extract dialogue states.
Outcome: The proposed model outperforms the state-of-the-art models in multi-domain dialogues with many turns of utterances.
Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth (2025.emnlp-main)

Copied to clipboard

Challenge: Despite excelling at many natural language processing tasks, large language models fail to grasp the layered semantics of Drivelological text.
Approach: They construct a benchmark dataset of over 1,200+ carefully curated and diverse examples across English, Mandarin, Spanish, French, Japanese, and Korean to examine their Drivelological characteristics.
Outcome: The proposed models lack conceptual understanding and lack conceptual and semantic accuracy.
Word Embedding and WordNet Based Metaphor Identification and Interpretation (P18-1)

Copied to clipboard

Challenge: Existing models cannot identify exact metaphorical words within a sentence . current models do not rely on hand-crafted knowledge for training .
Approach: They propose an unsupervised learning method that identifies and interprets metaphors at word-level without preprocessing.
Outcome: The proposed method outperforms baseline models in two translation systems for English to Chinese showing that it paraphrases metaphors into their literal counterparts.
Highly Efficient Knowledge Graph Embedding Learning with Orthogonal Procrustes Analysis (2021.naacl-main)

Copied to clipboard

Challenge: Knowledge Graph Embeddings (KGEs) have been explored in recent years due to their promise for a wide range of applications.
Approach: They propose a KGE framework which can reduce the training time and carbon footprint by orders of magnitudes compared with state-of-the-art approaches.
Outcome: The proposed framework reduces the training time and carbon footprint by orders of magnitudes compared with state-of-the-art approaches while producing competitive performance.
Development of a Benchmark Corpus to Support Entity Recognition in Job Descriptions (2022.lrec-1)

Copied to clipboard

Challenge: Existing tools for identifying and extracting salient entities from job descriptions are limited by the lack of publicly available training data.
Approach: They propose to use a standard definition of entities and a training corpus to develop a benchmark Entity Recognition (ER) model.
Outcome: The proposed model achieves an F1 score of 0.59 from 18.6k entities comprising five types (Skill, Qualification, Experience, Occupation, and Domain).
Metaphor Detection via Explicit Basic Meanings Modelling (2023.acl-short)

Copied to clipboard

Challenge: Existing methods for metaphor detection use the aggregated meaning of a word to approximate its basic meaning.
Approach: They propose a method which models the basic meaning of a word based on literal annotations and compares this with the contextual meaning in a target sentence to identify metaphors.
Outcome: The proposed method outperforms the state-of-the-art method significantly in the F1 score and even reaches the theoretical upper bound on the VUA18 benchmark.
Improving Chinese Story Generation via Awareness of Syntactic Dependencies and Semantics (2022.aacl-short)

Copied to clipboard

Challenge: Current neural models for Chinese story generation struggle to generate high-quality long text narratives due to ambiguity in syntactically parsing the Chinese language.
Approach: They propose a framework that enhances the feature capturing mechanism by informing the generation model of dependencies between words and additionally augmenting the semantic representation learning through synonym denoising training.
Outcome: The proposed framework outperforms the state-of-the-art Chinese generation models on all evaluation metrics, showing that it enhances dependency and semantic representation learning.
MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical Language (2024.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation methods focus on fluency and factual reliability, while neglecting figurative quality.
Approach: They propose a set of human evaluation metrics focused on the translation of figurative language and a parallel metaphor corpus generated by post-editing.
Outcome: The proposed evaluation protocol estimates four aspects of MT: Metaphorical Equivalence, Emotion, Authenticity, and Quality.
SLIDE: A Framework Integrating Small and Large Language Models for Open-Domain Dialogues Evaluation (2024.findings-acl)

Copied to clipboard

Challenge: Existing approaches to evaluate open domain dialogues have a one-to-many problem . existing approaches lack commonsense reasoning biases and perform poorly in domain-specific scenarios.
Approach: They propose a framework that leverages both a small, specialised model and LLMs for the evaluation of open-domain dialogues.
Outcome: The proposed framework achieves state-of-the-art performance in both classification and evaluation tasks and exhibits better correlation with human judgements.
Making Science Simple: Corpora for the Lay Summarisation of Scientific Literature (2022.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for lay summarisation are limited in size and scope, hindering the development of data-driven approaches.
Approach: They propose to use two new datasets for the lay summarisation of biomedical research articles to characterise their lay summaries.
Outcome: The proposed datasets are compared with existing datasets and show they can be leveraged to support different audiences and applications.
Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment (2025.coling-main)

Copied to clipboard

Challenge: Human values are inherently diverse, making it insufficient to align LLMs solely with general preferences.
Approach: They propose a flexible paradigm for individual preference alignment that disentangles preference representation from text generation in LLMs.
Outcome: The proposed method produces aligned quality and better than PEFT-based methods while reducing training time for each new individual preference by 80% to 90%.
CIF-Bench: A Chinese Instruction-Following Benchmark for Evaluating the Generalizability of Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: a recent study shows that large language models have limited generalization in low-resource languages like Chinese.
Approach: They propose to evaluate the zero-shot generalizability of large language models to the Chinese language . they release only half of the dataset publicly, with the remainder kept private .
Outcome: The Chinese Instruction-Following Benchmark evaluates the generalizability of LLMs to the Chinese language.
Improving Variational Autoencoder for Text Modelling with Timestep-Wise Regularisation (2020.coling-main)

Copied to clipboard

Challenge: Variational Autoencoders (VAEs) have been widely used in text modelling but posterior collapse is a problem when RNN-based models are employed.
Approach: They propose a timestep-wise regularisation VAE architecture which can effectively avoid posterior collapse when used in text modelling.
Outcome: The proposed model avoids posterior collapse and can be applied to any RNN-based VAE model.
EtriCA: Event-Triggered Context-Aware Story Generation Augmented by Cross Attention (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for story generation still suffer from problems of relevance and coherence.
Approach: They propose a novel neural generation model which maps contextual and event features to event sequences with a cross-attention mechanism and exploits logical relatedness between events.
Outcome: The proposed model outperforms state-of-the-art models on automatic and human evaluations and shows that it can leverage contextual and event features.
Length is a Curse and a Blessing for Document-level Semantics (2023.emnlp-main)

Copied to clipboard

Challenge: In recent years, contrastive learning (CL) has been extensively utilized to recover sentence and document-level encoding capability from pre-trained language models.
Approach: They propose a document-based contrastive learning framework that is length-agnostic self-reference based on document length.
Outcome: The proposed framework achieves state-of-the-art on the standard information retrieval benchmark.
How to Determine the Most Powerful Pre-trained Language Model without Brute Force Fine-tuning? An Empirical Survey (2023.findings-emnlp)

Copied to clipboard

Challenge: Transferability estimation has been a topic of great interest in computer vision fields . a lack of a comprehensive comparison between these estimation methods is a problem .
Approach: They conduct a thorough survey of existing methods to find the most suitable model . they also outline difficulties of consideration of training details and applicability to text generation .
Outcome: The proposed methods perform well with superiorities in effectiveness and efficiency.
Enhancing Biomedical Lay Summarisation with External Knowledge Graphs (2023.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to lay summarisation are reliant on the source article, which is unlikely to include all the information necessary for a lay audience.
Approach: They augment existing biomedical lay summarisation dataset with article-specific knowledge graphs that contain detailed information on relevant biomedically related concepts.
Outcome: The proposed methods improve readability and explanation of technical concepts by integrating graph-based domain knowledge within lay summarisation models.
ChatMusician: Understanding and Generating Music Intrinsically with LLM (2024.findings-acl)

Copied to clipboard

Challenge: Despite LLMs' impressive capabilities in musical knowledge, music reasoning remains an unsolved task.
Approach: They propose an open-source large language model (LLM) that integrates intrinsic musical abilities into LLaMA2 and GPT-3.5.
Outcome: The proposed model can understand and generate music with a pure text tokenizer without external multi-modal neural structures or tokenizers.
CAST: Corpus-Aware Self-similarity Enhanced Topic modelling (2025.naacl-long)

Copied to clipboard

Challenge: Existing topic modelling methods encode contextual information of documents while ignoring contextual details of candidate centroid words. Existing methods are limited by the contextualization gap.
Approach: They propose a topic modelling method that builds upon candidate centroid word embeddings contextualized on the dataset and a self-similarity-based method to filter out less meaningful tokens.
Outcome: The proposed method significantly enhances the coherence and diversity of generated topics, and handles noisy data, outperforming strong baselines.
Does Acceleration Cause Hidden Instability in Vision Language Models? Uncovering Instance-Level Divergence Through a Large-Scale Empirical Study (2025.emnlp-main)

Copied to clipboard

Challenge: Current acceleration evaluations focus on minimal overall performance degradation . however, accelerated models can exhibit significant changes in instance-level predictions .
Approach: They investigate whether accelerated vision-Language Models can still give the same answers as before . they found that accelerated models changed original answers up to 20% of the time .
Outcome: The results show that accelerated models changed their original answers up to 20% of the time.
Can MLLMs Understand the Deep Implication Behind Chinese Images? (2025.acl-long)

Copied to clipboard

Challenge: MLLMs perform poorly on traditional culture images, indicating limitations in understanding high-level semantics and lacking a deep knowledge base of Chinese traditional culture.
Approach: They propose to use Chinese images to assess MLLMs' higher-order perception and understanding of Chinese visual content.
Outcome: The proposed model incorporates images that represent Chinese traditional culture, such as famous Chinese traditional paintings, to ensure the authenticity of the Chinese context.
CMDAG: A Chinese Metaphor Dataset with Annotated Grounds as CoT for Boosting Metaphor Generation (2024.lrec-main)

Copied to clipboard

Challenge: Metaphors are a prominent linguistic device in human language and literature, as they add color, imagery, and emphasis to enhance effective communication.
Approach: They propose a large-scale high quality annotated Chinese Metaphor Corpus . they use a set of guidelines to ensure the accuracy and consistency of their annotations .
Outcome: The proposed corpus generates metaphors that resonate more with real-world intuition.
TranSHER: Translating Knowledge Graph Embedding with Hyper-Ellipsoidal Restriction (2022.emnlp-main)

Copied to clipboard

Challenge: Existing knowledge graph embedding methods restrict entities on hyper-ellipsoid surfaces, resulting in suboptimal knowledge graph completion.
Approach: They propose a score function that leverages relation-specific translations between head and tail entities to relax constraints on hyper-ellipsoid surfaces.
Outcome: The proposed method achieves state-of-the-art performance on link prediction and generalizes well to datasets in different domains and scales.
LVPruning: An Effective yet Simple Language-Guided Vision Token Pruning Approach for Multi-modal Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Multi-modal Large Language Models (MLLMs) incur significant computational overhead due to the large number of vision tokens processed, limiting their practicality in resource-constrained environments.
Approach: They propose a language-guided vision token pruning method that can be integrated into existing MLLMs with minimal architectural changes.
Outcome: The proposed method reduces vision tokens by 90% and preserves model performance.
The Iron(ic) Melting Pot: Reviewing Human Evaluation in Humour, Irony and Sarcasm Generation (2023.findings-emnlp)

Copied to clipboard

Challenge: Human evaluation is often considered to be the gold standard method of evaluating a Natural Language Generation system, but its quality is often brought into question.
Approach: They argue that the generation of more esoteric forms of language constitutes a subdomain where the characteristics of selected evaluator panels are of utmost importance.
Outcome: The proposed system generates coherent and well-formed text of a particular type, usually given an input such as a prompt, outline, topic, or data.
HiRAS: A Hierarchical Multi-Agent Framework for Paper-to-Code Generation and Execution (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to automate computational research use fixed sequential agent pipelines with weak global coordination, which limits their robustness and overall performance.
Approach: They propose a hierarchical multi-agent framework for end-to-end paper reproduction that employs supervisory manager agents to coordinate specialised agents across fine-grained stages.
Outcome: The proposed framework improves the paper2code benchmark and significantly reduces hallucination in the evaluation.
TwistList: Resources and Baselines for Tongue Twister Generation (2023.acl-short)

Copied to clipboard

Challenge: Previous work in phonetically-grounded language generation has focused on domains such as lyrics and poetry.
Approach: They propose to use TwistList to generate phonetically constrained tongue twisters, a large annotated dataset consisting of 2.1K+ human-authored examples.
Outcome: The proposed models perform better than pre-trained models with limited training and data and no explicit phonetic knowledge.
COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values (2026.findings-eacl)

Copied to clipboard

Challenge: Existing Chinese preference datasets suffer from limited scale, restricted domain coverage, and insufficiently rigorous data validation.
Approach: They propose an LLM-based data annotation pipeline with no human intervention to annotate Chinese preference datasets.
Outcome: The proposed pipeline outperforms existing Chinese preference datasets on AlignBench and Chinese Reward Benchmark.
Token-level Proximal Policy Optimization for Query Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have improved search engines and recommendation systems through their text understanding capabilities.
Approach: They propose a token-level proximal policy optimization approach to empower LLMs to perform better in query generation through fine-tuning.
Outcome: The proposed approach outperforms existing LLMs on an open-source and industrial dataset.
Comparing Apples to Oranges: A Dataset & Analysis of LLM Humour Understanding from Traditional Puns to Topical Jokes (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing work on humour explanation has focused on short pun-based jokes, but Large Language Models (LLMs) are not capable of generating adequate explanations of all joke types.
Approach: They compare the ability of Large Language Models (LLMs) to explain humour from simple puns to complex topical humor that requires esoteric knowledge of real-world entities and events.
Outcome: The proposed models are incapable of generating adequate explanations of all joke types, highlighting the narrow focus of most existing work on overly simple joke forms.
Leveraging Estimated Transferability Over Human Intuition for Model Selection in Text Ranking (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for text ranking are based on intuition, but their estimated transferability may not align well with the objectives of text ranking.
Approach: They propose to compute expected rank as transferability, explicitly reflecting the model’s ranking capability.
Outcome: The proposed method shows significant improvements over previous classification-oriented TE methods, human intuition, and ChatGPT with minor time consumption.
Lightweight Contenders: Navigating Semi-Supervised Text Mining through Peer Collaboration and Self Transcendence (2025.findings-naacl)

Copied to clipboard

Challenge: Existing frameworks for semi-supervised text mining with lightweight models are limited by label data scarcity.
Approach: They propose a framework for semi-supervised text mining with lightweight models . it incorporates online distillation to train lightweight student models by imitating the Teacher model .
Outcome: The proposed framework exhibits notable performance enhancements over existing frameworks.
Improving Biomedical Abstractive Summarisation with Knowledge Aggregation from Citation Papers (2023.emnlp-main)

Copied to clipboard

Challenge: Existing language models struggle to generate technical summaries that are on par with those produced by biomedical experts due to the lack of domain-specific background knowledge.
Approach: They propose a attention-based citation aggregation model that integrates domain-specific knowledge from citation papers and a large-scale biomedical summarisation dataset to build on.
Outcome: The proposed model outperforms state-of-the-art approaches and achieves substantial improvements in biomedical abstractive summarisation.
Cross-Lingual Word Embedding Refinement by ℓ1 Norm Optimisation (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods for building high-quality CLWEs learn mappings that minimise the l2 norm loss function but this optimisation objective has been shown to be sensitive to outliers.
Approach: They propose a simple post-processing step to improve cross-lingual word embeddings using the Manhattan norm goodness-of-fit criterion.
Outcome: The proposed approach outperforms four state-of-the-art baselines in bilingual lexicon induction and cross-lingual transfer tasks.
NGEP: A Graph-based Event Planning Framework for Story Generation (2022.aacl-short)

Copied to clipboard

Challenge: Current approaches to story generation are based on end-to-end neural generation models, such as BART, to generate event sequences.
Approach: They propose a novel event planning framework which generates an event sequence by performing inference on an automatically constructed event graph and enhances generalisation ability through a neural event advisor.
Outcome: The proposed framework outperforms state-of-the-art (SOTA) event planning approaches on multiple criteria and compares with existing models on the downstream task of story generation.
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval (2024.findings-acl)

Copied to clipboard

Challenge: Multi-modal information retrieval (MMIR) is a rapidly evolving field . current benchmarks for image-text pairings overlook the scientific domain .
Approach: They develop a scientific domain-specific MMIR benchmark to evaluate image-text pairings using open-access research paper corpora.
Outcome: The proposed benchmarks are based on 530K image-text pairs extracted from scientific documents with detailed captions.
End-to-End Sequential Metaphor Identification Inspired by Linguistic Theories (P19-1)

Copied to clipboard

Challenge: Existing sequence tagging models do not explicitly exploit linguistic theories of metaphor identification.
Approach: They propose to exploit linguistic theories of metaphor identification in deep neural networks to improve model performance.
Outcome: The proposed models achieve state-of-the-art in end-to-end metaphor identification on three datasets.
Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond (2025.findings-emnlp)

Copied to clipboard

Challenge: Comp-Comp is an iterative benchmarking framework grounded in the principles of comprehensiveness and compactness.
Approach: They propose a benchmark framework that incorporates the principle of comprehensiveness and compactness.
Outcome: The proposed framework is domain-agnostic and adaptable to a wide range of specialized fields.
Crafting Customisable Characters with LLMs: A Persona-Driven Role-Playing Agent Framework (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are capable of generating human-like text, but the potential for freely customisable characters remains underexplored.
Approach: They propose a framework which employs Large Language Models to create freely customisable characters through personalised characteristic feature injection.
Outcome: The proposed framework provides valuable insights for developing more accurate and customisable human simulacra.
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores (2024.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics for natural language generation tasks favor text generated by different LMs . human evaluation by experts is the most reliable approach, but it is costly and time-consuming .
Approach: They examine whether language model-driven evaluation metrics exhibit bias toward underlying language models in the context of summarization tasks.
Outcome: The proposed evaluation metrics tend to assign inflated scores to outputs generated by the very model they are based on.
MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: Current multimodal benchmarks focus on facts within individual images, but neglect associative relations among multiple images.
Approach: They propose a multi-image relational association task and a MMRA benchmark to evaluate LVLMs.
Outcome: The proposed benchmarks show that entity-level multi-image perception tasks pose greater challenges than image-level tasks.
X-ray Made Simple: Lay Radiology Report Generation and Robust Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: Technical language and templated nature of professional reports hinder patient comprehension and allow models to artificially boost lexical metrics such as BLEU by reproducing common report patterns.
Approach: They propose a layman's RRG framework that leverages layperson-friendly language to enhance patient accessibility and promote robust evaluation and report generation by encouraging models to focus on semantic accuracy over rigid templates.
Outcome: The proposed framework improves model performance with more layman-style data, compared to templated professional language and inflated lexical scores.
CM-Gen: A Neural Framework for Chinese Metaphor Generation with Explicit Context Modelling (2022.coling-1)

Copied to clipboard

Challenge: Nominal metaphors are commonly used in human language and have been shown to be effective in persuading, expressing emotion, and stimulating interest.
Approach: They propose a multitask framework which optimizes three tasks: NM identification, NM component identification, and NM generation.
Outcome: The proposed framework outperforms baselines on consistency and creativity on the NM generation task in Chinese.
Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts (2026.tacl-1)

Copied to clipboard

Challenge: Evaluating natural language generation systems is challenging due to the diversity of valid outputs.
Approach: They propose an inversion learning method that learns effective reverse mappings from model outputs back to their input instructions.
Outcome: The proposed method requires only a single evaluation sample and eliminates manual prompt engineering.
Metaphor Detection with Effective Context Denoising (2023.eacl-main)

Copied to clipboard

Challenge: Existing models focus on semantically relevant information and provide a target-oriented parse tree structure for metaphor detection.
Approach: They propose a new model which introduces a target-oriented parse tree structure for metaphor detection.
Outcome: The proposed model achieves state-of-the-art on several main metaphor datasets and compares with other methods.
FrameBERT: Conceptual Metaphor Detection with Frame Embedding Learning (2023.eacl-main)

Copied to clipboard

Challenge: Existing models for concept-level metaphor detection lack explicit knowledge of FrameNet . Metaphor detection is a pervasive linguistic device that is used in cognitive and communicative functions of language.
Approach: They propose a BERT-based model that explicitly learns FrameNet Embeddings for metaphor detection.
Outcome: The proposed model is more explainable and interpretable than existing models.
COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning (2025.findings-naacl)

Copied to clipboard

Challenge: Existing datasets for Chinese instruction tuning are not well-aligned with Chinese users’ interaction patterns.
Approach: They propose to use Chinese instruction tuning datasets to improve instruction fine-tuning for Chinese users.
Outcome: The proposed dataset shows that Chinese models achieve competitive performance in diverse benchmarks.
Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks (2025.coling-main)

Copied to clipboard

Challenge: Recent advances in natural language processing have highlighted the vulnerability of deep learning models to adversarial attacks.
Approach: They propose a benchmark for textual adversarial defence that evaluates state-of-the-art defence mechanisms across diverse datasets, models, and tasks.
Outcome: The proposed benchmark incorporates a wide range of datasets and evaluates state-of-the-art defence mechanisms.
LIME: Less Is More for MLLM Evaluation (2025.findings-acl)

Copied to clipboard

Challenge: Existing MLLM benchmarks and unified evaluation frameworks cannot accurately and efficiently reflect the ability of MLMLs.
Approach: They propose a semi-automated benchmark curated using a pipeline that filters out uninformative samples and eliminates answer leakage by focusing on tasks that require image-based understanding.
Outcome: The proposed benchmark reduces the number of samples by 76% and evaluation time by 77% while it can more effectively distinguish different models’ abilities.
DisCo: Distilled Student Models Co-training for Semi-supervised Text Mining (2023.emnlp-main)

Copied to clipboard

Challenge: Existing text mining models are fine-tuned by fine-timing a large pre-trained language model (PLM) in downstream tasks.
Approach: They propose a semi-supervised learning framework for fine-tuning a cohort of small student models generated from a large pre-trained language model using knowledge distillation.
Outcome: The proposed framework outperforms baseline models on semi-supervised text classification and extractive summarization tasks while maintaining comparable performance.
An Open-Source Data Contamination Report for Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing contamination analysis is conducted internally by large language model developers and lacks transparency and completeness.
Approach: They present a data contamination report for 15 popular large language models . they propose an open-source pipeline to perform contamination analysis on customised data .
Outcome: The proposed pipeline enables the community to perform contamination analysis on customised data and models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations