Papers by Chenghao Xiao
On Isotropy, Contextualization and Learning Dynamics of Contrastive-based Sentence Representation Learning (2023.findings-acl)
Copied to clipboard
| Challenge: | Incorporating contrastive learning objectives in sentence representation learning has yielded significant improvements on many sentence-level NLP tasks. |
| Approach: | They aim to examine why contrastive learning works for learning sentence-level semantics . they interpret successes through the geometry of the representation shifts based on isotropy . |
| Outcome: | The proposed model improves on many sentence-level NLP tasks, but it is not well understood why it works. |
Effective Distillation of Table-based Reasoning Ability from LLMs (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing work on table-based reasoning distillation has focused on smaller models with limited performance. |
| Approach: | They propose a table-based reasoning distillation approach to distill LLMs into smaller models . their results show that a 220 million parameter model fine-tuned using distilled data improves performance . |
| Outcome: | The proposed model improves on a scientific table-to-text generation dataset and surpasses specific LLMs. |
On the Rigour of Scientific Writing: Criteria, Analysis, and Insights (2024.findings-emnlp)
Copied to clipboard
| Challenge: | despite its importance, little work exists on modelling rigour in scientific writing . despite widespread use of term, scientific literature lacks definition of rigor . |
| Approach: | They propose a framework to automatically identify and define rigour criteria and assess their relevance in scientific writing. |
| Outcome: | The proposed framework can be tailored to the evaluation of scientific rigour for different areas. |
ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning (2025.emnlp-main)
Copied to clipboard
Yu Sun, Xingyu Qian, Weiwen Xu, Hao Zhang, Chenghao Xiao, Long Li, Deli Zhao, Wenbing Huang, Tingyang Xu, Qifeng Bai, Yu Rong
| Challenge: | Existing medical reasoning datasets are limited in scale and typically rely on incomplete data. |
| Approach: | They propose to use ReasonMed to train medical reasoning models using a multi-agent generation, verification, and refinement pipeline. |
| Outcome: | The largest medical reasoning dataset to date surpasses the prior best sub-10B models by 4.17% and even exceeds LLaMA3.1-70B on PubMedQA by 4.60%. |
AutoScraper: A Progressive Understanding Web Agent for Web Scraper Generation (2024.emnlp-main)
Copied to clipboard
Wenhao Huang, Zhouhong Gu, Chenghao Peng, Jiaqing Liang, Zhixu Li, Yanghua Xiao, Liqian Wen, Zulong Chen
| Challenge: | Existing methods for web scraping suffer from limited adaptability and scalability when faced with a new website. |
| Approach: | They propose a framework that generates web scrapers with large language models and a new executability metric to measure the performance of web scraper generation tasks. |
| Outcome: | The proposed framework can handle diverse web environments more efficiently. |
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation (2026.findings-acl)
Copied to clipboard
| Challenge: | Scientific rigour tends to be sidelined in favour of bold statements, leading authors to overstate claims beyond what their results support. |
| Approach: | They propose a multimodal framework that retrieves supporting evidence from a paper and assigns each claim an overstatement score. |
| Outcome: | The proposed framework retrieves supporting evidence from ICLR and NeurIPS papers and assigns each claim an overstatement score. |
Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages (2026.acl-short)
Copied to clipboard
| Challenge: | Existing studies show that performance across low-resource settings is variable, resulting in a significant barrier for the MT community. |
| Approach: | They propose to use FRED Difficulty Metrics to contextualize reported performance across different language pairs to determine whether breakthroughs reported in other contexts are artifacts of benchmark collection. |
| Outcome: | The proposed metrics explain a significant portion of result variability rather than model capability. |
Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth (2025.emnlp-main)
Copied to clipboard
| Challenge: | Despite excelling at many natural language processing tasks, large language models fail to grasp the layered semantics of Drivelological text. |
| Approach: | They construct a benchmark dataset of over 1,200+ carefully curated and diverse examples across English, Mandarin, Spanish, French, Japanese, and Korean to examine their Drivelological characteristics. |
| Outcome: | The proposed models lack conceptual understanding and lack conceptual and semantic accuracy. |
Analyzing LLMs’ Knowledge Boundary Cognition Across Languages Through the Lens of Internal Representations (2025.acl-long)
Copied to clipboard
| Challenge: | Understanding the knowledge boundaries of Large Language Models (LLMs) is crucial to prevent hallucination, but research on the knowledge boundary perceptions of LLMs has predominantly focused on English. |
| Approach: | They propose a training-free alignment method that effectively transfers knowledge boundary perception ability across languages, thereby helping reduce hallucination risk in low-resource languages. |
| Outcome: | The proposed method reduces hallucination risk in low-resource languages by fine-tuning on bilingual question pair translation. |
Length is a Curse and a Blessing for Document-level Semantics (2023.emnlp-main)
Copied to clipboard
| Challenge: | In recent years, contrastive learning (CL) has been extensively utilized to recover sentence and document-level encoding capability from pre-trained language models. |
| Approach: | They propose a document-based contrastive learning framework that is length-agnostic self-reference based on document length. |
| Outcome: | The proposed framework achieves state-of-the-art on the standard information retrieval benchmark. |
Understanding the Behaviors of Environment-aware Information Retrieval (2026.acl-long)
Copied to clipboard
| Challenge: | Recent retrieval-augmented generation approaches have demonstrated strong capability in handling complex queries. |
| Approach: | They propose a branching-based rollout technique that improves training stability . they find different retrievers exhibit distinct optimal query styles . |
| Outcome: | The proposed method improves training stability and improves retrieval-aware systems. |
CAST: Corpus-Aware Self-similarity Enhanced Topic modelling (2025.naacl-long)
Copied to clipboard
Yanan Ma, Chenghao Xiao, Chenhan Yuan, Sabine N Van Der Veer, Lamiece Hassan, Chenghua Lin, Goran Nenadic
| Challenge: | Existing topic modelling methods encode contextual information of documents while ignoring contextual details of candidate centroid words. Existing methods are limited by the contextualization gap. |
| Approach: | They propose a topic modelling method that builds upon candidate centroid word embeddings contextualized on the dataset and a self-similarity-based method to filter out less meaningful tokens. |
| Outcome: | The proposed method significantly enhances the coherence and diversity of generated topics, and handles noisy data, outperforming strong baselines. |
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval (2024.findings-acl)
Copied to clipboard
Siwei Wu, Yizhi Li, Kang Zhu, Ge Zhang, Yiming Liang, Kaijing Ma, Chenghao Xiao, Haoran Zhang, Bohao Yang, Wenhu Chen, Wenhao Huang, Noura Al Moubayed, Jie Fu, Chenghua Lin
| Challenge: | Multi-modal information retrieval (MMIR) is a rapidly evolving field . current benchmarks for image-text pairings overlook the scientific domain . |
| Approach: | They develop a scientific domain-specific MMIR benchmark to evaluate image-text pairings using open-access research paper corpora. |
| Outcome: | The proposed benchmarks are based on 530K image-text pairs extracted from scientific documents with detailed captions. |
Crafting Customisable Characters with LLMs: A Persona-Driven Role-Playing Agent Framework (2025.findings-emnlp)
Copied to clipboard
Bohao Yang, Dong Liu, Chenghao Xiao, Kun Zhao, Chen Tang, Chao Li, Lin Yuan, Yang Guang, Chenghua Lin
| Challenge: | Large Language Models (LLMs) are capable of generating human-like text, but the potential for freely customisable characters remains underexplored. |
| Approach: | They propose a framework which employs Large Language Models to create freely customisable characters through personalised characteristic feature injection. |
| Outcome: | The proposed framework provides valuable insights for developing more accurate and customisable human simulacra. |
X-ray Made Simple: Lay Radiology Report Generation and Robust Evaluation (2026.findings-acl)
Copied to clipboard
Kun Zhao, Chenghao Xiao, Sixing Yan, Haoteng Tang, William K. Cheung, Noura Al Moubayed, Liang Zhan, Chenghua Lin
| Challenge: | Technical language and templated nature of professional reports hinder patient comprehension and allow models to artificially boost lexical metrics such as BLEU by reproducing common report patterns. |
| Approach: | They propose a layman's RRG framework that leverages layperson-friendly language to enhance patient accessibility and promote robust evaluation and report generation by encouraging models to focus on semantic accuracy over rigid templates. |
| Outcome: | The proposed framework improves model performance with more layman-style data, compared to templated professional language and inflated lexical scores. |
Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts (2026.tacl-1)
Copied to clipboard
| Challenge: | Evaluating natural language generation systems is challenging due to the diversity of valid outputs. |
| Approach: | They propose an inversion learning method that learns effective reverse mappings from model outputs back to their input instructions. |
| Outcome: | The proposed method requires only a single evaluation sample and eliminates manual prompt engineering. |
Modeling Semantic Compositionality with Sememe Knowledge (P19-1)
Copied to clipboard
| Challenge: | Semantic compositionality (SC) is defined as the phenomenon that the meaning of a complex linguistic unit can be composed of the meanings of its constituents. |
| Approach: | They propose to incorporate sememes into SC models and employ them in learning multiword expressions. |
| Outcome: | The proposed models achieve significant performance boost compared to baseline methods without sememe knowledge. |