Papers by Teng Lin
Flames: Benchmarking Value Alignment of LLMs in Chinese (2024.naacl-long)
Copied to clipboard
Kexin Huang, Xiangyang Liu, Qianyu Guo, Tianxiang Sun, Jiawei Sun, Yaru Wang, Zeyang Zhou, Yixu Wang, Yan Teng, Xipeng Qiu, Yingchun Wang, Dahua Lin
| Challenge: | Existing benchmarks for large language models (LLMs) do not accurately uncover safety vulnerabilities in LLMs. |
| Approach: | They propose a value alignment benchmark called Flames that encompasses both harmlessness principles and a unique morality dimension that integrates specific Chinese values such as harmony. |
| Outcome: | The proposed model performs poorly on Flames, particularly in safety and fairness dimensions. |
PEMV: Improving Spatial Distribution for Emotion Recognition in Conversations Using Proximal Emotion Mean Vectors (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing research focuses on the analysis of contextual structure in dialogue and the interactions between different emotions. |
| Approach: | They propose a method that generates Proximal Emotion Mean Vectors (PEMVs) based on emotion feature queues to optimize the spatial representation of text features. |
| Outcome: | The proposed method achieves state-of-the-art performance on three widely used benchmark datasets. |
Empowering Healthcare Practitioners with Language Models: Structuring Speech Transcripts in Two Real-World Clinical Applications (2025.emnlp-industry)
Copied to clipboard
Jean-Philippe Corbeil, Asma Ben Abacha, George Michalopoulos, Phillip Swazinna, Miguel Del-Agua, Jerome Tremblay, Akila Jeeson Daniel, Cari Bader, Kevin Cho, Pooja Krishnan, Nathan Bodenstab, Thomas Lin, Wenxuan Teng, Francois Beaulieu, Paul Vozila
| Challenge: | Large language models (LLMs) have demonstrated strong performance on clinical natural language processing tasks across multiple medical benchmarks. |
| Approach: | They propose an agentic pipeline for generating realistic, non-sensitive nurse dictations, enabling structured extraction of clinical observations. |
| Outcome: | The proposed pipeline generates realistic, non-sensitive nurse dictations, enabling structured extraction of clinical observations. |
BLADE: Benchmarking Language Model Agents for Data-Driven Science (2024.findings-emnlp)
Copied to clipboard
Ken Gu, Ruoxi Shang, Ruien Jiang, Keying Kuang, Richard-John Lin, Donghe Lyu, Yue Mao, Youran Pan, Teng Wu, Jiaqian Yu, Yikun Zhang, Tianmai Zhang, Lanyi Zhu, Mike Merrill, Jeffrey Heer, Tim Althoff
| Challenge: | Language model-based agents can be used to conduct and support data-driven science, but evaluating them on open-ended tasks is challenging due to multiple valid approaches, partially correct steps, and different ways to express the same decisions. |
| Approach: | They propose a benchmark to automatically evaluate agents’ multifaceted approaches to open-ended research questions. |
| Outcome: | BLADE evaluates agents’ multifaceted approaches to open-ended research questions using data from 12 datasets and research questions drawn from existing scientific literature. |
Beyond Static Artifacts: An Evolutionary Framework for Synthetic Claim Generation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing claim detection benchmarks treat claims as static textual artifacts . current research ignores sociological etiology of how information naturally emerges and mutates . |
| Approach: | They propose a socially generative framework for synthetic claim generation . they propose utterance, proposition and context-based simulations to capture truth decay . |
| Outcome: | The proposed paradigm models claims as socially evolving entities . it allows precise simulation of truth decay and intervened propagation with multi-auditor oversight . |
DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph Refinement (2025.emnlp-main)
Copied to clipboard
| Challenge: | Current approaches typically merge sentence-level parsing outputs for discourse input, resulting in fragmented graphs and degraded downstream performance. |
| Approach: | They propose a task for discourse-level text scene graph parsing that merges sentence-level outputs for discourse input and propose 'DiscoSG' a dataset of 400 expert-annotated and 8,430 synthesised multi-sentence caption-graph pairs is used to test the new task. |
| Outcome: | The proposed task improves SPICE by 30% over the baseline while achieving 86 faster inference than existing models. |
MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answering (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) and Retrieval-augmented Generation (RAG) systems show promise, but their performance on cross-document MEQA remains underexplored due to the lack of tailored benchmarks. |
| Approach: | They propose a scalable multi-document, multi-entity benchmark to evaluate LLMs' capacity to retrieve, consolidate, and reason over scattered and dense information. |
| Outcome: | The proposed benchmarks show that even advanced models achieve only 59% accuracy on MEBench. |