Papers by Tan Lee

18 papers
Good Examples Make A Faster Learner: Simple Demonstration-based Learning for Low-resource NER (2022.acl-long)

Copied to clipboard

Challenge: Recent advances in prompt-based learning have shown strong results on few-shot text classification by using cloze-style templates.
Approach: They propose a demonstration-based learning method which lets the input be prefaced by task demonstrations for in-context learning.
Outcome: The proposed method improves on in-domain learning and domain adaptation in low-resource settings.
Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural Alignment (2026.acl-long)

Copied to clipboard

Challenge: Current alignment paradigms treat "human values" as a monolithic entity, ignoring the fact that many societies are a mosaic of diverse subgroups with distinct and sometimes conflicting values, preferences, and norms.
Approach: They examine whether Large Language Models can emulate distinct cultural values of subgroups . they use a global value survey to examine the value landscape of a multicultural society .
Outcome: The proposed model improves on unseen, out-of-distribution subgroups by 17.4% . the model widens the disparity between subgroup groups when measured by distance-aware metrics.
PodAgent: A Comprehensive Framework for Podcast Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing automatic audio generation methods struggle to generate podcast-like audio programs effectively.
Approach: They propose a framework for creating podcast-like audio programs that generates informative topic-discussion content by designing a multi-agent collaboration system, builds a voice pool and uses LLM-enhanced speech synthesis to generate expressive conversational speech.
Outcome: The proposed framework surpasses direct GPT-4 generation in topic-discussion dialogue content, and produces more expressive conversational speech.
DALK: Dynamic Co-Augmentation of LLMs and KG to answer Alzheimer’s Disease Questions with Scientific Literature (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in large language models have achieved promising performances across various applications, but the challenge of integrating long-tail knowledge continues to impede the seamless adoption of LLMs in specialized domains.
Approach: They propose a dynamic co-augmentation framework for the refinement of large language models and knowledge graphs in the context of Alzheimer's Disease.
Outcome: The proposed framework can be used to study Alzheimer's Disease (AD) using LLMs and KGs.
Metacognitive Self-Correction for Multi-Agent System via Prototype-Guided Next-Execution Reconstruction (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Model based multi-agent systems (MAS) excel at collaborative problem solving but remain brittle to cascading errors.
Approach: They propose a metacognitive framework that enables step-level error detection and self-correction in Large Language Model based multi-agent systems (MAS) .
Outcome: The proposed framework outperforms baselines on the Who When benchmark and delivers consistent gains on AgentErrorBench.
In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to long-term dialogue memory management fail to capture the natural semantic structure of conversations, leading to fragmented and incomplete representations.
Approach: They propose a mechanism that integrates forward- and backward-looking reflections into a personalized memory bank for effective future retrieval.
Outcome: The proposed mechanism outperforms state-of-the-art benchmarks on a long-term dialogue memory model.
MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing models lack cultural alignment across modalities and languages . a new framework to assess cultural awareness across linguistics and languages is needed .
Approach: They propose a framework that integrates tri-modally aligned cultural benchmarks and a five-dimensional evaluation protocol to assess cross-country awareness disparities.
Outcome: The proposed framework assesses cultural awareness disparities across modalities and languages . it is the first dataset aligned at the input level across text, image, and speech .
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics (2025.findings-emnlp)

Copied to clipboard

Challenge: PixelHumor is a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate LMMs’ ability to interpret multimodal humor and recognize narrative sequences.
Approach: PixelHumor is a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate LMMs’ ability to interpret multimodal humor and recognize narrative sequences.
Outcome: Experiments with state-of-the-art LMMs reveal that top models achieve only 61% accuracy in panel sequencing, far below human performance.
ReMedi: Reasoner for Medical Clinical Prediction (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to predicting future clinical outcomes from EHRs focus on enhancing medical knowledge through distillation or RAG while relying on the model’s internal ability to interpret contextual information.
Approach: They propose a framework for improving clinical outcome prediction from EHR using a sample regeneration mechanism that leverages ground-truth answers as hints to enhance reasoning.
Outcome: Experiments on multiple EHR prediction tasks show significant gains of up to 19.9% over state-of-the-art baselines in terms of F1 score, underscoring ReMedi’s effectiveness in real-world clinical prediction.
Persuasion Dynamics in LLMs: Investigating Robustness and Adaptability in Knowledge and Safety with DuET-PD (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) can struggle to balance gullibility to misinformation and resistance to valid corrections in persuasive dialogues.
Approach: They propose a framework evaluating multi-turn stance-change dynamics across dual dimensions: persuasion type and domain.
Outcome: The proposed framework improves LLM-3.1-8B-Instruct accuracy under misleading persuasion in safety contexts from 4.21% to 76.54%.
Task-Aware Resolution Optimization for Visual Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing visual large language models pre-assume a fixed resolution for downstream tasks, leading to sub-optimal performance.
Approach: They propose a formula to determine the optimal resolution for a given vision-language task . they then propose 'parameter-efficient' fine-tuning technique to extend the visual input resolution .
Outcome: The proposed method is based on rigorous experiments on vision-language tasks.
LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators (2025.emnlp-demos)

Copied to clipboard

Challenge: a new study shows that moderation systems that ignore localisation and low-resource variants risk degraded performance and exploitation in real-world deployments.
Approach: They propose a lightweight, multilingual moderation classifier tailored to Singapore's context . it uses pre-trained OpenAI embeddings and a multi-head ordinal classifier .
Outcome: The proposed classifier outperforms commercial and open-source models across 17 benchmarks.
Case-based Reasoning for Natural Language Queries over Knowledge Bases (2021.emnlp-main)

Copied to clipboard

Challenge: Using human-labeled examples, case-based reasoning can solve complex problems from scratch . case-Based reasoning is a paradigm that is used to solve complex problem .
Approach: They propose a neuro-symbolic CBR approach for question answering over large knowledge bases.
Outcome: The proposed approach outperforms the current state of the art on a CWQ dataset by 11% on accuracy.
BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing evaluations assess static recall or isolated visual grounding, leaving unanswered whether VLMs possess robust and transferable cultural understanding.
Approach: They propose a multimodal, multicultural benchmark to evaluate the robustness of everyday cultural knowledge in vision-language models across linguistic rephrasings and visual modalities.
Outcome: ‘BLEnD-Vis‘ constructs 313 culturally grounded question templates spanning 16 regions and generates three aligned multiple-choice formats.
Unmasking Implicit Bias: Evaluating Persona-Prompted LLM Responses in Power-Disparate Social Scenarios (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable capabilities in simulating human behaviour and social intelligence, but they risk perpetuating societal biases, especially when demographic information is involved.
Approach: They propose a framework that measures semantic shifts in responses and an LLM-judged Preference Win Rate to assess how demographic prompts affect response quality across power-disparate social scenarios.
Outcome: The proposed framework measures semantic shifts in responses and an LLM-judged Preference Win Rate (WR) to assess how demographic prompts affect response quality across power-disparate social scenarios.
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used to judge code, but their reliability remains poorly understood.
Approach: They propose a benchmark to evaluate Large Language Models as code judges . they find that small reasoning models outperform larger non-reasoning models .
Outcome: The proposed benchmark evaluates LLM-as-a-Judge models across three coding tasks.
HABERTOR: An Efficient and Effective Deep Hatespeech Detector (2020.emnlp-main)

Copied to clipboard

Challenge: HABERTOR model is a highly efficient and effective alternative to BERT for the hatespeech classification task.
Approach: They propose to modify BERT's HABERTOR model to generate its own vocabularies and pre-trained it using the largest scale hatespeech dataset.
Outcome: The proposed model is faster, more efficient and more robust than existing methods for hatespeech classification.
MobileQuant: Mobile-friendly Quantization for On-device Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have revolutionized language processing, but deployment on edge devices is costly in terms of memory, computation and energy.
Approach: They propose to reduce the number of bits used to represent weights and activations . they propose to use 8-bit activations to enable LLMs to fully exploit mobile-friendly hardware .
Outcome: The proposed method reduces the number of bits used to represent weights and activations . 8-bit activations are attractive for on-device deployment as they would exploit mobile-friendly hardware .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations